An agent swarm coordinates agents toward a shared task. A fleet can contain agents doing separate jobs. Some public research accounts report thousands of simultaneous workers; organizational adoption figures often count something else.
Browse the dated first-party reports below, checked by Codex (AI research agent) on September 8, 2026. Each preserves the unit, reporting window, outcome and limits. Research runs, deployed systems and product capacities are labeled separately. This selected record is not an industry-wide adoption survey.
15 of 15 observations
Agents working at once
Concurrency counts active workers at the same time. A reported capacity is a supported ceiling; a peak says how high one run reached. Neither establishes the average or the useful work produced.
| Organization and task | Reported scale | Result and limits |
|---|---|---|
| AnthropicC compiler experimentSoftware development · Research experiment | 16 parallel agents Claude instances working on one shared compiler project. Two-week experiment reported February 2026 Unit: concurrent agents · Reported run | Produced a compiler that built Linux 6.9 for three architectures. Limits: Human-designed tests and interventions shaped the work. GCC handled the x86 16-bit boot phase; the compiler remained incomplete. Anthropic · 2026-02-05 |
| CursorBrowser research harnessSoftware development · Research experiment | Hundreds of concurrent workers Workers contributing to the same experimental browser codebase. Nearly one week, reported January 2026 Unit: concurrent agents · Reported run | The team reported more than a million lines across roughly a thousand files. Limits: No exact concurrent count in this account. Code volume does not establish browser completeness or production readiness. Cursor · 2026-01-14 |
| Moonshot AIKimi K2.5 Agent SwarmResearch and office work · Research experiment | Up to 100 subagents Announced parallel swarm capacity, with up to 1,500 tool calls. K2.5 research preview; blog publication date not displayed Unit: subagent capacity · Product capacity | Moonshot reports execution up to 4.5 times faster than its single-agent baseline. Limits: Capacity and benchmark claims. This source does not establish a customer run with 100 simultaneous agents or disclose typical utilization. Moonshot AI · Publication date not stated |
| OpenAINavier–Stokes research agentsMathematics · Research experiment | Approximately 10,000 concurrent agents The group credited with the proposed proof; other groups explored related problems. September 1–5, 2026; 88 hours from initial launches Unit: concurrent agents · Reported run | OpenAI reports 17 additional hours for formalization and verification in Lean using GPT-6 Astra. Limits: Not independent mathematical acceptance. People redirected work and consolidated findings. Sustained concurrency and cost were not disclosed. OpenAI · 2026-09-08 |
Anthropic C compiler experiment
Software development · Research experiment
16 parallel agentsClaude instances working on one shared compiler project.
Two-week experiment reported February 2026
Unit: concurrent agents · Reported run
Produced a compiler that built Linux 6.9 for three architectures.
Limits: Human-designed tests and interventions shaped the work. GCC handled the x86 16-bit boot phase; the compiler remained incomplete.
Anthropic · 2026-02-05
Checked 2026-09-08 by Codex (AI research agent).Cursor Browser research harness
Software development · Research experiment
Hundreds of concurrent workersWorkers contributing to the same experimental browser codebase.
Nearly one week, reported January 2026
Unit: concurrent agents · Reported run
The team reported more than a million lines across roughly a thousand files.
Limits: No exact concurrent count in this account. Code volume does not establish browser completeness or production readiness.
Cursor · 2026-01-14
Checked 2026-09-08 by Codex (AI research agent).Moonshot AI Kimi K2.5 Agent Swarm
Research and office work · Research experiment
Up to 100 subagentsAnnounced parallel swarm capacity, with up to 1,500 tool calls.
K2.5 research preview; blog publication date not displayed
Unit: subagent capacity · Product capacity
Moonshot reports execution up to 4.5 times faster than its single-agent baseline.
Limits: Capacity and benchmark claims. This source does not establish a customer run with 100 simultaneous agents or disclose typical utilization.
Moonshot AI · Publication date not stated
Checked 2026-09-08 by Codex (AI research agent).OpenAI Navier–Stokes research agents
Mathematics · Research experiment
Approximately 10,000 concurrent agentsThe group credited with the proposed proof; other groups explored related problems.
September 1–5, 2026; 88 hours from initial launches
Unit: concurrent agents · Reported run
OpenAI reports 17 additional hours for formalization and verification in Lean using GPT-6 Astra.
Limits: Not independent mathematical acceptance. People redirected work and consolidated findings. Sustained concurrency and cost were not disclosed.
OpenAI · 2026-09-08
Checked 2026-09-08 by Codex (AI research agent).
Sessions and search attempts
A session or rollout is one stretch of agent work. A candidate program is one possible solution. These counts accumulate across a project, including retries and rejected work; compare the units before comparing the numbers.
| Organization and task | Reported scale | Result and limits |
|---|---|---|
| AnthropicC compiler experimentSoftware development · Research experiment | Nearly 2,000 sessions Successive sessions across the 16-worker project. Two weeks Unit: Claude Code sessions · Reported run | Reported consumption: 2 billion input tokens, 140 million output tokens, and just under $20,000 in API charges. Limits: Sessions accumulated over time. API charges exclude the researcher's labor and do not establish total project cost. Anthropic · 2026-02-05 |
| Edison Scientific / FutureHouseKosmosScientific research · Research experiment | Over 200 agent rollouts Literature search and data analysis coordinated through a shared world model. Per research run, up to 12 hours Unit: agent rollouts · Reported run | Average run: 1,500 papers read and 42,000 lines of analysis code executed. Scientist evaluation judged 79.4% of report statements accurate. Limits: Rollouts are cumulative. Concurrent worker count is not reported. The six-month human-equivalent claim is collaborators' estimation, not a controlled comparison. Edison Scientific and collaborators · 2025-11-04 |
| KlarnaAlphaEvolveFinancial services · Pilot | Nearly 6,000 candidate programs Search over code for one model's training pipeline, with reproducibility constraints. Three-week training optimization project Unit: candidate programs · Reported run | Training throughput rose from 49 to roughly 97 samples per second under those constraints. Limits: Candidate count is not agent concurrency. A faster, nondeterministic candidate was rejected; one experiment continued through 631 evaluations without improvement. Deployment coverage and search cost were not quantified. Klarna Engineering · 2026-03-30 |
Anthropic C compiler experiment
Software development · Research experiment
Nearly 2,000 sessionsSuccessive sessions across the 16-worker project.
Two weeks
Unit: Claude Code sessions · Reported run
Reported consumption: 2 billion input tokens, 140 million output tokens, and just under $20,000 in API charges.
Limits: Sessions accumulated over time. API charges exclude the researcher's labor and do not establish total project cost.
Anthropic · 2026-02-05
Checked 2026-09-08 by Codex (AI research agent).Edison Scientific / FutureHouse Kosmos
Scientific research · Research experiment
Over 200 agent rolloutsLiterature search and data analysis coordinated through a shared world model.
Per research run, up to 12 hours
Unit: agent rollouts · Reported run
Average run: 1,500 papers read and 42,000 lines of analysis code executed. Scientist evaluation judged 79.4% of report statements accurate.
Limits: Rollouts are cumulative. Concurrent worker count is not reported. The six-month human-equivalent claim is collaborators' estimation, not a controlled comparison.
Edison Scientific and collaborators · 2025-11-04
Checked 2026-09-08 by Codex (AI research agent).Klarna AlphaEvolve
Financial services · Pilot
Nearly 6,000 candidate programsSearch over code for one model's training pipeline, with reproducibility constraints.
Three-week training optimization project
Unit: candidate programs · Reported run
Training throughput rose from 49 to roughly 97 samples per second under those constraints.
Limits: Candidate count is not agent concurrency. A faster, nondeterministic candidate was rejected; one experiment continued through 631 evaluations without improvement. Deployment coverage and search cost were not quantified.
Klarna Engineering · 2026-03-30
Checked 2026-09-08 by Codex (AI research agent).
Output over time
Throughput connects output to a time window. Commits, proposed changes, and merged changes represent different stages of delivery. A peak hourly rate cannot stand in for a sustained rate.
| Organization and task | Reported scale | Result and limits |
|---|---|---|
| CursorBrowser research harnessSoftware development · Research experiment | Approximately 1,000 commits/hour at peak Research harness also accumulated 10 million tool calls during the week. Peak within a one-week run Unit: commits per hour · Reported run | Agents repeatedly contributed changes to an experimental browser. Limits: The system tolerated temporary errors. Commit rate does not measure accepted features, defect-free releases, or sustained throughput. Cursor · 2026-02-05 |
Cursor Browser research harness
Software development · Research experiment
Approximately 1,000 commits/hour at peakResearch harness also accumulated 10 million tool calls during the week.
Peak within a one-week run
Unit: commits per hour · Reported run
Agents repeatedly contributed changes to an experimental browser.
Limits: The system tolerated temporary errors. Commit rate does not measure accepted features, defect-free releases, or sustained throughput.
Cursor · 2026-02-05
Checked 2026-09-08 by Codex (AI research agent).
Use across an organization
Deployed systems and accepted work show where agents have entered an organization. They do not reveal how many agents collaborate on any one task or what share of all suitable work is automated.
| Organization and task | Reported scale | Result and limits |
|---|---|---|
| Cognition customersDevinSoftware development across industries · Operating use | Hundreds of thousands of merged PRs Vendor aggregate across thousands of customer companies. First 18 months after launch, reported November 2025 Unit: merged pull requests · Organization total | Documents accepted code changes across the customer base. Limits: Cognition also reports a 67% merge rate, without underlying attempt counts or a clear cohort window. Shared-task concurrency and production outcomes are not established. Cognition · 2025-11-14 |
| VerceleveBusiness operations · Operating use | More than 100 production agents Distinct agents serving roles across Vercel's business. Internal fleet described June 2026 Unit: deployed agents · Organization total | The source describes recurring use in analysis, support, sales, and content work. Limits: Deployed agents can have separate tasks and schedules. The figure does not establish simultaneous activity or a hundred-agent swarm on one task. Vercel · 2026-06-17 |
| BNYEliza digital employeesFinancial services · Operating use | Approximately 140 digital employees BNY's term for multi-agent systems working alongside human colleagues. First-quarter 2026 results Unit: deployed multi-agent solutions · Organization total | A deployment snapshot within the bank's AI program, reported separately from roughly 220 enterprise AI solutions in production. Limits: Each digital employee can contain multiple agents. This is neither a concurrent-worker count nor a measure of completed work per system. BNY · 2026-04-16 |
| BMWCodeRabbit code reviewAutomotive software · Operating use | More than 1,000 developers supported BMW's deployed source-code review system, supporting vehicle software development. August 2026; collaboration began more than two years earlier Unit: developers supported · Organization total | Reviews proposed changes for developers deciding what is ready to ship. Limits: Support coverage is not active-use frequency, agent count, or an end-to-end software factory. BMW-specific review volume and defect rates are not disclosed. BMW Group · 2026-08-12 |
| ItaúDevinFinancial services · Operating use | 75% of teams use Devin Vendor-hosted account featuring Itaú technology leaders. Undated customer account; measurement period not stated Unit: share of teams · Organization total | Documents organizational adoption; the account describes code maintenance and repair workflows. Limits: The total number of teams, definition of use, and reporting window are missing. It does not mean 75% of work is automated or reveal concurrency. Cognition / Devin · Publication date not stated |
Cognition customers Devin
Software development across industries · Operating use
Hundreds of thousands of merged PRsVendor aggregate across thousands of customer companies.
First 18 months after launch, reported November 2025
Unit: merged pull requests · Organization total
Documents accepted code changes across the customer base.
Limits: Cognition also reports a 67% merge rate, without underlying attempt counts or a clear cohort window. Shared-task concurrency and production outcomes are not established.
Cognition · 2025-11-14
Checked 2026-09-08 by Codex (AI research agent).Vercel eve
Business operations · Operating use
More than 100 production agentsDistinct agents serving roles across Vercel's business.
Internal fleet described June 2026
Unit: deployed agents · Organization total
The source describes recurring use in analysis, support, sales, and content work.
Limits: Deployed agents can have separate tasks and schedules. The figure does not establish simultaneous activity or a hundred-agent swarm on one task.
Vercel · 2026-06-17
Checked 2026-09-08 by Codex (AI research agent).BNY Eliza digital employees
Financial services · Operating use
Approximately 140 digital employeesBNY's term for multi-agent systems working alongside human colleagues.
First-quarter 2026 results
Unit: deployed multi-agent solutions · Organization total
A deployment snapshot within the bank's AI program, reported separately from roughly 220 enterprise AI solutions in production.
Limits: Each digital employee can contain multiple agents. This is neither a concurrent-worker count nor a measure of completed work per system.
BNY · 2026-04-16
Checked 2026-09-08 by Codex (AI research agent).BMW CodeRabbit code review
Automotive software · Operating use
More than 1,000 developers supportedBMW's deployed source-code review system, supporting vehicle software development.
August 2026; collaboration began more than two years earlier
Unit: developers supported · Organization total
Reviews proposed changes for developers deciding what is ready to ship.
Limits: Support coverage is not active-use frequency, agent count, or an end-to-end software factory. BMW-specific review volume and defect rates are not disclosed.
BMW Group · 2026-08-12
Checked 2026-09-08 by Codex (AI research agent).Itaú Devin
Financial services · Operating use
75% of teams use DevinVendor-hosted account featuring Itaú technology leaders.
Undated customer account; measurement period not stated
Unit: share of teams · Organization total
Documents organizational adoption; the account describes code maintenance and repair workflows.
Limits: The total number of teams, definition of use, and reporting window are missing. It does not mean 75% of work is automated or reveal concurrency.
Cognition / Devin · Publication date not stated
Checked 2026-09-08 by Codex (AI research agent).
Runtime and inference
Tokens measure model input or output; runtime measures elapsed agent activity. Track both with the task and outcome. Different models, tools, and accounting methods can produce very different totals for similar work.
| Organization and task | Reported scale | Result and limits |
|---|---|---|
| OpenAINavier–Stokes research agentsMathematics · Research experiment | Approximately 130 billion output tokens The Navier–Stokes effort, which also produced 2.7 million messages. No per-group breakdown is given. September 2026 Navier–Stokes effort Unit: output tokens · Reported run | Cumulative activity associated with the proposed proof. Limits: The wider multi-problem campaign used about 300 billion output tokens. Neither total measures discoveries or simultaneous agents; cost is undisclosed. OpenAI · 2026-09-08 |
| OpenAIInternal coding agentsAI research · Operating use | 3.1 agent-workdays per human workday Aggregate agent runtime normalized to an eight-hour workday. Research organization, mid-August 2026 Unit: runtime ratio · Organization total | Shows how much agent activity had entered daily research work. Limits: Runtime is not equivalent productive labor. Coverage is incomplete. Separately, January–July task analysis found interventions in more than half of successful tasks estimated at 4–8 human hours. OpenAI · 2026-09-06 |
OpenAI Navier–Stokes research agents
Mathematics · Research experiment
Approximately 130 billion output tokensThe Navier–Stokes effort, which also produced 2.7 million messages. No per-group breakdown is given.
September 2026 Navier–Stokes effort
Unit: output tokens · Reported run
Cumulative activity associated with the proposed proof.
Limits: The wider multi-problem campaign used about 300 billion output tokens. Neither total measures discoveries or simultaneous agents; cost is undisclosed.
OpenAI · 2026-09-08
Checked 2026-09-08 by Codex (AI research agent).OpenAI Internal coding agents
AI research · Operating use
3.1 agent-workdays per human workdayAggregate agent runtime normalized to an eight-hour workday.
Research organization, mid-August 2026
Unit: runtime ratio · Organization total
Shows how much agent activity had entered daily research work.
Limits: Runtime is not equivalent productive labor. Coverage is incomplete. Separately, January–July task analysis found interventions in more than half of successful tasks estimated at 4–8 human hours.
OpenAI · 2026-09-06
Checked 2026-09-08 by Codex (AI research agent).
How to compare a new claim
Start with the unit and the time window. Then ask what entered the system, what counted as accepted work, and how much human effort was needed. Keep failed attempts, repair, and review in the calculation. “Not publicly stated” is a useful answer when the source omits a field.
For a claim about concurrency, look for average utilization as well as a peak. For a claim about adoption, look for the denominator: eligible teams, suitable tasks, or total accepted changes. For a speed claim, keep the quality and reproducibility requirements fixed. A run that finds more candidates may still spend most of its time revisiting weak ideas.
This collection is a selected public record. Companies disclose different things, and some disclose very little. Absence here does not establish absence of use. Old observations remain useful when their measurement window stays visible; a newly checked source does not turn an old figure into a current one.
The software factory metrics guide explains how to measure useful output, review effort, and total cost within your own team. The company cases explain the workflows behind the numbers.