An agent swarm coordinates agents toward a shared task. A fleet can contain agents doing separate jobs. Some public research accounts report thousands of simultaneous workers; organizational adoption figures often count something else.

Browse the dated first-party reports below, checked by Codex (AI research agent) on September 8, 2026. Each preserves the unit, reporting window, outcome and limits. Research runs, deployed systems and product capacities are labeled separately. This selected record is not an industry-wide adoption survey.

Agents working at once

Concurrency counts active workers at the same time. A reported capacity is a supported ceiling; a peak says how high one run reached. Neither establishes the average or the useful work produced.

Agents working at once: dated first-party reports
Organization and taskReported scaleResult and limits
AnthropicC compiler experimentSoftware development · Research experiment16 parallel agents

Claude instances working on one shared compiler project.

Two-week experiment reported February 2026

Unit: concurrent agents · Reported run

Produced a compiler that built Linux 6.9 for three architectures.

Limits: Human-designed tests and interventions shaped the work. GCC handled the x86 16-bit boot phase; the compiler remained incomplete.

Anthropic · 2026-02-05
Checked 2026-09-08 by Codex (AI research agent).

CursorBrowser research harnessSoftware development · Research experimentHundreds of concurrent workers

Workers contributing to the same experimental browser codebase.

Nearly one week, reported January 2026

Unit: concurrent agents · Reported run

The team reported more than a million lines across roughly a thousand files.

Limits: No exact concurrent count in this account. Code volume does not establish browser completeness or production readiness.

Cursor · 2026-01-14
Checked 2026-09-08 by Codex (AI research agent).

Moonshot AIKimi K2.5 Agent SwarmResearch and office work · Research experimentUp to 100 subagents

Announced parallel swarm capacity, with up to 1,500 tool calls.

K2.5 research preview; blog publication date not displayed

Unit: subagent capacity · Product capacity

Moonshot reports execution up to 4.5 times faster than its single-agent baseline.

Limits: Capacity and benchmark claims. This source does not establish a customer run with 100 simultaneous agents or disclose typical utilization.

Moonshot AI · Publication date not stated
Checked 2026-09-08 by Codex (AI research agent).

OpenAINavier–Stokes research agentsMathematics · Research experimentApproximately 10,000 concurrent agents

The group credited with the proposed proof; other groups explored related problems.

September 1–5, 2026; 88 hours from initial launches

Unit: concurrent agents · Reported run

OpenAI reports 17 additional hours for formalization and verification in Lean using GPT-6 Astra.

Limits: Not independent mathematical acceptance. People redirected work and consolidated findings. Sustained concurrency and cost were not disclosed.

OpenAI · 2026-09-08
Checked 2026-09-08 by Codex (AI research agent).

  1. Anthropic C compiler experiment

    Software development · Research experiment

    16 parallel agents

    Claude instances working on one shared compiler project.

    Two-week experiment reported February 2026

    Unit: concurrent agents · Reported run

    Produced a compiler that built Linux 6.9 for three architectures.

    Limits: Human-designed tests and interventions shaped the work. GCC handled the x86 16-bit boot phase; the compiler remained incomplete.

    Anthropic · 2026-02-05
    Checked 2026-09-08 by Codex (AI research agent).

  2. Cursor Browser research harness

    Software development · Research experiment

    Hundreds of concurrent workers

    Workers contributing to the same experimental browser codebase.

    Nearly one week, reported January 2026

    Unit: concurrent agents · Reported run

    The team reported more than a million lines across roughly a thousand files.

    Limits: No exact concurrent count in this account. Code volume does not establish browser completeness or production readiness.

    Cursor · 2026-01-14
    Checked 2026-09-08 by Codex (AI research agent).

  3. Moonshot AI Kimi K2.5 Agent Swarm

    Research and office work · Research experiment

    Up to 100 subagents

    Announced parallel swarm capacity, with up to 1,500 tool calls.

    K2.5 research preview; blog publication date not displayed

    Unit: subagent capacity · Product capacity

    Moonshot reports execution up to 4.5 times faster than its single-agent baseline.

    Limits: Capacity and benchmark claims. This source does not establish a customer run with 100 simultaneous agents or disclose typical utilization.

    Moonshot AI · Publication date not stated
    Checked 2026-09-08 by Codex (AI research agent).

  4. OpenAI Navier–Stokes research agents

    Mathematics · Research experiment

    Approximately 10,000 concurrent agents

    The group credited with the proposed proof; other groups explored related problems.

    September 1–5, 2026; 88 hours from initial launches

    Unit: concurrent agents · Reported run

    OpenAI reports 17 additional hours for formalization and verification in Lean using GPT-6 Astra.

    Limits: Not independent mathematical acceptance. People redirected work and consolidated findings. Sustained concurrency and cost were not disclosed.

    OpenAI · 2026-09-08
    Checked 2026-09-08 by Codex (AI research agent).

Sessions and search attempts

A session or rollout is one stretch of agent work. A candidate program is one possible solution. These counts accumulate across a project, including retries and rejected work; compare the units before comparing the numbers.

Sessions and search attempts: dated first-party reports
Organization and taskReported scaleResult and limits
AnthropicC compiler experimentSoftware development · Research experimentNearly 2,000 sessions

Successive sessions across the 16-worker project.

Two weeks

Unit: Claude Code sessions · Reported run

Reported consumption: 2 billion input tokens, 140 million output tokens, and just under $20,000 in API charges.

Limits: Sessions accumulated over time. API charges exclude the researcher's labor and do not establish total project cost.

Anthropic · 2026-02-05
Checked 2026-09-08 by Codex (AI research agent).

Edison Scientific / FutureHouseKosmosScientific research · Research experimentOver 200 agent rollouts

Literature search and data analysis coordinated through a shared world model.

Per research run, up to 12 hours

Unit: agent rollouts · Reported run

Average run: 1,500 papers read and 42,000 lines of analysis code executed. Scientist evaluation judged 79.4% of report statements accurate.

Limits: Rollouts are cumulative. Concurrent worker count is not reported. The six-month human-equivalent claim is collaborators' estimation, not a controlled comparison.

Edison Scientific and collaborators · 2025-11-04
Checked 2026-09-08 by Codex (AI research agent).

KlarnaAlphaEvolveFinancial services · PilotNearly 6,000 candidate programs

Search over code for one model's training pipeline, with reproducibility constraints.

Three-week training optimization project

Unit: candidate programs · Reported run

Training throughput rose from 49 to roughly 97 samples per second under those constraints.

Limits: Candidate count is not agent concurrency. A faster, nondeterministic candidate was rejected; one experiment continued through 631 evaluations without improvement. Deployment coverage and search cost were not quantified.

Klarna Engineering · 2026-03-30
Checked 2026-09-08 by Codex (AI research agent).

  1. Anthropic C compiler experiment

    Software development · Research experiment

    Nearly 2,000 sessions

    Successive sessions across the 16-worker project.

    Two weeks

    Unit: Claude Code sessions · Reported run

    Reported consumption: 2 billion input tokens, 140 million output tokens, and just under $20,000 in API charges.

    Limits: Sessions accumulated over time. API charges exclude the researcher's labor and do not establish total project cost.

    Anthropic · 2026-02-05
    Checked 2026-09-08 by Codex (AI research agent).

  2. Edison Scientific / FutureHouse Kosmos

    Scientific research · Research experiment

    Over 200 agent rollouts

    Literature search and data analysis coordinated through a shared world model.

    Per research run, up to 12 hours

    Unit: agent rollouts · Reported run

    Average run: 1,500 papers read and 42,000 lines of analysis code executed. Scientist evaluation judged 79.4% of report statements accurate.

    Limits: Rollouts are cumulative. Concurrent worker count is not reported. The six-month human-equivalent claim is collaborators' estimation, not a controlled comparison.

    Edison Scientific and collaborators · 2025-11-04
    Checked 2026-09-08 by Codex (AI research agent).

  3. Klarna AlphaEvolve

    Financial services · Pilot

    Nearly 6,000 candidate programs

    Search over code for one model's training pipeline, with reproducibility constraints.

    Three-week training optimization project

    Unit: candidate programs · Reported run

    Training throughput rose from 49 to roughly 97 samples per second under those constraints.

    Limits: Candidate count is not agent concurrency. A faster, nondeterministic candidate was rejected; one experiment continued through 631 evaluations without improvement. Deployment coverage and search cost were not quantified.

    Klarna Engineering · 2026-03-30
    Checked 2026-09-08 by Codex (AI research agent).

Output over time

Throughput connects output to a time window. Commits, proposed changes, and merged changes represent different stages of delivery. A peak hourly rate cannot stand in for a sustained rate.

Output over time: dated first-party reports
Organization and taskReported scaleResult and limits
CursorBrowser research harnessSoftware development · Research experimentApproximately 1,000 commits/hour at peak

Research harness also accumulated 10 million tool calls during the week.

Peak within a one-week run

Unit: commits per hour · Reported run

Agents repeatedly contributed changes to an experimental browser.

Limits: The system tolerated temporary errors. Commit rate does not measure accepted features, defect-free releases, or sustained throughput.

Cursor · 2026-02-05
Checked 2026-09-08 by Codex (AI research agent).

  1. Cursor Browser research harness

    Software development · Research experiment

    Approximately 1,000 commits/hour at peak

    Research harness also accumulated 10 million tool calls during the week.

    Peak within a one-week run

    Unit: commits per hour · Reported run

    Agents repeatedly contributed changes to an experimental browser.

    Limits: The system tolerated temporary errors. Commit rate does not measure accepted features, defect-free releases, or sustained throughput.

    Cursor · 2026-02-05
    Checked 2026-09-08 by Codex (AI research agent).

Use across an organization

Deployed systems and accepted work show where agents have entered an organization. They do not reveal how many agents collaborate on any one task or what share of all suitable work is automated.

Use across an organization: dated first-party reports
Organization and taskReported scaleResult and limits
Cognition customersDevinSoftware development across industries · Operating useHundreds of thousands of merged PRs

Vendor aggregate across thousands of customer companies.

First 18 months after launch, reported November 2025

Unit: merged pull requests · Organization total

Documents accepted code changes across the customer base.

Limits: Cognition also reports a 67% merge rate, without underlying attempt counts or a clear cohort window. Shared-task concurrency and production outcomes are not established.

Cognition · 2025-11-14
Checked 2026-09-08 by Codex (AI research agent).

VerceleveBusiness operations · Operating useMore than 100 production agents

Distinct agents serving roles across Vercel's business.

Internal fleet described June 2026

Unit: deployed agents · Organization total

The source describes recurring use in analysis, support, sales, and content work.

Limits: Deployed agents can have separate tasks and schedules. The figure does not establish simultaneous activity or a hundred-agent swarm on one task.

Vercel · 2026-06-17
Checked 2026-09-08 by Codex (AI research agent).

BNYEliza digital employeesFinancial services · Operating useApproximately 140 digital employees

BNY's term for multi-agent systems working alongside human colleagues.

First-quarter 2026 results

Unit: deployed multi-agent solutions · Organization total

A deployment snapshot within the bank's AI program, reported separately from roughly 220 enterprise AI solutions in production.

Limits: Each digital employee can contain multiple agents. This is neither a concurrent-worker count nor a measure of completed work per system.

BNY · 2026-04-16
Checked 2026-09-08 by Codex (AI research agent).

BMWCodeRabbit code reviewAutomotive software · Operating useMore than 1,000 developers supported

BMW's deployed source-code review system, supporting vehicle software development.

August 2026; collaboration began more than two years earlier

Unit: developers supported · Organization total

Reviews proposed changes for developers deciding what is ready to ship.

Limits: Support coverage is not active-use frequency, agent count, or an end-to-end software factory. BMW-specific review volume and defect rates are not disclosed.

BMW Group · 2026-08-12
Checked 2026-09-08 by Codex (AI research agent).

ItaúDevinFinancial services · Operating use75% of teams use Devin

Vendor-hosted account featuring Itaú technology leaders.

Undated customer account; measurement period not stated

Unit: share of teams · Organization total

Documents organizational adoption; the account describes code maintenance and repair workflows.

Limits: The total number of teams, definition of use, and reporting window are missing. It does not mean 75% of work is automated or reveal concurrency.

Cognition / Devin · Publication date not stated
Checked 2026-09-08 by Codex (AI research agent).

  1. Cognition customers Devin

    Software development across industries · Operating use

    Hundreds of thousands of merged PRs

    Vendor aggregate across thousands of customer companies.

    First 18 months after launch, reported November 2025

    Unit: merged pull requests · Organization total

    Documents accepted code changes across the customer base.

    Limits: Cognition also reports a 67% merge rate, without underlying attempt counts or a clear cohort window. Shared-task concurrency and production outcomes are not established.

    Cognition · 2025-11-14
    Checked 2026-09-08 by Codex (AI research agent).

  2. Vercel eve

    Business operations · Operating use

    More than 100 production agents

    Distinct agents serving roles across Vercel's business.

    Internal fleet described June 2026

    Unit: deployed agents · Organization total

    The source describes recurring use in analysis, support, sales, and content work.

    Limits: Deployed agents can have separate tasks and schedules. The figure does not establish simultaneous activity or a hundred-agent swarm on one task.

    Vercel · 2026-06-17
    Checked 2026-09-08 by Codex (AI research agent).

  3. BNY Eliza digital employees

    Financial services · Operating use

    Approximately 140 digital employees

    BNY's term for multi-agent systems working alongside human colleagues.

    First-quarter 2026 results

    Unit: deployed multi-agent solutions · Organization total

    A deployment snapshot within the bank's AI program, reported separately from roughly 220 enterprise AI solutions in production.

    Limits: Each digital employee can contain multiple agents. This is neither a concurrent-worker count nor a measure of completed work per system.

    BNY · 2026-04-16
    Checked 2026-09-08 by Codex (AI research agent).

  4. BMW CodeRabbit code review

    Automotive software · Operating use

    More than 1,000 developers supported

    BMW's deployed source-code review system, supporting vehicle software development.

    August 2026; collaboration began more than two years earlier

    Unit: developers supported · Organization total

    Reviews proposed changes for developers deciding what is ready to ship.

    Limits: Support coverage is not active-use frequency, agent count, or an end-to-end software factory. BMW-specific review volume and defect rates are not disclosed.

    BMW Group · 2026-08-12
    Checked 2026-09-08 by Codex (AI research agent).

  5. Itaú Devin

    Financial services · Operating use

    75% of teams use Devin

    Vendor-hosted account featuring Itaú technology leaders.

    Undated customer account; measurement period not stated

    Unit: share of teams · Organization total

    Documents organizational adoption; the account describes code maintenance and repair workflows.

    Limits: The total number of teams, definition of use, and reporting window are missing. It does not mean 75% of work is automated or reveal concurrency.

    Cognition / Devin · Publication date not stated
    Checked 2026-09-08 by Codex (AI research agent).

Runtime and inference

Tokens measure model input or output; runtime measures elapsed agent activity. Track both with the task and outcome. Different models, tools, and accounting methods can produce very different totals for similar work.

Runtime and inference: dated first-party reports
Organization and taskReported scaleResult and limits
OpenAINavier–Stokes research agentsMathematics · Research experimentApproximately 130 billion output tokens

The Navier–Stokes effort, which also produced 2.7 million messages. No per-group breakdown is given.

September 2026 Navier–Stokes effort

Unit: output tokens · Reported run

Cumulative activity associated with the proposed proof.

Limits: The wider multi-problem campaign used about 300 billion output tokens. Neither total measures discoveries or simultaneous agents; cost is undisclosed.

OpenAI · 2026-09-08
Checked 2026-09-08 by Codex (AI research agent).

OpenAIInternal coding agentsAI research · Operating use3.1 agent-workdays per human workday

Aggregate agent runtime normalized to an eight-hour workday.

Research organization, mid-August 2026

Unit: runtime ratio · Organization total

Shows how much agent activity had entered daily research work.

Limits: Runtime is not equivalent productive labor. Coverage is incomplete. Separately, January–July task analysis found interventions in more than half of successful tasks estimated at 4–8 human hours.

OpenAI · 2026-09-06
Checked 2026-09-08 by Codex (AI research agent).

  1. OpenAI Navier–Stokes research agents

    Mathematics · Research experiment

    Approximately 130 billion output tokens

    The Navier–Stokes effort, which also produced 2.7 million messages. No per-group breakdown is given.

    September 2026 Navier–Stokes effort

    Unit: output tokens · Reported run

    Cumulative activity associated with the proposed proof.

    Limits: The wider multi-problem campaign used about 300 billion output tokens. Neither total measures discoveries or simultaneous agents; cost is undisclosed.

    OpenAI · 2026-09-08
    Checked 2026-09-08 by Codex (AI research agent).

  2. OpenAI Internal coding agents

    AI research · Operating use

    3.1 agent-workdays per human workday

    Aggregate agent runtime normalized to an eight-hour workday.

    Research organization, mid-August 2026

    Unit: runtime ratio · Organization total

    Shows how much agent activity had entered daily research work.

    Limits: Runtime is not equivalent productive labor. Coverage is incomplete. Separately, January–July task analysis found interventions in more than half of successful tasks estimated at 4–8 human hours.

    OpenAI · 2026-09-06
    Checked 2026-09-08 by Codex (AI research agent).

How to compare a new claim

Start with the unit and the time window. Then ask what entered the system, what counted as accepted work, and how much human effort was needed. Keep failed attempts, repair, and review in the calculation. “Not publicly stated” is a useful answer when the source omits a field.

For a claim about concurrency, look for average utilization as well as a peak. For a claim about adoption, look for the denominator: eligible teams, suitable tasks, or total accepted changes. For a speed claim, keep the quality and reproducibility requirements fixed. A run that finds more candidates may still spend most of its time revisiting weak ideas.

This collection is a selected public record. Companies disclose different things, and some disclose very little. Absence here does not establish absence of use. Old observations remain useful when their measurement window stays visible; a newly checked source does not turn an old figure into a current one.

The software factory metrics guide explains how to measure useful output, review effort, and total cost within your own team. The company cases explain the workflows behind the numbers.