Banks, trading firms and manufacturers are publishing detailed accounts of agents that write, test, review and improve software. Scientific teams are applying similar methods to experiments and proofs. The useful comparison starts with the work: a code migration, a reviewed change, a planning algorithm or a research finding.
This reference follows those uses across industries. The company library examines individual operating systems in depth. The scale reference records numerical observations, including the largest concurrent runs in this collection. Together they show where industrialized agent work is appearing and what its public evidence can establish.
Industry overview
The examples below were checked on September 8, 2026. They are selected public accounts, not a representative survey. Their presence establishes documented activity; it cannot tell us what percentage of an industry has adopted software factories.
| Industry | Organizations and work | What the evidence establishes |
|---|---|---|
| Trading | Jane Street: internal coding tools and agent feedback | An internal harness used to build tools; no public swarm-size or firm-wide adoption figure |
| Banking | Capital One, Itaú, Nubank: security repair, development and migrations | Named operating workflows; metrics cover different populations and tasks |
| Financial operations | BNY: Eliza and digital employees | A reported portfolio of deployed AI solutions, including multi-agent systems |
| Automotive | BMW, Mercedes-Benz: code review and modernization | Deployed review at BMW; a vendor-reported modernization pilot and rollout at Mercedes-Benz |
| Chemicals and supply chains | BASF: evolved planning algorithms | Initial historical simulations with a measured objective; live network-wide autonomy is not established |
| Payments and machine learning | Klarna: training-code optimization | Thousands of candidate programs evaluated against speed and reproducibility constraints |
| Enterprise platforms | Palantir: AI FDE | A documented application-building agent; Palantir's internal adoption and swarm size remain undisclosed here |
| Industrial engineering | Siemens: simulation-code maintenance | A research prototype with specialist agent roles |
| Life sciences | Daiichi Sankyo and Bayer Crop Science: Co-Scientist | Named private-preview research use, reported by the provider |
| Scientific research | OpenAI and other research systems: proofs, analysis and hypothesis search | Documented research runs; scientific acceptance and business adoption require their own evidence |
Banking and trading: agents inside established engineering systems
Jane Street is a useful case because its development environment is unusually specific. Its AIDE case explains how an internal harness gives agents tools and feedback that fit the firm's software. For a product leader, the transferable question is how well an agent can work inside the systems the organization already depends on.
Capital One illustrates several task shapes within one company. Its case covers security investigation and repair, structured development plans, and a completed data-analysis workflow. Keeping those scopes separate makes the account more useful: repository coverage in a security system says little about the speed of an unrelated data task.
Itaú's case adds an adoption perspective. Its reported use spans development, testing, migrations and security repair. The denominator matters throughout: a share of teams is an adoption measure, while the share of scanner findings repaired describes one workflow's output.
Nubank supplies a particularly clear migration pattern. Its customer account describes independent subtasks, examples for training and evaluation, and engineers who review and merge the resulting changes. It reports an eight-to-twelvefold improvement in engineering-time efficiency on the delegated scope. The roughly 100,000 data-class implementations describe the overall migration project, not a verified count completed by agents. This is a historical project account, with no publication date displayed.
For comparison, ask which part of the work can be repeated with little variation, how failed attempts are handled, and whether the reported savings include human review. Those questions travel well between a trading firm and a retail bank.
BNY: counting deployed solutions
BNY's 2025 annual report defines its “digital employees” as multi-agent solutions. Its Q1 2026 presentation, page 4, reports approximately 220 enterprise AI solutions in production and 140 digital employees. It names payments processing, anomaly detection and onboarding among the work areas.
This is evidence of agents entering business operations. The 140 figure counts systems, each of which may contain several agents and run many times. It does not disclose how many workers are active together, or how many collaborate on any one task. A deployment portfolio and a research swarm need different columns in a comparison.
Automotive: review and modernization
BMW describes CodeRabbit supporting more than 1,000 software developers after over two years of collaboration. Its August 2026 account places agent review in vehicle software development, using repository context, requirements and test results to assess proposed changes. That establishes a deployed review layer. It leaves the size of any wider coding-agent operation open.
At Mercedes-Benz, Cognition's April 2026 account describes a four-week pilot involving more than 200,000 lines of COBOL. It reports eight days of modernization work against an eight-month estimate, followed by a wider product rollout. The estimate is a useful planning comparison, with no controlled causal result or public concurrency count attached.
These cases help a reader locate where agents enter an existing development process. BMW's evidence concerns assessing changes; the Mercedes-Benz example concerns transforming an older system. Neither requires every stage of development to become autonomous before it is worth studying.
BASF and Klarna: searching for better algorithms
BASF and Google describe a system that repeatedly changes a planning program and tests it against historical supply-chain data. Their May 2026 account reports thousands of experiments and an improvement over the initial model. The work was described as initial simulations. BASF's large production network explains the problem's importance; its size should not be read as the deployment coverage of the experiment.
Klarna describes a related loop for machine-learning training code. Over three weeks, AlphaEvolve evaluated nearly 6,000 candidate programs. The engineering account reports throughput rising from 49 to roughly 97 samples per second under reproducibility constraints. Faster candidates that broke those constraints were rejected.
The shared pattern is a searchable space of programs and a test that can discriminate between them. A team can spend substantial compute exploring alternatives when each improvement will be reused. To judge the economics, readers still need the search cost, the cost of checking a candidate, and the number of future uses. Candidate count alone cannot supply that answer.
Palantir: building inside the business platform
Palantir's AI FDE documentation describes an agent that builds pipelines, edits code and data models, and creates applications inside Foundry. It can run previews, inspect build checks and use the results to decide its next action. Changes are proposed through branches or pull requests by default, within the user's permissions.
This makes Palantir relevant to software factories across industries: the work can include the business's data structures and workflows alongside application code. Public product documentation establishes those capabilities. It does not establish how widely Palantir itself uses the system internally or the size of its agent teams. A detailed customer operating account would answer a different question and deserves its own evidence.
Industrial and scientific research
Siemens' Simcenter prototype divides code maintenance among agents that select documentation, update a simulation macro and explain the changes. The March 2025 post explicitly calls this research exploration. Three named roles do not establish three simultaneous workers or a deployed product.
Google's May 2026 science announcement names Daiichi Sankyo, Bayer Crop Science and U.S. National Labs as Co-Scientist users in private preview. That is a meaningful sign of scientific uptake, with outcomes and scale still requiring partner-specific evidence.
Research also changes what counts as a finished result. A program may pass its tests; a proposed proof needs mathematical scrutiny; a biological hypothesis needs appropriate experimental validation. The scale reference preserves the reported task and outcome beside each number so these differences survive comparison.
How to follow industry adoption
Track the depth of use within each organization: access to tools, a recurring workflow, multiple teams using it, and a documented share of work passing through it. Keep research experiments and commercial product capabilities labeled alongside operating accounts. These descriptions can coexist without implying a maturity ranking.
To estimate penetration across an industry, a study would need a defined population, a sampling method and a consistent definition of qualifying use. SWFT's public-source collection supplies examples and mechanisms. It also identifies the missing evidence that a broader study would need: adoption denominators, sustained output, rejected work, human effort and total cost.
The next useful step is to compare the company workflows, inspect reported scale, or use the measurement guide to evaluate one inside your own organization.