# How companies build AI software factories

Compare how 21 technology companies move work from request to checked result. Every case explains the operating system in plain language, names the AI-engineering patterns, and keeps company claims separate from what remains unknown.

21 company cases · sources checked through September 1, 2026 by a second AI editorial-review agent (Codex) · company claims stay labeled · alphabetical, not ranked

## Compare company software factories

The same six fields appear for every company. Reported use tells you whether the public evidence describes an operating system, one bounded production line, an early rollout, or a product the company also uses itself. Access tells you what a reader can actually inspect or buy.

First-party report marks evidence supplied by the company. Observed artifact marks something readers can inspect directly. Neither label is a quality score, and no blank cell means hidden certainty: unknowns are written out.

| Company and system | Main job | Reported use | Human checkpoint | Access | Source window |
| --- | --- | --- | --- | --- | --- |
| [Airbnb](/companies/airbnb-llm-migration) · LLM migration pipeline | Move 3,500 React test files through staged refactoring and validation. | **Bounded production line.** One completed production migration program. | People tuned the line and finished the final 3% of files. | **Not publicly available.** The migration system is described publicly; its code is internal. | First-party report; 2026-09-01 |
| [Amazon](/companies/amazon-q-migrations) · Q Developer code transformation | Upgrade production Java applications at fleet scale. | **Bounded production line.** Tens of thousands of migrations; the product path is changing. | Developers review and accept each transformation before it lands. | **Existing customers.** New Q Developer accounts and subscriptions stopped May 15, 2026. Eligible existing customers retain transformation access through April 30, 2027; AWS directs new modernization work to AWS Transform. | First-party report; 2026-09-01 |
| [Anthropic](/companies/anthropic-ai-native-sdlc) · Claude Code and Claude Tag | Carry work from repository instructions through layered review and incident response. | **Product with internal use.** Public products plus extensive first-party internal use. | Risk determines review, approval, deployment, and incident permissions. | **Mixed.** Claude products are public; Anthropic's internal Tag and integrations are private. | First-party report; 2026-09-01 |
| [Block](/companies/block-builderbot-goose-buzz) · Builderbot, Goose, and Buzz | Turn Slack requests and tickets into tested pull requests and shared agent work. | **Internal operation.** Builderbot is rolled out internally; Buzz remains early-stage dogfood. | People steer work, review changes, and approve consequential actions. | **Mixed.** Goose, Builderbot, and Buzz have public repositories; internal services differ. | First-party report + Observed artifact; 2026-09-01 |
| [Cloudflare](/companies/cloudflare-ai-engineering-stack) · Internal AI engineering stack | Give company agents one identity, context, tool, cost, and review control plane. | **Internal operation.** Company-wide internal system across thousands of repositories. | People own approval while risk rules can block serious findings. | **Internal.** The operating stack is internal; some underlying platform parts are public products. | First-party report; 2026-09-01 |
| [Cognition](/companies/cognition-builds-devin) · Devin | Move bugs, design audits, and routine fixes from event to reviewed pull request. | **Product with internal use.** Cognition publishes a detailed account of Devin building Devin. | People define playbooks, review changes, and own releases. | **Commercial.** Devin is commercial; the internal code and operating data are private. | First-party report; 2026-09-01 |
| [Cursor](/companies/cursor-agent-factory) · Cloud agents and security agents | Build parallel pull requests with proof artifacts, then run event-driven review. | **Product with internal use.** Customer-facing agents and internal review and security lines. | People review artifacts and own merge, severity, and production decisions. | **Mixed.** Cursor is commercial; a security-automation reference implementation is public. | First-party report + Observed artifact; 2026-09-01 |
| [Dropbox](/companies/dropbox-nova) · Nova and Deflaker | Run interactive coding and repeatable repair jobs on one shared agent platform. | **Internal operation.** Internal platform used for CI repair, flaky tests, and migrations. | The surrounding workflow controls tests, branches, publication, and retry limits. | **Not publicly available.** Nova and Deflaker are internal systems described in a public engineering article. | First-party report; 2026-09-01 |
| [GitHub](/companies/github-copilot-agent) · Copilot coding agent | Turn assigned issues into draft pull requests inside the normal repository workflow. | **Product with internal use.** Public product used inside the core repository that builds github.com. | A different person approves the work; repository and CI rules still apply. | **Commercial.** Copilot coding agent is a commercial GitHub capability. | First-party report; 2026-09-01 |
| [Linear](/companies/linear-coding-sessions) · Coding Sessions | Use issues and team chat as dispatch and review surfaces for coding agents. | **Product with internal use.** A recent public product with concrete internal design and triage examples. | People review the diff, proof artifacts, and merge decision in Linear. | **Commercial.** Coding Sessions is a commercial Linear capability; the implementation is private. | First-party report; 2026-09-01 |
| [LinkedIn](/companies/linkedin-capt) · CAPT | Turn company knowledge and internal tools into reusable, executable playbooks. | **Internal operation.** Internal system reported in use by more than 1,000 engineers. | Domain experts author and maintain the playbooks; engineers own results. | **Not publicly available.** CAPT is internal and built on the open Model Context Protocol. | First-party report; 2026-09-01 |
| [Meta](/companies/meta-capacity-agents) · Capacity Efficiency agents | Turn production performance signals into diagnosed, review-ready code changes. | **Internal operation.** Specialized internal line for performance regressions and opportunities. | Engineers review, deploy, and remain responsible for the proposed fix. | **Not publicly available.** The agent platform and its skills are internal. | First-party report; 2026-09-01 |
| [OpenAI](/companies/openai-codex-symphony) · Codex and Symphony | Pull tasks from a project board into isolated agent workspaces and review loops. | **Internal operation.** Internal teams plus one fully agent-written experimental product repository. | People prioritize work, set acceptance criteria, and validate outcomes; pull-request review may be agent-to-agent. | **Mixed.** Codex CLI and Symphony are public; internal repositories and operating data are not. | First-party report + Observed artifact; 2026-09-01 |
| [Ramp](/companies/ramp-inspect) · Inspect | Run background coding with full company context in prepared remote environments. | **Internal operation.** Internal operating system with multiple specialized factory lines. | People review and land changes; automated lines stop at review-ready work. | **Not publicly available.** Inspect is internal; an independent inspired implementation is public. | First-party report + Observed artifact; 2026-09-01 |
| [Replit](/companies/replit-self-driving-company) · Internal agent fleets | Launch parallel agent work across company systems, check results, and learn from feedback. | **Product with internal use.** Customer product plus a first-party internal agent-of-agents system. | People set policy, review risk-rated changes, and judge product outcomes. | **Commercial.** Replit Agent is public and commercial; internal orchestration and data are private. | First-party report; 2026-09-01 |
| [Shopify](/companies/shopify-river-aquifer) · World, River, and Aquifer | Run public-to-the-company agent sessions and specialist lines on one durable substrate. | **Internal operation.** Large internal operation; the shared Aquifer substrate is still rolling out. | People join visible threads, redirect work, and review pull requests and findings. | **Mixed.** Roast is open source; River, Aquifer, and Dispatch are internal. | First-party report; 2026-09-01 |
| [Spotify](/companies/spotify-honk) · Honk and Fleetshift | Apply agent judgment to repository-wide maintenance and data migrations. | **Internal operation.** Internal agent work built on Spotify's existing fleet-management systems. | People initiate work and own review, merge, and unresolved edge cases. | **Mixed.** Honk is internal; Fleetshift is available as a managed Backstage product. | First-party report; 2026-09-01 |
| [Stripe](/companies/stripe-minions) · Minions | Turn one unattended delegation into a checked branch for two human review gates. | **Internal operation.** High-volume internal background-agent operation. | The initiator inspects the result and another engineer reviews the pull request. | **Not publicly available.** Minions and its Stripe-specific infrastructure are internal. | First-party report; 2026-09-01 |
| [Tempo](/companies/tempo-software-factory) · Custom Agents | Turn product and error signals into prioritized, parallel agent work and review. | **Early rollout.** Concrete public examples; organization-wide operating scale is not stated. | People prioritize the queue and review the pull requests or media output. | **Commercial.** Tempo's agent products are commercial; internal operating details are limited. | First-party report; 2026-09-01 |
| [Uber](/companies/uber-software-factory) · Software Factory | Coordinate coding agents, shared context, managed jobs, checks, and result economics. | **Internal operation.** Company operating model described across multiple factory layers. | People define valuable work and own quality, acceptance, and release. | **Not publicly available.** Uber's factory infrastructure is internal. | First-party report; 2026-09-01 |
| [Vercel](/companies/vercel-eve) · eve | Give many production agents durable runs, sandboxes, permissions, traces, and evals. | **Internal operation.** Shared framework reported across more than 100 internal production agents. | Engineers own release controls, executable guardrails, and production outcomes. | **Mixed.** The eve framework is open source; Vercel's internal agents and connections are private. | First-party report + Observed artifact; 2026-09-01 |

## Read the company cases

Each case follows the workflow, the proof, the human checkpoint, the public artifact, and the limit of the available evidence.

- [Airbnb's LLM migration factory: 3,500 test files with a proof loop](/companies/airbnb-llm-migration) — Airbnb gave every test file a visible state, used ordinary tools to check model-generated conversions, improved common failures in batches, and sent the final 3 percent to people.
- [Amazon Q migrations: an AI factory for fleet-wide Java upgrades](/companies/amazon-q-migrations) — Amazon turned Java upgrades into repeated analyze, plan, transform, build, test, and review jobs, then began moving customers to successor products.
- [Anthropic's AI-native SDLC: agents write, review, test, and investigate code](/companies/anthropic-ai-native-sdlc) — Anthropic connects Claude Code and Claude Tag to remote environments, specialist review agents, deterministic checks, risk-based approval, and incident response.
- [Block's AI software factory: Goose, Builderbot, and the Buzz workspace](/companies/block-builderbot-goose-buzz) — Block connects an open agent engine, a shared workflow that turns Slack requests into reviewed code changes, layered tests and review, and an early multiplayer workspace.
- [Cloudflare's AI engineering stack: a control plane for company-wide agents](/companies/cloudflare-ai-engineering-stack) — Cloudflare centralizes identity, model access, cost, company context, approved internal tools, write controls, and risk-based AI code review across thousands of repositories.
- [How Cognition uses Devin to build Devin](/companies/cognition-builds-devin) — Cognition starts Devin from chats, tickets, alerts, schedules, and APIs, then connects implementation to review, CI repair, bug investigation, and reusable playbooks.
- [Cursor's agent factory: cloud computers, Bugbot review, and security loops](/companies/cursor-agent-factory) — Cursor combines isolated cloud coding agents, an eight-pass pull-request reviewer, and event-driven security agents with evidence and human escalation paths.
- [Dropbox Nova: agents propose changes while deterministic systems control proof](/companies/dropbox-nova) — Dropbox runs several coding agents through one internal cloud platform, with exact code snapshots, declared checks, bounded retries, one branch, and publication outside the agent.
- [How GitHub uses Copilot coding agent to build github.com](/companies/github-copilot-agent) — GitHub assigns issues in its core repository to Copilot, receives first-pass pull requests, and keeps the decision to merge, revise, or close with human engineers.
- [Linear coding sessions: from product issue to reviewed pull request](/companies/linear-coding-sessions) — Linear carries issue and customer context into a managed coding sandbox, then returns code and can include checks and visual evidence in the same workflow for review.
- [LinkedIn CAPT: turning company knowledge into executable agent playbooks](/companies/linkedin-capt) — LinkedIn gives existing agents approved internal tools and reusable playbooks, then uses verifier-driven loops for specialized production machine-learning work.
- [Meta's capacity agents: production signals become review-ready fixes](/companies/meta-capacity-agents) — Meta combines regression detection, profiling data, code history, reusable expert skills, agent implementation, and human review in one performance-efficiency line.
- [OpenAI's Codex factory: harness engineering and Symphony orchestration](/companies/openai-codex-symphony) — OpenAI made one repository legible and testable by Codex, then used Symphony to turn project issues into isolated, continuously managed agent work.
- [Ramp Inspect: how a background coding agent became factory infrastructure](/companies/ramp-inspect) — Inspect pairs a prepared remote development environment with company context, executable feedback, human acceptance, and specialized production lines.
- [Replit's self-driving company: manager agents launch parallel work](/companies/replit-self-driving-company) — Replit gives employees manager agents that can start several agents in parallel, while the wider system checks results and routes exceptions back to people.
- [Shopify River and Aquifer: one durable platform for many software agents](/companies/shopify-river-aquifer) — Shopify runs River in public Slack threads while rolling Aquifer out profile by profile as a shared foundation for sessions, sandboxes, controls, and agent profiles.
- [Spotify Honk: adding an agent to a software factory that already worked](/companies/spotify-honk) — Spotify kept existing systems for choosing repositories, coordinating work, checking results, and merging safe changes, then used Honk for complex transformations.
- [Stripe Minions: how developer infrastructure became an AI software factory](/companies/stripe-minions) — Stripe combines ready-to-use development computers, scoped context, curated internal tools, repeatable workflow steps, limited automated build-and-test repair, and two human review gates.
- [Tempo's software factory: turning product signals into review-ready PRs](/companies/tempo-software-factory) — Tempo routes product and error signals through human prioritization, then Custom Agents work in parallel, open pull requests, and return decisions to people.
- [Inside Uber's AI software factory: measuring cost and quality at scale](/companies/uber-software-factory) — Uber connects coding-agent adoption with reusable skills, managed workloads, quality signals, and cost and quality per completed outcome.
- [Vercel eve: the shared infrastructure behind more than 100 production agents](/companies/vercel-eve) — Vercel standardized file-based agent definitions, resumable runs, sandboxes, permissioned connections, activity records, repeatable behavior checks, previews, and rollback in one platform.

Cases are included when a first-party source describes repeatable intake, agent work, verification, and handoff. General coding-assistant adoption is not enough. [Read the editorial method.](/methods)
