Compare how 21 technology companies move work from request to checked result. Every case explains the operating system in plain language, names the AI-engineering patterns, and keeps company claims separate from what remains unknown.
21 company cases · sources checked through September 1, 2026 by a second AI editorial-review agent (Codex) · company claims stay labeled · alphabetical, not ranked
Compare company software factories
The same six fields appear for every company. Reported use tells you whether the public evidence describes an operating system, one bounded production line, an early rollout, or a product the company also uses itself. Access tells you what a reader can actually inspect or buy.
First-party report marks evidence supplied by the company. Observed artifact marks something readers can inspect directly. Neither label is a quality score, and no blank cell means hidden certainty: unknowns are written out.
The table can scroll sideways when the window is narrow. Tab to the comparison, then use Left and Right Arrow. Home and End move to the first and last columns.
Company software factory comparison, source window through September 1, 2026
Upgrade production Java applications at fleet scale.
Bounded production lineTens of thousands of migrations; the product path is changing.
Developers review and accept each transformation before it lands.
Existing customersNew Q Developer accounts and subscriptions stopped May 15, 2026. Eligible existing customers retain transformation access through April 30, 2027; AWS directs new modernization work to AWS Transform.
Airbnb gave every test file a visible state, used ordinary tools to check model-generated conversions, improved common failures in batches, and sent the final 3 percent to people.
Main job
Move 3,500 React test files through staged refactoring and validation.
Compare details
Reported use
Bounded production line One completed production migration program.
Human checkpoint
People tuned the line and finished the final 3% of files.
Access
Not publicly available The migration system is described publicly; its code is internal.
Amazon turned Java upgrades into repeated analyze, plan, transform, build, test, and review jobs, then began moving customers to successor products.
Main job
Upgrade production Java applications at fleet scale.
Compare details
Reported use
Bounded production line Tens of thousands of migrations; the product path is changing.
Human checkpoint
Developers review and accept each transformation before it lands.
Access
Existing customers New Q Developer accounts and subscriptions stopped May 15, 2026. Eligible existing customers retain transformation access through April 30, 2027; AWS directs new modernization work to AWS Transform.
Anthropic connects Claude Code and Claude Tag to remote environments, specialist review agents, deterministic checks, risk-based approval, and incident response.
Main job
Carry work from repository instructions through layered review and incident response.
Compare details
Reported use
Product with internal use Public products plus extensive first-party internal use.
Human checkpoint
Risk determines review, approval, deployment, and incident permissions.
Access
Mixed Claude products are public; Anthropic's internal Tag and integrations are private.
Block connects an open agent engine, a shared workflow that turns Slack requests into reviewed code changes, layered tests and review, and an early multiplayer workspace.
Main job
Turn Slack requests and tickets into tested pull requests and shared agent work.
Compare details
Reported use
Internal operation Builderbot is rolled out internally; Buzz remains early-stage dogfood.
Human checkpoint
People steer work, review changes, and approve consequential actions.
Access
Mixed Goose, Builderbot, and Buzz have public repositories; internal services differ.
Cloudflare centralizes identity, model access, cost, company context, approved internal tools, write controls, and risk-based AI code review across thousands of repositories.
Main job
Give company agents one identity, context, tool, cost, and review control plane.
Compare details
Reported use
Internal operation Company-wide internal system across thousands of repositories.
Human checkpoint
People own approval while risk rules can block serious findings.
Access
Internal The operating stack is internal; some underlying platform parts are public products.
Cognition starts Devin from chats, tickets, alerts, schedules, and APIs, then connects implementation to review, CI repair, bug investigation, and reusable playbooks.
Main job
Move bugs, design audits, and routine fixes from event to reviewed pull request.
Compare details
Reported use
Product with internal use Cognition publishes a detailed account of Devin building Devin.
Human checkpoint
People define playbooks, review changes, and own releases.
Access
Commercial Devin is commercial; the internal code and operating data are private.
Cursor combines isolated cloud coding agents, an eight-pass pull-request reviewer, and event-driven security agents with evidence and human escalation paths.
Main job
Build parallel pull requests with proof artifacts, then run event-driven review.
Compare details
Reported use
Product with internal use Customer-facing agents and internal review and security lines.
Human checkpoint
People review artifacts and own merge, severity, and production decisions.
Access
Mixed Cursor is commercial; a security-automation reference implementation is public.
Dropbox runs several coding agents through one internal cloud platform, with exact code snapshots, declared checks, bounded retries, one branch, and publication outside the agent.
Main job
Run interactive coding and repeatable repair jobs on one shared agent platform.
Compare details
Reported use
Internal operation Internal platform used for CI repair, flaky tests, and migrations.
Human checkpoint
The surrounding workflow controls tests, branches, publication, and retry limits.
Access
Not publicly available Nova and Deflaker are internal systems described in a public engineering article.
GitHub assigns issues in its core repository to Copilot, receives first-pass pull requests, and keeps the decision to merge, revise, or close with human engineers.
Main job
Turn assigned issues into draft pull requests inside the normal repository workflow.
Compare details
Reported use
Product with internal use Public product used inside the core repository that builds github.com.
Human checkpoint
A different person approves the work; repository and CI rules still apply.
Access
Commercial Copilot coding agent is a commercial GitHub capability.
Linear carries issue and customer context into a managed coding sandbox, then returns code and can include checks and visual evidence in the same workflow for review.
Main job
Use issues and team chat as dispatch and review surfaces for coding agents.
Compare details
Reported use
Product with internal use A recent public product with concrete internal design and triage examples.
Human checkpoint
People review the diff, proof artifacts, and merge decision in Linear.
Access
Commercial Coding Sessions is a commercial Linear capability; the implementation is private.
LinkedIn gives existing agents approved internal tools and reusable playbooks, then uses verifier-driven loops for specialized production machine-learning work.
Main job
Turn company knowledge and internal tools into reusable, executable playbooks.
Compare details
Reported use
Internal operation Internal system reported in use by more than 1,000 engineers.
Human checkpoint
Domain experts author and maintain the playbooks; engineers own results.
Access
Not publicly available CAPT is internal and built on the open Model Context Protocol.
Meta combines regression detection, profiling data, code history, reusable expert skills, agent implementation, and human review in one performance-efficiency line.
Main job
Turn production performance signals into diagnosed, review-ready code changes.
Compare details
Reported use
Internal operation Specialized internal line for performance regressions and opportunities.
Human checkpoint
Engineers review, deploy, and remain responsible for the proposed fix.
Access
Not publicly available The agent platform and its skills are internal.
Replit gives employees manager agents that can start several agents in parallel, while the wider system checks results and routes exceptions back to people.
Main job
Launch parallel agent work across company systems, check results, and learn from feedback.
Compare details
Reported use
Product with internal use Customer product plus a first-party internal agent-of-agents system.
Human checkpoint
People set policy, review risk-rated changes, and judge product outcomes.
Access
Commercial Replit Agent is public and commercial; internal orchestration and data are private.
Shopify runs River in public Slack threads while rolling Aquifer out profile by profile as a shared foundation for sessions, sandboxes, controls, and agent profiles.
Main job
Run public-to-the-company agent sessions and specialist lines on one durable substrate.
Compare details
Reported use
Internal operation Large internal operation; the shared Aquifer substrate is still rolling out.
Human checkpoint
People join visible threads, redirect work, and review pull requests and findings.
Access
Mixed Roast is open source; River, Aquifer, and Dispatch are internal.
Spotify kept existing systems for choosing repositories, coordinating work, checking results, and merging safe changes, then used Honk for complex transformations.
Main job
Apply agent judgment to repository-wide maintenance and data migrations.
Compare details
Reported use
Internal operation Internal agent work built on Spotify's existing fleet-management systems.
Human checkpoint
People initiate work and own review, merge, and unresolved edge cases.
Access
Mixed Honk is internal; Fleetshift is available as a managed Backstage product.
Tempo routes product and error signals through human prioritization, then Custom Agents work in parallel, open pull requests, and return decisions to people.
Main job
Turn product and error signals into prioritized, parallel agent work and review.
Compare details
Reported use
Early rollout Concrete public examples; organization-wide operating scale is not stated.
Human checkpoint
People prioritize the queue and review the pull requests or media output.
Access
Commercial Tempo's agent products are commercial; internal operating details are limited.
Airbnb gave every test file a visible state, used ordinary tools to check model-generated conversions, improved common failures in batches, and sent the final 3 percent to people.
Anthropic connects Claude Code and Claude Tag to remote environments, specialist review agents, deterministic checks, risk-based approval, and incident response.
Block connects an open agent engine, a shared workflow that turns Slack requests into reviewed code changes, layered tests and review, and an early multiplayer workspace.
Cloudflare centralizes identity, model access, cost, company context, approved internal tools, write controls, and risk-based AI code review across thousands of repositories.
Cognition starts Devin from chats, tickets, alerts, schedules, and APIs, then connects implementation to review, CI repair, bug investigation, and reusable playbooks.
Cursor combines isolated cloud coding agents, an eight-pass pull-request reviewer, and event-driven security agents with evidence and human escalation paths.
Dropbox runs several coding agents through one internal cloud platform, with exact code snapshots, declared checks, bounded retries, one branch, and publication outside the agent.
GitHub assigns issues in its core repository to Copilot, receives first-pass pull requests, and keeps the decision to merge, revise, or close with human engineers.
Linear carries issue and customer context into a managed coding sandbox, then returns code and can include checks and visual evidence in the same workflow for review.
LinkedIn gives existing agents approved internal tools and reusable playbooks, then uses verifier-driven loops for specialized production machine-learning work.
Meta combines regression detection, profiling data, code history, reusable expert skills, agent implementation, and human review in one performance-efficiency line.
Replit gives employees manager agents that can start several agents in parallel, while the wider system checks results and routes exceptions back to people.
Shopify runs River in public Slack threads while rolling Aquifer out profile by profile as a shared foundation for sessions, sandboxes, controls, and agent profiles.
Spotify kept existing systems for choosing repositories, coordinating work, checking results, and merging safe changes, then used Honk for complex transformations.
Tempo routes product and error signals through human prioritization, then Custom Agents work in parallel, open pull requests, and return decisions to people.