A real company doubled engineering output with AI (2.09x). But per-reviewer load doubled and automated review overtook human review. AI lifts production faster than inspection, so the bottleneck moves to review. Scale the checking with the producing, or you just ship unreviewed work faster.
- A 2026 field study of a real company that mandated doubling engineering output found it worked: output reached 2.09x the baseline, driven by AI adoption. But the bottleneck moved.
- Across 802 developers and 196,212 pull requests, per-reviewer load roughly doubled and automated review overtook human review. The machines picked up the checking, because the humans could not keep pace.
- AI lifts the rate of production far more easily than the rate of inspection. Output is elastic; human judgement is not. Push output, and review becomes the binding constraint.
- Automated review is fine for mechanical checks and blind to judgement calls. On the surface nothing broke (merge and revert rates held steady), but revert rates only catch the failures you notice.
- Scale the checking with the producing: measure review capacity, decide what only a human may approve, and fund inspection with the savings AI creates on production.
You rolled out AI, and the output climbed. More shipped per person than ever, the dashboards looked wonderful, and for a while that was the whole story. Then you looked at the reviews. The same people who used to check the work were now drowning in twice as much of it, and somewhere along the way the checking had quietly started to be done by machines, because the humans simply could not keep up. You got the productivity you asked for. You are a good deal less sure you still have the quality.
That is not a hypothetical. A 2026 longitudinal study followed a real, AI-forward company that set an explicit "2x mandate", to double the merged pull requests per engineer. Over sixteen months it tracked 802 developers and 196,212 pull requests, and the mandate worked: output reached 2.09 times the pre-mandate baseline, and the gain tracked AI adoption and grew with continued use. It is one of the largest productivity gains from an AI-tools deployment documented in the field. This is a single-company case study rather than a universal law, and its authors are careful about that, but the pattern it exposes is the one every leader chasing AI-driven output should sit with.
What actually happened when output doubled?
The revealing part is not the doubling. It is what the doubling did to review. As output climbed, the load on each reviewer roughly doubled too, and automated review overtook human review. The people did not somehow check twice as much work; the machines absorbed the overflow. On the surface nothing broke, merge and revert rates held steady, which is genuinely reassuring for routine work. But it means the human act of looking at the work, judging it, catching the thing that is subtly wrong, was quietly displaced by automation as a side effect of pushing the numbers.
The general principle underneath is simple and easy to miss. AI raises the rate at which you can produce things far more easily than the rate at which you can inspect them. Generating is now cheap; the cost of using a model has fallen roughly 280-fold in eighteen months, by Stanford's measure. Checking is not cheap in the same way, because a human verifying whether an output is right still takes human time. Double the output and you double the review burden without doubling the reviewers. Something gives, and what gives is the depth, or the humanity, of the check.
Scale your AI output without quietly losing the check
The AI Strategy Session helps you lift productivity with AI while keeping review capacity and quality intact, not one at the expense of the other, in ninety minutes.
Book your Strategy SessionIs automated review actually a problem?
Not in itself, but it should be a decision you make on purpose, not one you inherit through overload. Automated checks are strong on the mechanical: does it compile, does it match the rule, does it pass the test. They are blind to the judgement calls: is this the right thing to have built, is it subtly but importantly wrong, does it do something no rule anticipated. In the study, revert rates stayed flat, which tells you the obvious failures were still being caught. It tells you nothing about the class of problems automated review cannot see, now being shipped at twice the rate. That is the quiet risk: not a visible break, but an invisible one, produced faster.
How do I scale output without losing the check?
By treating review capacity as something you manage deliberately, on the same footing as output.
- Measure review capacity, not just throughput. Track how much is being reviewed, and how deeply, alongside how much is being produced. If output doubles and reviewers do not, you are carrying a hidden deficit.
- Decide what only a human may approve. Draw the line explicitly: which changes and decisions require human judgement regardless of how confident the automation is. Make it a rule, not a reflex.
- Fund the checking side of the equation. If AI makes producing cheap, spend some of that saving on inspection, more reviewers, better tooling, and the time to actually look, rather than banking all of it as speed.
- Split the work by type. Let automation handle the mechanical, rule-based checks and reserve people for "is this right, is this wise." Designing where each belongs is part of designing the loop the work runs in.
- Watch what automation cannot see. Revert rates catch the obvious. Add checks for the subtle failures, the wrong thing built well, the quiet error that passes every automated gate.
AI makes producing twice as fast. It does not make checking twice as fast. Push output and the bottleneck moves to review. Protect that, or you are just shipping unreviewed work faster.
What does this change for me as a leader?
It changes which number you trust. The productivity figure, the 2x, is the seductive part and the least complete. Doubling output is real and worth having, but the question a good leader asks straight after is: at twice the volume, who is still able to check this, and how would I know if the honest answer had become "no one." Output you cannot inspect is not throughput you can trust; it is just risk, accumulating faster.
This is the discipline behind the end of business as usual: AI changes the rate of the work, and your job is to make sure the whole system, including the checking, keeps up, not just the part that shows on the dashboard. Build review capacity to match production capacity, and the speed-up becomes a durable advantage rather than a quiet pile of unexamined work. These are the same assurance questions a board should be asking when output suddenly jumps.
| Source | Finding on AI output and the review bottleneck |
|---|---|
| Enterprise "2x mandate" field study (2026) | Across 802 developers and 196,212 pull requests over 16 months, output reached 2.09x the pre-mandate baseline, driven by AI adoption |
| Same study, on oversight | Per-reviewer load roughly doubled and automated review overtook human review, while merge and revert rates held steady |
| Stanford HAI, 2025 AI Index | The cost of using a model fell around 280-fold in 18 months, producing output is now cheap, while human verification still costs human time |
Frequently asked questions
Does AI really double productivity?
Why does AI move the bottleneck to review?
How can leaders scale AI output without losing quality control?

About the author
British technology futurist, AI keynote speaker and advisor. Thirty years across enterprise technology and AI strategy, helping leaders navigate the future of work. The futurist who died.