The bottleneck moved: what leadership has to change.
Your people are producing more than they did a year ago, yet your P&L does not show it. That gap has a mechanism, and it sits in a part of the work generative AI leaves untouched.
Producing work got cheaper. Judging work did not. Everything below follows from that asymmetry, and so does a good part of the distance between a pilot that worked and a result finance will sign off. Multiply what gets produced, leave the layer that approves it untouched, and you have built an operational bottleneck.
The gains are real. Start there.
Three randomised trials, two of them with money as the outcome. A Fortune 500 software firm gave 5,172 support agents a generative AI assistant: 15% more issues resolved per hour, 34% for novices, retention up and customer sentiment improved (Quarterly Journal of Economics, 2025). One of the world's largest cross-border retail platforms ran randomised experiments across seven workflows: a pre-sale chatbot lifted sales 16.3%, through more customers converting rather than bigger baskets (Fang, Yuan, Zhang, Donati and Sarvary, working paper 2026). And 640 Kenyan entrepreneurs were randomly given a GPT-4 business adviser: the strongest performers grew revenue and profit by just over 15% (Management Science, 2025).
Read the other half of each. In the support centre the most experienced agents gained close to nothing; the tool lifted the people who were still learning. On the platform, one of the seven workflows showed no detectable effect at all, and a second, advertising titles, a negative estimate too noisy to read. In Kenya there was no average effect: the weakest performers did 8% worse, and the difference lay in which advice each owner acted on. The P&L gain is real, and it is decided at the point of judgment.
Felt speed and measured speed disagree
METR ran a randomised trial in 2025 with 16 experienced open-source developers on 246 real tasks in repositories they knew well. With AI tools allowed, they took 19% longer. They had forecast a 24% speed-up. Afterwards they still believed they had been 20% faster. Wrong before, and wrong after.
The people living through the slowdown reported a speed-up. That is the kind of thing an adoption survey measures.
Where the time actually went
Google Cloud's 2025 DORA report, roughly 5,000 technology professionals worldwide, finds that higher AI adoption is associated with a rise in delivery throughput and a rise in delivery instability at the same time. Their explanation is the point: the time saved in generation goes straight back out on verification and prompting. 30% of developers report little to no trust in AI-generated code.
The human side of that check degrades where you need it most. Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real AI-assisted tasks (CHI 2025). Higher confidence in the AI goes with less critical thinking. Higher confidence in oneself goes with more. AI output is fluent, confident and plausible, which the automation-bias literature (Parasuraman and Manzey, 2010) identifies as precisely the profile that suppresses checking.
The first instinct is to add reviewers, or to point AI at the checking. Adding reviewers does not scale against multiplied volume, and in our experience AI checking AI shares the failure mode of the work it is checking. The way out is deciding which work needs checking at all, and how deeply.
A year of harder scaling, the same financial result
McKinsey's State of AI in 2026, 1,719 participants fielded in May and June 2026, reports that 80% of respondents say AI improved their individual productivity, while 37% say it contributed positively to their organisation's EBIT. The EBIT share is flat against a year earlier. Over the same year, enterprise-wide scaling went from 38% to 44%, and the high performers who attribute 5% or more of EBIT to AI stayed at about 6%.
That is the asymmetry showing up in the accounts. It is a survey, so it measures perception of impact rather than audited impact. It is also the widest sample anyone has.
The payroll data say the same thing without asking anyone how they feel. Humlum and Vestergaard linked adoption surveys of 25,000 workers at 7,000 Danish workplaces to payroll records. Two years after AI chatbots arrived in the most exposed occupations, earnings and hours had not moved, with confidence intervals ruling out anything above 2%, and that held for intensive users, early adopters and workplaces that had invested heavily. Users reported saving about 3% of their time. Individuals feel the gain. Payrolls have yet to record it.
A rival explanation deserves naming, because it fits the same gap. Earlier general purpose technologies showed the same lag, and the Productivity J-Curve (Brynjolfsson, Rock and Syverson, 2021) describes why: the complementary investments that make a technology pay (process redesign, co-invention, human capital) land before any measured productivity does. Brynjolfsson, Hitt and Yang put the scale at ten dollars of organisational capital for every dollar of technology. On that reading the gap closes with patience. The two explanations point at the same bill and disagree about the line item. The J-curve says buy the organisational capital. The evidence above says the part of it currently binding is the capacity to approve. Under either reading, more licences is the one purchase that changes nothing.
The person carrying it has a job title
Split the same survey by level. 47% of midlevel managers and individual contributors report at least one negative effect of AI on their work, against 31% of executives and senior managers. Among midlevel managers, 20% say AI hurts their ability to think critically against 9% of C-level respondents, 19% report anxiety about their career prospects, and 14% say they are overwhelmed by the volume of AI output. Midlevel managers absorb the verification load. No dashboard shows it.
Underneath it, the training ground is thinning. Stanford's Digital Economy Lab, working from ADP payroll data covering millions of workers, finds no evidence of widespread, economy-wide job displacement. Inside that, employment for 22 to 25 year olds in the most AI-exposed occupations now stands 19% below where it would be had it kept pace with their less-exposed peers. Entry-level work is where senior judgment was historically manufactured. Automate it away while judgment is the binding constraint, and you are spending the capacity you will need to approve with.
What a leadership team changes
Four shifts. Each has an observable test.
From buying the tool to rebuilding the work. Nearly three-quarters of McKinsey's high performers (n = 92) have fundamentally redesigned workflows because of AI, against a quarter of everyone else (n = 1,429). The test is the split of your AI budget between licences and redesign.
From measuring production to measuring approved output. Put decision latency on the dashboard next to volume. The test is whether anyone in the room can say how long a decision currently waits between ready and approved.
From trusting felt gains to demanding measured ones. Baseline before rollout, and let a null result stand. The test is whether a pilot has ever been stopped on the numbers.
From delegating verification downward to designing where it sits. Classify decisions by consequence, then state per class whether a human sits in the loop, above it, or out of it, with one owner by name. The test is whether that classification exists on paper. For high-risk systems this is already law, before it is our advice: EU AI Act Article 14 requires that a person can understand the output, intervene, override and stop the system. The Digital Omnibus, in force since July 2026, moved the high-risk deadlines to December 2027 and August 2028, so treat it as the direction of travel and the minimum bar.
If you take one, take the third. The others cannot hold without measurement, and the evidence says felt experience points the wrong way here. It is also the shift that self-corrects, because whoever starts measuring discovers the rest.
Of the eleven practices McKinsey found separating the 6% from everyone else, one sits under Strategy: they have determined how and when model outputs need a human in the loop. The firms making money from AI are the ones that have decided where verification sits; McKinsey reports the correlation, not the direction.
The studies locate the constraint; they do not test these shifts. That is what the tests are for: run them on your own numbers.
Four questions for your next leadership meeting
- What share of our AI budget goes to licences, and what share to redesigning the work?
- How long does a decision wait here between ready and approved?
- Has a pilot ever been stopped on the numbers?
- If we rebuilt this process AI-first, which decisions would still need a person, and who would own each of them by name?
Where we come in
We run the four shifts above with leadership teams, inside their own work, on the numbers it produces. First your leaders run it with us. Whether they run the next one without us is what we come back to check.
How we work with leadership teams →
← All field notes