Hand an AI assistant to 5,172 customer support agents and issues resolved per hour go up 15%. Split that average by skill and it falls apart. The least-skilled fifth gained 36%. The most experienced gained a little speed and lost a little quality. The whole of that 15% came from closing a gap at the bottom, while the top barely moved. It comes from a staggered rollout at one firm, peer-reviewed in the Quarterly Journal of Economics in February 2025, and it’s been quoted for the 15% ever since.
Why It Matters
There’s a second line in the abstract that rarely survives the retelling. The gains were largest for what the authors call moderately rare problems, where the human had less experience but the system still had enough training data. So the same shape shows up twice, once in the people and once in the problems. The instinct is to point AI at the highest-volume work, since that’s where the savings look biggest on a slide, and this result sits somewhere else entirely. Your highest-volume work is the work your people have done a thousand times, so they’re already fast, and that’s where the study found the tool adding least. The genuinely unusual problems have too little behind them for the system to have learned much. The gain sits in between. Wherever human experience runs thin, the tool pays. Where it’s deep, it doesn’t.
The Decision
So the question worth asking is what you pointed your AI at, and why. Volume is a defensible answer. It’s measurable, it’s where the headcount sits, and it builds a clean business case. But volume tracks what your people do most, which is rarely the same list as what they know least. One firm, one job, and the authors are blunt that this shouldn’t be generalized. Treat it as a question, not a number.
What To Do This Week
- Take your largest AI deployment and ask its owner one question. Did we point this at our highest-volume work, or at the work our people see least often?
- Ask a frontline manager which questions their team looks up rather than answers. What they name is where experience runs thin, in their own words.
- Ask your support or ops lead for the last five things your AI tool got wrong. If they’re the unusual ones, that’s the boundary the researchers describe, and a better prompt won’t move it.
What Not To Do
Don’t buy on the average. A single 15% across everyone is the least useful form of this result, and it’s the form that reaches you in the vendor deck. Don’t point the tool at your most routine work just because that work is the easiest to measure. And don’t hold up your most experienced people as the proof case. In this study they gained least, and on a few quality measures they slipped.
Signal Boost
Brynjolfsson, Li and Raymond, Generative AI at Work - the full peer-reviewed paper from the Quarterly Journal of Economics, open access and opens straight to the document. Worth it for the authors’ own limits on the finding, stated more bluntly than most write-ups manage.