← All articles Part of: AI & the Workforce

What Stays Expensive

8 min read
Hand-drawn ink and crayon editorial illustration on a warm cream background, in steel blue with two small mustard yellow marks. A single tall wooden ladder stands upright and roughly centred in the frame, seen straight on, its two rails drawn as rough weathered poles with visible grain and dark charcoal outlines. The left rail leans very slightly inward toward the top so the ladder narrows as it rises. Eight evenly spaced horizontal rungs run between the rails, each a thick blue-grey bar with pale highlights along its upper edge. The top end of each rail is capped with a small mustard yellow mark. The ladder stands on a low mound of loose blue scrubby ground drawn with short scratchy strokes, its two feet planted a little apart, and a soft pale blue shadow spreads across the ground to the right. The rest of the background is empty cream. No text appears in the image.

Key takeaways

  • In the peer-reviewed version of Generative AI at Work, published in the Quarterly Journal of Economics in 2025, access to an AI assistant raised issues resolved per hour by 15% on average across 5,172 customer-support agents at one Fortune 500 firm.
  • The lowest skill quintile gained 36%. The most skilled saw no significant productivity change and, among mixed results on the study's other measures, small but statistically significant decreases in resolution rates and customer satisfaction. The authors say the findings should not be generalized across occupations or AI systems.
  • The researchers describe the mechanism as consistent with the idea that these tools may function by exposing lower-skill workers to the best practices of higher-skill workers. That is their reading of why, hedged twice, and not a measured finding.
  • The paper notes that top performers contribute many of the examples used to train the system, which raises their value to the firm, while their own measured performance is the least improved by it.
  • A throughput dashboard rolled out alongside an assistant will show your most experienced people as the ones who gained least. That is what the study would predict and not evidence that they resisted.

Somewhere in the last year or so, a piece of work landed on your desk and you couldn’t quite tell how much of it was the person.

It wasn’t bad. It was clean, organised roughly the way you’d have organised it, using the vocabulary of somebody three levels further along. Nothing in it was wrong. And it came from someone who, a while back, would have needed two more drafts and a conversation with you to get there.

That is not a complaint about the work, and it isn’t a story about anyone cutting corners. It’s a pricing problem. For most of your career you’ve read output as a proxy for the person who produced it, and the proxy was reliable, because producing the memo required knowing the things you needed them to know. The tool broke that link, and it broke it from the bottom up.

So here’s the rule underneath everything that follows. When a tool lifts the floor faster than it lifts the ceiling, every signal you were using to spot talent gets noisier, and the ones left standing are the ones the tool can’t produce.

Which ones those are is an empirical question, and there’s a piece of research that answers part of it unusually well.

The number that matters is the one that went down

Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered deployment of a generative AI assistant to 5,172 customer-support agents at a Fortune 500 firm that sells business-process software. The work circulated for two years as a working paper and was then published, peer reviewed, in the Quarterly Journal of Economics in 2025.

Access to the tool raised productivity, measured as issues resolved per hour, by 15% on average. Split that by skill and the average turns out to be hiding the finding. Workers in the lowest skill quintile gained 36%. For the most skilled workers, the researchers report that AI assistance “does not lead to any significant change in productivity.”

Then there’s the part that gets left out of the retelling. For the highest-skilled agents the paper reports mixed results across its other measures, including “small but statistically significant decreases” in resolution rates and customer satisfaction. In the abstract’s own summary, the least experienced improved both the speed and quality of their work, while “the most experienced and highest-skilled workers see small gains in speed and small declines in quality.”

One objection they closed off before anyone could raise it. The obvious reading of a top-end dip is mean reversion, that agents who happened to be measured high just before the rollout would drift down again anyway. They tested for it, and report a consistent linear increase in productivity across all skill levels afterwards, with no strong evidence of mean reversion.

So the honest version of this result is not that the tool helped everyone and helped beginners most. It’s that the tool moved people toward a common standard from both directions.

The researchers offer a reading of why, and they hedge it carefully, twice. Their results are “consistent with the idea that generative AI tools may function by exposing lower-skill workers to the best practices of higher-skill workers.” Elsewhere they describe suggestive evidence that adoption “drives convergence in communication patterns: low-skill agents begin communicating more like high-skill agents.” Worth holding those qualifiers, because everything downstream leans on them. The percentages are measurements. The explanation of what the tool was doing is the authors’ inference about the cause, and they were careful to say so.

One more thing about the timing, which cuts against the instinct to dismiss an old study. The rollout ran primarily through the autumn of 2020 and the winter of 2021, on a system built on GPT-3. That’s several generations behind anything you’d deploy today. This pattern was already visible with a much weaker tool, which makes the result a floor rather than a ceiling.

What actually got cheaper

Here is where this stops being a report of a study and starts being an argument, and it’s worth marking the join. Everything above happened in chat-based technical support for a stable software product, which is close to the easiest case a pattern-matching tool will ever get. Whether the same split runs through analysis, or advisory work, or your own management layer is not something this paper measured. What follows is a reading.

It helps to stop saying expertise as though it were one thing.

Part of it is accumulated knowledge. What good looks like, the standard phrasing, the framework that fits this category of problem, the twenty examples you’ve seen before. That part is transferable by description, which is exactly what a language model is good at supplying.

The other part is harder to name and it’s the part that decides outcomes. Knowing which of the twenty examples this one actually resembles. Knowing when the standard answer doesn’t apply here. Carrying context that was never written down anywhere. Recognising that a convincing answer is still the wrong answer, and saying so while everyone else in the room is nodding.

The first part just got much cheaper to supply. Nothing in this study suggests the second did, and the flat top end is at least consistent with that, though it does not prove it: the study measured resolutions per hour, not judgment, so reading the experts’ flatness as evidence about judgment assumes their edge was judgment rather than familiarity with one product’s quirks. Worth being honest that the assumption is mine.

That asymmetry is the repricing. And most of what an organisation does about talent, in hiring, in review, in promotion, was built when the two halves were bundled together and could safely be treated as one thing.

Six things that change

Stop reading output as a proxy for the person

The most common review conversation asks somebody to walk you through what they produced. That question now points at the part of the work with the most help behind it.

Ask for the decisions instead. What did you decide not to do. What did you escalate rather than resolve. Where did the first answer look right and turn out not to be, and what tipped you off. Those questions cost nothing to start asking and they reach the half of the work that hasn’t been repriced. They have an honest failure mode worth naming: somebody who genuinely made no interesting calls will have nothing to say, and that’s information rather than a broken question.

Change what your interview samples

If a work sample can be produced by a tool in twenty minutes, it has stopped separating candidates, and a cleaner submission may only mean better prompting.

The fix isn’t to ban the tool from the process, which mostly selects for people willing to say they didn’t use it. Give them the tool, and give them a case where the obvious answer is wrong for a reason only context supplies.

Here’s a shape that works, and it costs an afternoon to build. Take a real request that came into your business, one where the by-the-book answer turned out to be wrong. Strip out the single thing that made it wrong, the contract clause, the customer’s history, the regulatory quirk, and hand over what’s left along with the tool. A candidate who returns a polished, confident, wrong answer has told you one thing. A candidate who returns the same answer and adds that they’d want to check three specific things before it went out has told you the thing you’re actually hiring for. You’re scoring the question they ask, not the document they hand back, and the tool can’t ask it for them because it doesn’t know what’s missing either.

This lands on a process most companies stopped examining a while ago. The hiring freezes that funded AI budgets went through in large numbers while the underlying work design mostly stayed put, which means a lot of organisations changed who they hire before changing how they choose.

Write the second axis into the ladder

Most competency frameworks have one direction of travel. You climb by producing more, faster, with less supervision. That ladder made sense when throughput and judgment grew together, and it quietly stops making sense when a tool supplies throughput to people who haven’t built the other half yet.

Naming the second axis explicitly is unglamorous work, and it prevents an expensive mistake: promoting on the metric that moved, then wondering why the calls got worse. If you want somewhere crude to start, it isn’t seniority and it isn’t quality of output, both of which the tool now flatters. It’s closer to how often this person is right about what doesn’t apply here, and how early they say it. That’s a rough instrument and it beats leaving the axis unnamed, because an axis nobody has named is an axis nobody is assessed on. It matters more in a flatter organisation, not less, because what goes out with a manager layer is judgment redundancy along with the coordination overhead. Fewer people are checking, so more depends on whether the ones left were chosen for checking.

Look at what your dashboard is about to say about your best people

This one follows directly from the study’s structure, and it gets missed because it looks like an adoption problem.

Roll out an assistant, measure throughput before and after, and your most experienced people will show the smallest gain. In the study the pattern by tenure was cleanly monotonic, with the newest agents gaining most and the researchers reporting no effect on resolutions per hour for agents past a year. Look at quality alongside it and some of the strongest may tick slightly down, which is what the researchers found on resolution rates and satisfaction for the highest-skilled group.

Their explanation is worth reading slowly, because it isn’t the flattering one. They suggest AI recommendations “may distract top performers or lead them to choose the faster or less cognitively taxing option (following suggestions) rather than taking the time to come up with their own responses.” Not resistance, then, and not coasting either. Something closer to the opposite: your best people taking the assistant’s answer when they’d previously have written a better one.

Read the dashboard the wrong way and you’ll spend a quarter running adoption interventions at exactly the people who need the reverse conversation. Worth knowing what your own AI reporting will show before it shows it to somebody who doesn’t have this context.

Know what your top performers are contributing that your metrics can’t see

The paper has a finding almost nobody quotes, and it’s the one with the sharpest edge for anybody running a team.

The system was trained on the firm’s own successful conversations, so the best agents were the ones supplying the examples it learned from. The researchers put it plainly: top performers “contribute many of the examples used to train the AI system we study. This increases their value to the firm.” And then, in the same breath, they note that access to AI suggestions “may lead them to put less effort into coming up with new solutions.”

Read those two sentences together and you have a compensation problem with no established answer. The people whose judgment is being distilled into the tool are creating value that shows up in everybody else’s numbers and in none of their own. They may also be the people the tool quietly makes lazier, which is the loop the researchers flag when they note that “addressing this outcome is potentially important because the conversations of top agents are used for ongoing AI training.” The authors go as far as saying that pay policies rewarding contribution to model training “could be important.” That is a two-word hedge and I read more into it than it strictly says, so treat the emphasis as mine and the observation as theirs.

There is no settled answer to this and I am not going to invent one. But knowing that your throughput metrics systematically under-report your best people is enough to stop you making a decision on those metrics alone.

Fund the review step, because it’s where the difference now surfaces

If work arrives looking more capable than the person who produced it, the place the difference gets found is review. Not the drafting, not the tooling, the moment somebody senior reads it and decides whether it’s right.

That step is the one almost nobody funds as work in its own right. It sits in the gaps of other people’s calendars, it’s invisible in every plan, and it’s the first thing to compress when a team gets busier, which is exactly what happens when the drafting gets faster. An organisation leaning harder on judgment while never giving review a name, an owner and protected time is relying on its most repriced asset being donated in the margins.

The scope, stated plainly

One firm. One occupation. One tool, and an old one. A measure built on issues resolved per hour. The authors are unusually direct about this, writing that their findings “apply for a particular AI tool, used in a single firm, in a single occupation, and should not be generalized across all occupations and AI systems.” They add a specific caution that their setting had a stable product and a stable set of support questions, and that where the environment moves faster the effect could run differently, including AI “promoting outdated practices observed in historical training data.”

None of that is a reason to ignore the study. All of it is a reason not to quote a percentage at anybody as a forecast for your team. What this research supplies is a clean look at a shape that’s genuinely hard to see from inside one organisation: a capability arriving unevenly, landing hardest where the gap was knowledge, and leaving a different gap untouched.

The part that doesn’t have a move attached

There’s a version of this article that ends with a tidy checklist and a suggestion to review your comp bands. That version would be overclaiming.

Here is what can be said. The knowledge half of expertise is repricing quickly and visibly. The other half isn’t, and nobody yet has a good instrument for measuring it, which is why every organisation keeps using proxies that are getting noisier while the decisions stay just as consequential. Convergence toward a common standard is a real gain across most of the distribution, and at the top of it the same convergence showed up as a small drop in quality and, on the researchers’ reading, less effort spent reaching for a better answer. The study is honest that it saw both.

Which makes the useful question less about what to measure and more about what you’re willing to notice. Somewhere in your organisation this week, somebody is going to look at an answer that reads perfectly and say it doesn’t fit here. That’s the whole thing. It won’t show up on any dashboard, it never has, and it’s a larger share of what you’re paying for than it was five years ago.

Questions this article gets

Does the 36% mean my juniors are about to get 36% better?

No, and the authors rule that reading out themselves. They write that their findings apply for a particular AI tool, used in a single firm, in a single occupation, and should not be generalized across all occupations and AI systems. The 36% is the productivity lift for the lowest skill quintile at one Fortune 500 business-process software company, measured as issues resolved per hour. What travels is the shape rather than the number: the gain landed where the constraint was knowledge, and it did not land where the constraint was something else.

Is the answer to stop hiring juniors, or to stop hiring seniors?

Neither, and both readings come from treating the finding as a headcount question rather than a pricing one. If the knowledge half of expertise gets cheaper to supply, a junior becomes useful faster, which is an argument for hiring them. The judgment half gets scarcer relative to demand, which is an argument for paying for it. The mistake to avoid is quietly continuing to price both halves the way you did before, which is what happens by default because nobody has to decide anything to keep doing it.

The study ran on a GPT-3 era tool five years ago. Is it still relevant?

It is a real limitation and worth stating plainly: the rollout ran on a GPT-3 era tool through late 2020 and early 2021, and nobody has run this study again on a current model. What the age buys you is that the effect showed up before any organisation had built habits around these tools, so it caught the pattern clean rather than tangled up with how people have since learned to use them. Whether a stronger model would close the gap at the top rather than widen it is an open question, and this study cannot answer it in either direction.

Our best people barely use AI. Isn't that the real problem?

Check what your metrics are about to say before you decide it. In the study the most skilled workers saw no significant productivity change and, among mixed results on the other measures, small but statistically significant decreases in resolution rates and customer satisfaction. The researchers' own explanation is not that those workers resisted the tool but close to the reverse, that AI recommendations may lead top performers to take the less cognitively taxing option. If that is what is happening in your org, an adoption push makes it worse rather than better, because the problem is not that they are ignoring the assistant.

What is the cheapest place to start?

Look at what your review conversations actually ask for. Most of them ask people to walk through what they produced, which is now the part with the most help behind it. Asking instead what they decided not to do, what they escalated, and where they overrode the first answer they got costs nothing, needs no new process, and can start at the next one in your calendar.

Read the original post on LinkedIn