SecondSource logo

SecondSource

Archives
Log in
Subscribe
July 18, 2026

SecondSource Deep Dive · 2026/7/17|Deep Dive #9 — What AI Levels, the Market Reprices

The short version

Three years ago, the most-cited finding in all of AI-and-work research said this: give consultants GPT-4, and the bottom half of performers improve 43% on quality while the top half gain only 17% — the gap between them collapsing from 22 percentage points to 4. "AI is a skill leveler" has been the single most quoted line in corporate workforce decks ever since. This March, the study behind it cleared peer review and landed in the management journal Organization Science (INFORMS). But over the same stretch in which that result hardened, three opposing lines of evidence grew up alongside it: in real-world business decisions with no answer key, AI made the strong stronger and the weak weaker; in the labor market, the entry-level jobs most easily leveled are the ones in relative decline; and the most famous piece of counter-evidence ("top performers benefit most") was disavowed by MIT as resting on data that cannot be trusted. This essay puts three years of evidence on one table. The leveling is real, and more real than most people assume. But what it levels is output quality, not market value. The slice of work that got leveled is being repriced.

For anyone who has to make a decision this quarter, here is the whole essay in one move: stop agonizing over whether to train employees on AI — the leveling mostly happens on its own — and instead sort your role list into two piles by a single question: does this job have a checkable quality template? Pay for the template pile will drift toward the price of an AI subscription, so stop buying it with premium salaries. Move the training and raise budget toward developing people into verification and gatekeeping work: the jobs outside the template. The three thousand or so words that follow are the evidence for that paragraph. If all you needed was the first decision, you already have it.

The experiment that just got its degree

Foundation first. In September 2023, researchers from Harvard, Wharton, and MIT partnered with Boston Consulting Group to run an experiment on 758 consultants (roughly 7% of the firm's workforce), randomly assigning who could and couldn't use GPT-4 across 18 tasks designed to mirror real consulting work. Consultants with AI completed 12.2% more tasks and worked 25.1% faster, with quality about 40% higher. The famous part is the distribution: consultants who scored in the bottom half of a baseline assessment improved 43% on quality, while the top half improved only 17%, and the average gap between the two groups shrank from 22 percentage points to 4. This is the birthplace of the phrase "skill leveling."

Why did this one study get cited more than any other? Because it assembled nearly every ingredient of credibility at once: large sample, pre-registered, run inside a real company, on tasks designed by that company's own people. The one thing it lacked was peer review. That gap closed in March 2026, when the paper formally appeared in Organization Science (INFORMS). Numbers that have been pasted into slide decks for three years are now textbook-grade evidence.

The paper carries its own shadow, and always has. The researchers deliberately designed one task to sit outside AI's capability range — and on that task, consultants using AI were 19 percentage points less likely to get the right answer; in the journal abstract's own words, AI users were "less likely to produce correct solutions." The authors named the phenomenon the "jagged frontier": AI's strengths and weaknesses are not distributed the way human intuition expects, and the boundary is invisible. From day one, the leveling held only inside the frontier.

Wherever there's a template, it repeats

If leveling had shown up only at BCG, it would be an anecdote. It isn't. Four independent research teams, on four different kinds of work, measured the same shape — and all four papers have now cleared peer review:

Writing. MIT researchers had 453 college-educated professionals do writing tasks drawn from real occupational work: those using ChatGPT took about 40% less time and scored 18% higher on quality, and the initially weaker writers benefited most (Science).

The support floor, same story. Economists at Stanford and MIT obtained the complete work records of 5,172 customer-support agents at a publicly traded software company: an AI assistant lifted issues resolved per hour by 15% on average, with the gains concentrated among novices and lower-skilled agents: the most junior group improved around 30%, and agents two months into the job performed on par with colleagues six months in without AI (QJE). Roughly speaking, new hires climbed the learning curve three times faster.

Law bends the shape a little. Law professors at the University of Minnesota had students write legal analysis with GPT-4: students at the bottom of the class improved dramatically while students at the top slightly declined; the quality gains overall were slight and inconsistent, but the speed gains were large and consistent (Minnesota Law).

Creative writing. Given five AI-generated ideas, initially less creative writers gained about 10.7% in novelty and 11.5% in usefulness — enough that it "effectively equalizes the creativity scores across less and more creative writers." But the same study measured leveling's shadow: "generative AI–enabled stories are more similar to each other than stories by humans alone" — the group's collective diversity went down even as individual scores went up (Science Advances).

Line the four up and the common denominator surfaces: every one of these tasks has a template for what a good answer looks like. A competent consulting deck, a well-handled support reply, a properly structured legal memo — there is a known shape to aim at. Where a template exists, AI amounts to distributing that template to everyone at commodity prices, and the people who started furthest from the template naturally get pulled up the most. The creative-writing study shows you the cost: when everyone receives the same template, the output starts reading like one author. The same mechanism produces both: the leveling and the homogenization.

The line even extends to team structure. In 2025, Wharton and Harvard researchers — a group overlapping heavily with the core team behind the BCG experiment — ran a field experiment at Procter & Gamble: one person working with AI matched the performance of a two-person team working without it (Import AI). The leveling reached two gaps at once: between individuals, and between an individual and a team. This one is still a single venue and a relayed source, so it carries less weight than the four above; the direction is the same.

Where there's no answer key, the gap widens

Change the kind of task, and the result doesn't just weaken — it inverts.

A field experiment published in Management Science in 2025 randomized 640 Kenyan small-business owners; the treatment group got a GPT-4 business adviser available on demand over WhatsApp, with outcomes tracked over several months (Management Science; Berkeley Haas). The average effect: close to zero. Split the sample and the zero comes apart: owners who were already running their businesses well benefited by just over 15%, while those running them poorly did about 8% worse. That is the exact opposite of the BCG direction. The gap didn't compress, it widened, and the bottom didn't merely gain less: it got absolutely worse.

The mechanism is the most valuable part. The researchers found that high and low performers asked the AI questions of similar quality and received advice of similar quality; the difference showed up in selection and execution. Good operators knew which advice was worth acting on and could actually carry it out; struggling operators often walked away holding advice they couldn't implement, or spent their effort on the wrong problem entirely. AI made "advice" cheap and abundant. The judgment required to cash advice in was not leveled by a single cent.

Software engineering supplies the second counterexample. METR (Model Evaluation and Threat Research), an AI evaluation research group, had 16 veteran open-source developers work on repositories they had maintained for years, across 246 real issues, with AI-tool access randomized per task: with AI they were 19% slower. Yet before the study, the developers predicted AI would make them 24% faster — and afterward, they still estimated it had made them 20% faster (METR). Sixteen people in a single domain is limited weight. But "perception and measurement point in opposite directions" is, by itself, a warning shot at every self-reported AI productivity number in circulation.

Medicine sits on the same side of the ledger: since 2025, most studies have leaned toward the finding that LLM assistance does not meaningfully improve physicians' diagnostic reasoning — directionally supporting the view that these tools shouldn't diagnose autonomously without a physician as gatekeeper — and AI's errors still need someone able to catch them. We have not verified these medical studies against their primary sources one by one; we cite them as directional corroboration only and attach no numbers.

Put this section next to the previous one and a candidate dividing line emerges: whether the task has a checkable quality template. A consulting deck has one; a support reply has one. "What should my small business do next" has none — and neither does the correct fix for a hard ticket in a mature codebase. Kenya and BCG also differ on venue realism, tracking duration, and subject population — three confounds nobody has excluded — so singling out "the template" is a reading, not a verdict. It is simply the one line that currently explains both sides of the evidence at once, and the first of the observation windows at the end of this essay exists to test it. Where a template exists, AI pulls people toward the template; where none exists, AI hands you raw material that takes judgment to convert — and judgment is what Wharton professor Ethan Mollick was pointing at back in 2022, in his newsletter One Useful Thing, when he wrote that "creative AI will not make a non-expert an expert" (One Useful Thing). That 2022 sentence got its hardest empirical validation in Kenya, in 2025.

The task got leveled; the job is disappearing

Experiments hand people tasks; the labor market prices jobs. The third line of evidence comes from payroll files, and it measures a different layer than either section above.

Stanford's Digital Economy Lab tracked millions of American workers through ADP payroll records. The November 2025 version of the results: in the occupations most exposed to AI, relative employment of entry-level workers aged 22 to 25 declined 16%, while senior employment in the same occupations held stable — and the decline concentrated in one kind of exposure. In the paper's words: "entry-level employment has declined in applications of AI that automate work, but not those that most augment it" (Stanford Digital Economy Lab).

Resist reading that as an AI bloodbath. Over the same period, as relayed in a Dwarkesh Podcast interview, Yale's Budget Lab had scanned aggregate employment data and concluded that 33 months after ChatGPT's launch, "you really have to squint to see anything happening" in the way of AI-driven labor-market disruption (Dwarkesh Podcast); and Danish researchers who linked the official wage records of 25,000 workers found that AI chatbots had no significant effect on earnings or hours in any occupation they examined (Humlum & Vestergaard). These three datasets are not actually fighting: one measures the aggregate, one measures relative change by age within occupations, one measures incumbent wages. Aggregate calm and structural reshuffling can both be true at once.

But the structural signal (entry level takes the pressure first) collides head-on with the leveling narrative. If AI genuinely lifts novices to average, why are novice jobs the first to vanish? Freelance platforms posted the preview as early as the start of 2024: after ChatGPT's release, "posting for jobs that could be done with AI declined 21%," with "a 17% drop in graphic design jobs after image-creating AIs were released" (One Useful Thing). That one is relayed research — we have not traced the original paper — so read it as a directional signal.

There is a reading that holds all three datasets at once, and it hides in economics' oldest lesson: leveling is commoditization. When anyone with an AI subscription can deliver average quality, average quality stops commanding a price. What an employer's salary buys is no longer "a person who can do this task to the average standard" — AI supplies that slice directly — but "a person who can vouch for the result and carry responsibility for it." The traditional function of an entry-level job was precisely to trade templatable work for a seat at the table; when the template is supplied at subscription prices, the terms of that trade collapse. The causal chain is at present an inference, not a verified finding. Payroll data is observational — the interest-rate cycle and the post-pandemic correction in tech hiring remain alternative explanations nobody has ruled out, and the paper claims only to have controlled for firm-level shocks; nobody has run the randomized experiment that would rule them out. Moreover, "relatively fewer entry jobs" is evidence about quantity; "pay converging toward tool cost" still has no direct hourly-wage or billing-rate data behind it, and market pricing often lags a year or two. It is the most economical unified explanation on the table — not a bridge that has been load-tested.

Interlude: the kingmaker's crown was forged

Before assembling the final picture, one piece of evidence has to be taken off the table in public view. There has always been a third possible script: AI hands a small number of virtuosos order-of-magnitude gains and mints a new class of winners. This "kingmaker" story got its most beautiful evidence in late 2024: a paper by MIT doctoral student Toner-Rodgers claimed to have tracked 1,018 materials scientists at a large corporate lab, finding that AI nearly doubled the output of top researchers while the bottom third barely benefited. "Only those who know the domain can drive AI" — a perfect narrative, and it was cited everywhere.

In May 2025, MIT formally stated it had "no confidence in the provenance, reliability or validity of the data," asked for the paper to be withdrawn from the preprint server, and the author left MIT (TechCrunch). The kingmaker side's most-cited empirical anchor was fabricated.

This episode earns its own section because what it exposes is not one person's failure but the field's demand for narrative — strong enough that a made-up distribution story people wanted to hear could walk deep into a top journal's submission pipeline. And it calls for some honesty on our part: if you define "kingmaker" loosely as "the gap widens," then the divergence in the Kenya experiment is itself randomized-experiment evidence pointing in the kingmaker direction, and we can't pretend not to see it. We file Kenya under the judgment layer because the gap there widens from the capacity to select and execute advice. The classification follows the mechanism — not the fact that a few people capture order-of-magnitude excess returns. The narrow kingmaker script — a handful of AI virtuosos doubling output or better and taking the market — is left with only selection-riddled market observations: a consultancy reporting, for instance, a 56% wage premium for workers with "AI skills," where who counts as AI-skilled overlaps heavily with role and seniority, so no causality can be read off it — and the original report page blocks crawlers, so we could verify only secondhand accounts. The narrow kingmaker scenario currently has no clean evidence behind it, and stays unproven either way.

The final shape, answered by layer

With the forged anchor off the table, the board is clean enough to answer the question the 2023 researchers themselves posed. They flagged three possible end states for the distribution: a "leveler" where only low performers gain, an "escalator" where everyone rises proportionally, and a "kingmaker" where a few AI virtuosos take all (One Useful Thing). Three years on, here is where the race stands, in one picture:

Skill leveler (b. 2023, peer-reviewed 2026)
  (where a quality template exists, AI
  pulls the weak toward it — replicated
  across writing/support/law/creative)
 └ The final-shape contest (ongoing)
    (what is leveled work still worth —
    commoditized, escalator, kingmaker?)  ?
 ├ vs Path A, escalator: all rise in step
    (only RCT on seniors, METR: −19%;
    n=16, one domain — thin, zero support)
 └ vs Path B, kingmaker: few AI power
    users take all (Toner-Rodgers anchor
    disavowed by MIT; only selection-
    biased market data, no clean evidence)

Looking back, the snag was that the original question stacked three different layers into one. Unstack them, and each layer already has its answer:

At the task layer, leveling holds within the tested range — it is no longer a hypothesis. On tasks with a quality template, "the weakest gain most" has replicated across writing, customer support, law, and creative work, with all five papers now through peer review. Two flags belong on the record: the law study's quality gain was slight (the effect there was mostly speed); and the headline compression, 22 percentage points down to 4, has been measured in exactly one venue, BCG, with no independent replication of the compression itself. The direction is hard. The magnitude is still a single-venue number.

At the judgment layer, leveling fails and the gap widens. On tasks with no answer key — open-ended business decisions, hard problems in mature systems, diagnoses that need a gatekeeper — AI amplifies existing differences, because what it supplies is raw material, and the judgment that converts raw material into results has not been leveled. The escalator scenario finds zero support at this layer: the only randomized experiment to directly test senior practitioners measured the reverse — though with 16 people in a single domain, its weight is limited.

At the market layer, the question itself has changed. The leveled tasks commoditize and entry-level roles absorb the pressure; the human premium migrates to judgment, verification, and responsibility. The same author who wrote in 2022 that AI would not turn non-experts into experts described his own working method in 2026 this way: "I am closer to a patron. I describe what I want, I pay for it, and I judge the result" (One Useful Thing). Expert dependence now sits at the acceptance step rather than the production step. In business-strategy terms this is a textbook value-chain rearrangement: AI commoditizes the production link, and commoditizing one link raises the value of its complements: here, the capacity to gatekeep and to be accountable. That cuts in opposite directions for two classes of products (what follows is judgment, not verified market data): tools that sell "do the task to average quality" — self-serve support bots and their kin — face a one-way slope of price compression; tools that sell "decide which answer to use, and who answers for the error" are the ones positioned to charge more.

So the question actually worth staring at is not the three-way choice of "hire cheaper people and arm them with AI, or cut headcount, or groom a few virtuosos." It is whether each salary you pay is buying the leveled slice or the unlevelable slice. The market price of the first will keep converging toward an AI subscription fee. The scarcity of the second has only just begun to be priced.

Where we land

AI levels the skill distribution at the task layer, on tasks with a quality template — that direction has now replicated across domains in peer-reviewed journals and is no longer a hypothesis (the compression magnitude remains a single-venue number). The economic meaning of leveling is commoditization: pay for the leveled slice of labor converges toward tool cost (this is an inference — direct price evidence is still pending), the relative squeeze on entry-level jobs is its first observable signal, and the human premium migrates to what won't templatize — judgment, verification, and responsibility. The distribution question has shifted from "who gets leveled" to "what is the leveled slice still worth, and who owns the slice that wasn't."

What would prove this wrong — three observation windows this publication has set for itself, on a 12-month clock: (1) if a clean leveling replication appears in a template-free domain (open-ended business judgment, research problem selection), "the judgment layer doesn't level" gets revised; (2) the commoditization split-reading is overturned if the entry-level employment decline shows up at equal magnitude in AI-augmentation occupations, not just automation ones — that would be aggregate substitution, not repricing; (3) if large-sample wage data, after controlling for role and tenure, still measures a persistently widening premium for AI users, the kingmaker script upgrades from "no clean evidence" to "happening," and the shape question reopens.

What this means for you

The same conclusions land as different moves depending on where you sit.

People who manage and set pay: the thing to change is where the budget goes: re-sort the role list by "does it have a quality template." For the roles inside the template, the AI subscription fee is the market's pricing anchor. The highest-return use of a training budget is not teaching people to use AI — the leveling largely happens on its own — but developing and moving people into verification and gatekeeping positions.

If you're early in your career, face one fact squarely: "AI lifts you to average" is true — and precisely because of that, average is no longer a selling point. The narrowing of the entry ladder's bottom rungs is a number in payroll files, not a mood. Your bargaining power lives entirely outside the template: judging which answer is right, and carrying the responsibility when the answer is wrong.

People who read labor data for a living: remember that "AI is gutting white-collar work" and "AI has had no effect" can both be backed by data at the same time — the difference is whether the measurement is of the aggregate or of the structure. The next time a startling number crosses your desk, ask three things first: what layer it's measuring, whether the number is relative or absolute, and which version of the paper it comes from.

If you make decisions from research, note that this field's demand for narrative has been strong enough to produce an outright fabrication — and strong enough that the same paper's numbers drift between versions. Before citing anything on "who benefits from AI," check whether it has been published, whether it has been retracted, and which version you are holding. Every effect size in this essay has been checked against a primary page; the few claims for which we could reach only secondhand or relayed sources — the consultancy wage premium, the freelance-platform postings, the medical gatekeeping studies — are flagged as such in the text.


Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.


EN English edition|繁 中文版 Traditional Chinese →

Don't miss what's next. Subscribe to SecondSource:
← Newer SecondSource Daily Brief · July 18, 2026 | IBM's Worst Day in 115 Years as a Public Company — Will the Business That AI Is Taking Ever Come Back? Older → SecondSource Daily Brief · 2026/7/17 | "AI narrows the quality gap to average" just passed peer review — but where work has no answer key, the gap widens instead
buttondown.com
Powered by Buttondown, the easiest way to start and grow your newsletter.