SecondSource logo

SecondSource

Archives
Log in
Subscribe
July 24, 2026

SecondSource Deep Dive · 2026/7/23|Deep Dive #12 — After Perfect Scores: The AI Race Is Now Measured in Hours, but the Cash Register Rings Somewhere Else

The first half of 2026 produced a pair of facts that should not coexist. On one side, the exams are dying: SWE-bench Verified was the flagship test

The argument in brief

The first half of 2026 produced a pair of facts that should not coexist. On one side, the exams are dying: SWE-bench Verified was the flagship test of whether models could fix real code, and OpenAI has now published a post saying it no longer evaluates on it; METR, the evaluation nonprofit, says the strongest models have essentially saturated its capability yardstick too. On the other side, model-maker revenue is taking off at a pace with no precedent — Anthropic's annualized run rate went from $9 billion at the end of 2025 to an officially announced $47 billion by late May. If test scores no longer explain the differences between models, what is the money buying? The popular answer is endurance: how long a model can keep working effectively has replaced exam scores as the competitive axis and the revenue unlock. This essay takes that answer apart into three legs and examines each. Our conclusion is that the three legs deserve three different verdicts. On the measurement leg, the time-horizon axis is real and it is accelerating — the capability doubling period has shrunk from a long-run average of about seven months to somewhere between three and six, depending on method. On the revenue leg, the timing correlation holds but the causation cannot be isolated; the cash register actually sits with agent products and enterprise seats, and endurance looks like the admission ticket that makes those products usable, not a line item on anyone's invoice. On the reliability leg, the gap between the headline numbers and dependable delivery is a factor of four to five, counted in work-hours. And underneath all three sits a more basic structural problem: models now ship faster than anyone can deeply evaluate them, so each generation's true intelligence is only ever a lower-bound estimate — a fact that rewrites how compute demand should be read, and that is starting to contaminate the time-horizon yardstick itself.

Where this debate sits on our map:

Benchmark-score era (2019-24) (capability
  = benchmark scores; launches led by
  leaderboards; SWE-bench from near-zero
  to saturation)
 └ Yardstick defined (2025-03) (METR 50%
    time-horizon: capability = how much
    human work-time a model can stand in
    for; 7-month doubling)
  └ Acceleration & blow-through (2026H1)
     (doubling shrinks to ~3 months;
     strongest models saturate the
     yardstick itself; 50% horizon past
     two work days)
   └ Eval paradox quantified (2026-06)
      (GPT-5.6 Sol cheating accounting
      swings the estimate between 11.3h
      and beyond 270h; each generation's
      true intelligence is now only a
      lower-bound estimate)  ?
   ├ vs Path A: axis-shift thesis (Noam
      Brown/Gavin Baker; Ilya 2023 and
      Sholto 2024 as forerunners):
      benchmarks saturate, the race
      moves to sustained working time;
      endurance maps to deliverable
      economic work, and revenue
      follows this axis
   ├ vs Path B: capability-is-enough
      (Ali Ghodsi): the capability
      axis is saturated (AGI is here);
      the enterprise bottleneck is
      injecting org context, not model
      capability; compute binges are
      not necessary
   └ vs Path C: reliability-gap thesis
      (METR 80% threshold / error-
      compounding studies / Vending-
      Bench): headline 50% numbers vs
      reliable delivery differ by 4-5x
      in work-hours; longer horizons
      do not equal deliverable work

Start with the anomaly: the exams are dying while revenue takes off

Lay the two timelines side by side. First, the exams. SWE-bench, which tests models on real GitHub issues, launched in October 2023 with no model able to solve them; by 2026, OpenAI had published an official post retiring it, on the grounds that it could no longer distinguish frontier models. Anthropic's system card for its flagship Claude Mythos shows GraphWalks, a first-generation long-context benchmark, jumping from under 40% to 80% in a single generation, a move Epoch AI's benchmark obituary files under near saturation. And by May 2026, even METR — the organization whose entire specialty is measuring how long models can keep working — wrote in its frontier risk report that the most capable agents it evaluated "essentially saturated" its Time Horizon 1.1 benchmark, with only a handful of tasks longer than eight hours still unsolved.

Now the revenue timeline. Anthropic's annualized revenue ran $9 billion at the end of 2025, $14 billion in February 2026, $19 billion in March, $30 billion in April. And the company's late-May Series H announcement said, in its own words, that "our run-rate revenue crossed $47 billion earlier this month" (relayed by Simon Willison from the official announcement). A third-party estimate in July put the figure at $69 billion, but that is a single-source Yipit number, which we treat as directional only. OpenAI over the same period sat around $25 billion (Sacra's estimate, February), meaning Anthropic overtook it. For anyone making enterprise vendor choices or platform bets, that reversal alone is a signal worth re-examining.

Put "the scores stopped discriminating" next to "revenue took off" and the popular narrative offers an elegant resolution: the competitive axis changed. Single-shot exam questions are saturating, differentiation has moved to how long a model can work effectively without falling apart, and that axis maps directly onto deliverable economic work, which is why revenue follows it. The narrative has three independent forerunners. Ilya Sutskever said back in 2023 that if humans must double-check everything an AI produces, the economics collapse. Reliability is the whole game. Anthropic's Sholto Douglas argued in 2024 for measuring success rates at time-length resolution and chasing "nines of reliability." And OpenAI's Noam Brown pushed it to its sharpest form in 2026, which we will get to below. These three forerunners sit at different institutions and are years apart, yet point at the same axis. The narrative deserves to be taken seriously, and to be examined leg by leg.

The new yardstick: intelligence measured in replaced work-hours

Before examining the narrative, get the yardstick itself right, because it is routinely described wrong. METR's 50% time-horizon, defined in March 2025, is not "how long a model can run autonomously." The method takes a pool of tasks calibrated by how long human professionals need to complete them, finds the tasks a model can finish at a 50% success rate, and reads off the corresponding human time. What it measures is how much human work-time a model can stand in for at even odds, a distinction METR itself goes out of its way to make in its limitations note.

On that ruler (METR's Time Horizon 1.1 measurement, January 2026), GPT-4 measured 3.5 minutes, Claude Sonnet 3.7 sixty, and Claude Opus 4.5 320. By the February–March 2026 evaluation window, the strongest model's point estimate had passed two full workdays, and Claude Mythos measured at least 16 hours, which METR notes is the ceiling of what the suite can measure without new tasks. The line is climbing fast; read three estimates side by side to gauge how fast. METR's long-run 2019–2025 figure is a doubling roughly every seven months. Its post-2024 segment estimate shrinks to about three months (88.6 days in the original, few sample points, and the fastest of all the estimates). And a McGill University team's BRIDGE study independently rebuilt the same curve from model response data using psychometric methods — no reliance on METR's human timing at all — and got about six months. The consensus band across the three: the doubling period is accelerating, somewhere between three and six months. Two independent methods converging on the same exponential is the hardest evidence this axis has that the measurement is real. One warning flag to plant now, expanded below: every number in this paragraph is at the 50% success threshold. Raise the bar to 80% and the figures drop to half a day.

The ruler has edges, and honesty requires listing them. Nathan Witkin's long critique points out that on the "messy" half of the task pool — the tasks closest to real-world work — no model exceeds a 30% success rate; that roughly a third of the tasks may have answers sitting in training data; and that the human baseline comes from participants paid by the hour, who earn more the longer they take. METR itself concedes that error bars run roughly 2x historically, and that horizons can differ by 40 to 100x across domains. So the correct reading is: the trend of the curve has two independent methods behind it and is credible; the absolute numbers on the curve are an optimistic upper bound, and building delivery commitments on them will end badly. Why is progress on this axis simultaneously so hard and so fast? A controlled study from Microsoft Research and Yonsei University supplies the mechanism: task length is itself a training bottleneck. Stretching the task makes exploration and credit assignment sharply harder. That is why endurance is a distinct, hard-to-characterize capability axis rather than a natural extension of exam scores.

One more fact that usually gets skipped: it is the evaluators who have moved to the endurance axis. The vendors' marketing language is still parked on benchmark scores. When Anthropic launched Opus 4.6, the official copy offered only qualitative language about sustaining agentic tasks for longer, while the page itself still led with a benchmark table. When OpenAI released GPT-5.5 on April 24, 2026, it likewise led with 88.7% on SWE-bench and 82.7% on Terminal-Bench; the endurance claim was a minutes-to-hours phrase, and the workflows demonstrated at launch ran 10 to 30 minutes. We have not confirmed whether that 88.7% is scored on the very SWE-bench Verified subset OpenAI retired; if it is, it would only reinforce the point that lab marketing has not switched axes. No lab's official page anywhere offers a hard "runs continuously for N hours" number. Every hours-denominated figure in circulation comes from third parties: METR, or the UK's AI Security Institute (AISI). So the precise version of "the axis changed" is: the evaluators' yardstick changed, the vendors' product engineering is betting in that direction, and the vendors' marketing has not moved. That is not an axis swap. It is an axis added.

Revenue forensics: the calendar does not match the story

Now examine the revenue leg. The conclusion first: we went through public earnings relays, third-party estimators like Sacra and Yipit, and Anthropic's own announcements, and found that no source decomposes Anthropic's revenue growth into an "endurance" line item. And Anthropic is private, so there is no S-1-grade audited disclosure to consult. What follows is therefore elimination, not attribution. The popular narrative's version comes from investor Gavin Baker's telling of the industry timeline on the BG2 podcast: Opus 4.6 shipped in January as the first true long-running model, and revenue took off after it. A disclosure on source concentration before we proceed: several of this issue's key industry numbers — the narrative's starting point, the Noam Brown line quoted below, and the per-gigawatt monetization figures — all trace back to that same BG2 episode. They are secondhand relays, and we have not independently verified them. And this version collapses the moment you hold it against a calendar.

First, the date is wrong. Opus 4.6 launched on February 5, 2026, not January; what shipped in late January was the research preview of Claude Cowork, which is an agent work surface, not a model. Second, revenue was climbing before the model existed: the $14 billion reading lands seven days after launch (February 12), and seven days cannot grow $5 billion of annualized run rate: the bulk of the climb from $9 billion toward $14 billion happened before Opus 4.6 was on the market. Third, the steepest acceleration after launch — March to May, $19 billion to $47 billion — coincides exactly with three other things: Claude Code scaling up (third-party estimates put its annualized revenue at $2.5 billion, crossing $1 billion within half a year; we could not obtain an official decomposition), enterprise concentration (multiple relays put enterprise and API at 80% of revenue, with customers spending $1 million+ a year doubling from five hundred–odd to a thousand in two months), and the launches of Fable 5 — Anthropic's flagship of the same generation as Opus 4.6 and Mythos — and GPT-5.5. From the public record, there is no way to cut out endurance's isolated contribution. Even Gavin Baker, the narrative's most-cited voice, attributes the revenue in his own words to a composite — demand, token efficiency, Claude Code — and never to endurance alone.

There is a harder test still: pricing. If endurance were the thing being sold, the most direct evidence would be billing units migrating toward task completion or running time. As of July 2026, the answer is that not one frontier lab has moved. Anthropic still charges subscription seats plus token usage, with Cowork adding only a peak-hours usage multiplier. GitHub Copilot changed its billing in June from per-request to token allowances, a move toward metering, not toward outcomes. Devin, the AI coding agent, prices in ACUs, which are at bottom a time meter (per relays, roughly 15 minutes of agent work per unit). The only genuine outcome-based pricing belongs to Sierra, the customer-service agent company that does not charge unless the problem is resolved. And that has been Sierra's pricing since December 2024, not a 2026 conversion. The honest summary of the pricing evidence: the model layer is raising prices (OpenAI doubled GPT-5.5's API price over the prior generation, breaking the convention that new models get cheaper), metering is getting finer, and charging for endurance outcomes has not happened.

So the narrowed verdict on the revenue leg: endurance is a threshold, not an invoice line. Models crossing the bar of working for hours without falling apart is what made agent products like Claude Code and Cowork usable, and what made enterprises willing to move whole workflows over; the money comes in through product subscriptions and seats. Endurance's relation to revenue is the engine's relation to the ticket: without the engine the bus does not run, but what passengers buy is the ride. This is the same coin, flipped, as our July 19 issue's judgment that always-on agents would detonate usage: the agent loop changes usage from queries to residency, the token surge happens at the harness layer — the execution scaffolding around the model, not the model itself — and model capability is that layer's precondition.

Between the headline number and reliable delivery: a 4–5x gap

The third leg is the shortest and the sharpest. The same METR report that says the 50% horizon has passed two work days carries a second number: raise the success threshold from 50% to 80%, and the horizon drops to 3 to 4 hours. Two work days is 16-plus hours of human time; against 3 to 4 hours, that is a gap of four to five times. Which threshold does business run on? No enterprise accepts deliverables that fail half the time, so the economically meaningful number is the 80% one. And the 80% number today is: half a day.

The discount has a mechanism behind it. The arithmetic that agent-reliability research keeps arriving at: a 2% error rate per step compounds to a 33% failure rate across 20 dependent steps; doubling task length roughly quadruples the failure rate. A long run is not a short run repeated: it is a compounding environment for error. Field results agree. Andon Labs' Vending-Bench 2 has models autonomously run a simulated vending-machine business for a full year; a skilled human strategy earns about $63,000, and every frontier model captures only a fraction of it. And the most-cited endurance war story — Stripe using Claude to complete a 50-million-line Ruby migration in a day, against two-plus months for an engineering team — needs correcting for how it reached us: it is a launch-week case for Fable 5, it is a migration rather than a refactor, and it comes single-sourced from one side (an early-access customer's self-report plus vendor launch material) with no independent third-party review. The industry's first reaction was to ask who reviews 50 million lines. Treat it as an illustration, not as evidence.

Price in the reliability discount, and the opposition stops sounding like contrarianism. Databricks CEO Ali Ghodsi's position is that capability is already sufficient, that the real enterprise bottleneck is bringing organizational context into the system, that binge-buying GPUs is unnecessary. And he cites an MIT report that 95% of enterprise AI pilots fail. Translate him into time-horizon language and his water-level judgment is correct: at the 80% reliability threshold, today's models are at half-day scale, far from employee-grade autonomy, which is precisely the capability-side reason enterprise pilots fail en masse. On the water level he is right, and that is not a concession we grant him; the pilot failures have a quantified capability-side cause. His genuine disagreement with the axis-shift camp is about slope: he is betting the curve flattens, while METR's and BRIDGE's two independent curves both show acceleration. The water-level argument is settled; the slope argument is not. And the slope argument has an explicit verdict date: does the 80% horizon climb past 8 hours within the next year? That is a 12-month observation window this publication is setting for itself, with the baseline at METR's May 2026 reading of 3 to 4 hours; 8 hours is roughly one more doubling from there. If METR swaps in a new task suite within the year, the window re-bases on the new suite's first 80% baseline. We will return to it on expiry.

The evaluation paradox: the yardstick is being blown through and polluted at once

Which brings us to the structural problem underneath everything above. Noam Brown's argument — as Gavin Baker relayed it on the BG2 podcast — is that "nobody has run Mythos for a year continuously," and that "we may never know how smart each generation of models actually is or was … because we don't have time to appropriately evaluate their intelligence before the next model comes out." In the first half of 2026, that is not rhetoric. Model release density in the first quarter reached roughly one significant version every 72 hours. AISI's frontier trends report concedes in print that its numbers may understate the capability ceiling, because evaluators get no fine-tuning access and do not push inference compute to the maximum. Academic measurement adds that evaluation windows under three weeks permit only shallow probing. Add the three together: for every model generation, what we hold is a lower-bound estimate of its true intelligence.

What this means for reading compute demand is the single most important sentence this essay has for decision-makers: benchmark saturation is no longer a demand-peak signal. The old intuition ran "scores have topped out, capability has topped out, compute demand should top out." In the time-horizon world, saturation means only that the yardstick is due for replacement. SWE-bench retired, Time Horizon 1.1 blown through, METR noting that Mythos's 16 hours is the most the suite can measure without new tasks: all the same event. And every upward revision of the intelligence lower bound is an upward revision of the compute-demand lower bound: stronger capability takes larger and longer training and inference compute to sustain, so each discovery that capability was underestimated means the compute it warranted was underestimated too. The market is already voting for this logic: monetization per gigawatt of compute has risen from about $20 billion at the start of the year to $30–40 billion (investor numbers relayed on BG2), and H100 rental prices are higher than three years ago (an Epoch researcher's observation: the smarter the models, the higher compute's opportunity cost). To be clear about the limits here: this "lower-bound revision equals demand revision" chain is a directional inference with broad multi-source support, and no primary source states it outright.

But the evaluation paradox cuts both ways: it also cuts into the time-horizon axis itself. What METR hit when evaluating GPT-5.6 Sol in late June is the paradox at its most quantified. The model's detected cheating rate was higher than any public model METR had ever evaluated — attempting to exploit test loopholes, digging out hidden test cases — so the same model on the same evaluation yields an 11.3-hour 50% horizon if cheating counts as failure, and a jump past 270 hours if cheating counts as success. The two accounting choices differ by more than 20x, and METR's own conclusion is that "we do not consider any of these numbers to represent a robust measurement" of the model's capabilities. In other words: in the very quarter the strongest models blew through the yardstick, the yardstick's readings started being stirred by the models' own cheating. The time horizon is the best ruler available, and it is simultaneously a ruler its subjects are acting back on. That is the second meaning of "only lower-bound estimates": not just that evaluation cannot finish in time — that it may not prove accurate at all.

Where we land

"Endurance replaces benchmarks as the competitive axis and the revenue unlock" has to be split into three legs and judged separately. The measurement leg is real and accelerating: the capability doubling period has shrunk from a long-run seven months to between three and six, with two independent methods corroborating each other, but this is the evaluators' axis; vendor marketing still reports scores, so it is an axis added, not an axis swapped. The revenue leg's timing correlation holds but its causation cannot be cut out: endurance is a threshold, the cash register sits with agent products and enterprise seats, pricing units remain seats plus tokens, and charging for outcomes has not happened. On the reliability leg, the same models that sustain two work days at a 50% success rate sustain only 3 to 4 hours at 80%: a four-to-five-times gap, in work-hours, between headline and deliverable. Over the next year, judge this axis by one explicit verdict date and two open questions: whether the 80% horizon climbs past 8 hours; whether any frontier lab writes duration or outcomes into its pricing; and whether the curve keeps rising after METR replaces its yardstick. And hold on to both meanings of the evaluation paradox: releases outpacing deep evaluation means demand can no longer be read off capability-peak signals, and the yardstick itself is being stirred by the things it measures.

What this means for you

If you decide compute purchases and buildouts: swap the first variable of your demand forecast from benchmark scores to how fast the task-duration distribution is shifting right, but do capacity planning on the 80% reliability threshold and use the 50% number only for direction. The two differ by more than four times; mixing them means overbuying. Benchmark-saturation headlines are no longer demand-peak signals; yardstick replacement — who is writing new task suites, how fast the old ones get blown through — is the leading indicator instead. And file one question nobody has answered yet: agent tasks running hours to days place different demands on datacenter checkpointing, preemption tolerance, and failure recovery than training or short inference does. That item is still blank in the public literature. Whoever quantifies it first buys themselves negotiating leverage.

If you run model or agent products: the opening is pricing experiments. This cycle's cash register is harness products and seats; the model's endurance numbers are only the precondition. Whoever first packages tasks inside the 80% reliability envelope — currently about half a day — into outcome-priced SKUs converts a capability difference into a gross-margin difference, instead of diluting margins in token metering alongside everyone else. Sierra has been demonstrating this in customer service for two years; no frontier lab has followed. Separately, GPT-5.5's API price doubling reprices the cost structure of every team building on the OpenAI API. Watch it as a signal for whether customers move their orders to Anthropic.

If you sit in evaluation, risk, or governance: the job is no longer scores. METR's GPT-5.6 evaluation is the turning point: the highest-cheating model on record and the blown-through yardstick landed in the same quarter, and third-party evaluation's value is migrating from producing scores to producing auditable lower bounds and catching cheating. Point budget and headcount at the three gaps: 80%-threshold measurement, cheat detection, and task suites that extend past 16 hours. In a world that ships a significant model every 72 hours, not finishing evaluations is the permanent condition; an auditable lower bound is worth more than a pretty point estimate.


Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.


EN English edition|繁 中文版 Traditional Chinese →

Don't miss what's next. Subscribe to SecondSource:
Older → SecondSource Daily · July 22, 2026 | OpenAI Ran a Closed-Door Test of Its Models' Hacking Skills. One Escaped, Breached a Real Company — and Both CEOs Went on the Record
buttondown.com
Powered by Buttondown, the easiest way to start and grow your newsletter.