SecondSource logo

SecondSource

Archives
Log in
Subscribe
July 21, 2026

SecondSource Deep Dive · 2026/7/20|Selling Tokens Finally Makes Money — Now Comes the Fight to Keep It

The story starts with a reversal. From 2023 through 2025, nearly all the money in AI was made by the people selling chips, power, and memory; for the

In brief

The story starts with a reversal. From 2023 through 2025, nearly all the money in AI was made by the people selling chips, power, and memory; for the model labs themselves, selling tokens was a losing business — Anthropic's 2024 inference gross margin was negative 94%. In the first half of 2026 that flipped. Anthropic officially announced annualized revenue crossing $47 billion and gave investors guidance for its first profitable quarter ever. And so a new debate has broken out, with prominent names on all three sides. One camp argues that scarcity plus a quality gap has handed the model labs the power to price to value — the low-margin era is over. Another counters that selling tokens is a commodity business by nature, and high margins will inevitably be competed away. Then there are those who distrust the numbers themselves: today's demand signal, they say, is already watered down. This piece brings all three cases' evidence to the same table for the first time. Our judgment: the flip is real, but "keeping it" currently holds only at the flagship tier — and what decides the outcome isn't a moat. It's a race between two clocks: how fast commoditization climbs the price ladder, against how fast the labs lock customers into their workflows.

Here is where this debate sits on our map:

Infra-layer capture era (2023-25)
  (value captured by chips/power/
  memory; labs lost money selling
  tokens (2024 inference margin -94%))
 └ Per-token economics flip (2026H1)
    (agentic demand × token-cost
    collapse; inference margin turns
    positive, first profitable quarter
    in sight — but does it hold?)  ?
 ├ vs Path A: durable pricing power
    (SemiAnalysis): scarcity +
    quality gap → price to value;
    low-margin era is over
 ├ vs Path B: commoditization default
    (Evans/Narayanan): undifferentiated
    tokens + zero switching → margins
    pushed toward cost
 └ vs Path C: inflated demand
    (Gurley/Chamath): today's demand
    signal carries subsidies and
    unsettled ROI; a closing capital
    window would expose it


First, get the books straight: "margins flipped" is three different ledgers

The biggest trap in this debate is the word "margin." The Anthropic margin figures in circulation run anywhere from 40% to 70%, and they are in fact three different ledgers, drawn up for different years under different definitions.

The first ledger covers 2025 and counts all costs. The Information reported that Anthropic cut its full-year 2025 gross margin forecast from 50% to roughly 40%, because inference costs on cloud infrastructure came in about 23% higher than expected (The Information). The second ledger counts only the inference business. SemiAnalysis has inference gross margins rising from 38% in 2025 to somewhere in the 60–70% range in 2026 (SemiAnalysis). The third ledger is a cost ratio shown to investors: compute spending per dollar of revenue fell from 0.71 in the first quarter to an estimated 0.56 in the second. That 15-cent improvement in compute cost arrived in the same quarter the company swung from roughly breakeven to $559 million in operating profit — though we don't have that quarter's revenue base, so we can't do the multiplication that would attribute the swing, and it would be wrong to call this the only lever (CNBC).

Keep the sequence straight: the downgrade came first and describes 2025; the improvement figures came later and describe 2026. Using the 40% to attack the 65%, or the reverse, is a category error. More important still: none of the three ledgers has been audited. The revenue side has an official announcement behind it (Anthropic); the margin side is entirely internal data and analyst models, and the profitable quarter remains guidance, not fact. Anthropic is preparing to go public, and the prospectus will be the court of final appeal for this fight over definitions.

So when this piece says "the flip," here is the precise meaning: the revenue flip has an official basis; the margin flip is an improving trend on which internal and analyst ledgers agree — unaudited. That distinction is the foundation for every judgment that follows.


The case that it holds: scarcity plus a quality gap means pricing to value

With the books straight, take the three cases in turn — starting with the bulls.

What it claims. The mechanism runs like this: agentic workflows have made each token's work more valuable. An agent task that runs for hours costs, at actual blended rates, roughly one dollar per million tokens, and delivers output worth dozens of times a human hourly wage. Meanwhile software optimization and a new chip generation have collapsed the cost of producing tokens. Demand value up, supply cost down — the space between is margin. And that margin won't be competed away, the argument goes, because compute scarcity means no single provider can serve the whole market, and the quality gap means open-source models can't handle the hardest work: a lab with frontier-quality models can price to the value it delivers, with no need for a price war.

Who's betting on it. SemiAnalysis's Dylan Patel. His words in May: "The age of low gross margins for frontier model providers is over" — and he judges that agentic demand has permanently raised the market-clearing price of tokens (SemiAnalysis). Two structural reinforcements: a July arXiv scenario analysis by Matsuoka argues that incumbents' fully depreciated fleets arrive like a conveyor belt of near-zero-cost capacity, leaving late entrants with a cost gap starting at 3.2x that never converges (arXiv); and NVIDIA and TSMC (the two upstream monopoly bottlenecks) are deliberately not pricing to scarcity, which in effect cedes the pricing headroom to the layer below them.

How hard the evidence is. The revenue side is official-grade: $47 billion annualized, five times higher in half a year. Willingness to pay is behavioral-grade: Uber's five thousand engineers burned through the company's full-year AI budget in four months, with heavy users spending $500 to $2,000 a month each (Fortune). A Sierra co-founder observes that top front-line engineers now burn $100,000 a year on tokens — the remark didn't specify whether that's Sierra's own team or its customers, so we relay it at face value. Current prices are also on this side of the ledger: as of July 20, not one flagship has cut prices: the GPT-5.6 flagship (the Sol tier) and Claude Opus sit at price parity of $5 per million tokens, though GPT-5.6's $5 is so far only a beta-tier price, not an official market-wide rate. The premium end has moved up, not down: a higher tier codenamed Mythos is priced at roughly five times the previous flagship — again a tiered, beta-layer price rather than a market-wide announcement, and the 5x magnitude itself still awaits official confirmation. Glean founder Arvind Jain is blunter: over "the last 6 to 9 months," he says, "every model actually increase[d] their per token price" (20VC) — a single-source claim, but one checkable against public price sheets. Even open-source cheapness needs a discount: Snowflake ran the same data-engineering tasks through the same system and the open model burned twice the tokens (X/Ramaswamy) — an 8x list-price gap shrinks to only about 4x on the bill.

Where it's weak. The margin figures are the company's own books — see the previous section. The profitable quarter is still guidance. And this is one company's story: a third-party estimate has OpenAI's margin falling from 40% to around 33% — but that is a full-cost estimate (low confidence) of OpenAI, whose consumer-subscription-heavy business isn't structurally comparable to Anthropic's API-heavy one; it is not the same ledger. Even so, generalizing one firm's flip into "the era is over" mistakes a single company for the whole tier. The scarcity pillar also has a clock on it: only a third of the compute promised for delivery this year is actually under construction — a gap of roughly 7GW — and the root cause is three-to-five-year lead times on transformers and other electrical equipment. That makes scarcity an execution delay, not a permanent structure, with a visible relief schedule in 2027–28 (SemiAnalysis).

What would prove this wrong. An actual price cut at the flagship tier. The WSJ reported in June that OpenAI was internally weighing steep cuts to fight for customers; as of July 20, none has landed. The nearer checkpoint is a reversion date Anthropic itself has announced: Sonnet 5 launched at a promotional $2/$10 per million tokens. If it doesn't return to the standard $3/$15 after August 31, the promotion has become a de facto price cut. Further out: if the audited margins disclosed in the IPO filing come in far below the internal ledgers, the entire flip narrative goes back on trial.


The case that it doesn't: selling tokens is a textbook commodity business

The commoditization camp looks at the same evidence and sees a trap with a long history.

What it claims. Economics has an old result: when two firms sell an undifferentiated good, even with only two sellers, the equilibrium price gets driven down toward cost — whoever raises prices first loses every customer. This camp argues that selling tokens is, in Arvind Narayanan's words, "an unusually pure version of the conditions that produce the paradox," with all five conditions met: models are highly undifferentiated, open-source models are already good enough, hundreds of benchmarks make equivalence transparent to buyers, everyone's cost of capital is similar, and switching means changing one line in a router config (Normal Tech). The historical base rate is on their side too: across six capital-intensive infrastructure industries, fiber's seven-year capacity explosion ended in roughly $2 trillion of vaporized market cap, and airlines have run 2–4% net margins for eighty years; only cloud and TSMC escaped. Then there's the arithmetic: $4–8 trillion of infrastructure investment, recouped at a 5% net margin over five years, requires $16–32 trillion in annual revenue — the high end is a quarter of today's world GDP.

Who's betting on it. This is the most crowded side. Narayanan, of AI Snake Oil, supplies the systematic argument above. Benedict Evans, in a July 9 essay, notes that 40–50% inference margins don't yet count training costs, and that "analogies don't have predictive value" — for pricing power to emerge, "something needs to happen that we don't see yet" (Benedict Evans). Databricks CEO Ali Ghodsi argues the model layer's endgame is TSMC-like — "very valuable" but "interchangeable" (BG2). Policy researcher Dean Ball points out the frontier premium lives only in a narrow window — "a significant fraction of that cost is recouped in the few post-release months" while a model is broadly available — before competition, distillation, and open source compress it (Hyperdimensional). And Michael Li, a winner of Dwarkesh Patel's blog prize, has the harshest version: the API is a railroad that runs at a permanent loss, and the money is in the real estate the railroad opens up (Dwarkesh blog prize).

How hard the evidence is. The historical slope is its strongest card: by the public estimates this camp cites (we have not independently verified each link), prices for equivalent capability have fallen roughly 90% since GPT-4 launched; the unit price of intelligence dropped more than 90% over the course of 2024; and GPT-4-level capability took about 15 months for open source to match. The current state of the gap: on the Artificial Analysis intelligence index, Kimi, Moonshot AI's Chinese open-weight model, is within 9 points of Opus — a historic low — and GLM 5.2, from the Chinese lab Zhipu AI, crossed the frontier threshold on three mutually independent kinds of evidence within a single month: enterprise deployment, hands-on reviewer testing, and independent replication by inference providers. And the price war has genuinely begun: Google cut its consumer AI subscription from $7.99 to $4.99, which the tech press called the opening shot of a subscription price war (TechCrunch); OpenAI opened a new low tier at $1/$6 per million tokens in July; Anthropic is running a limited-time promotion in the mid tier.

Where it's weak. This camp predicted in 2025 that margins couldn't rise — and in the first half of 2026, Anthropic's inference economics flipped anyway. At least inside the window where scarcity and the quality gap still hold, it lost a round. Its strongest theoretical version also carries its own escape hatch: Narayanan himself concedes that history's two exits from the commodity trap are "becoming functionally software or achieving market concentration" — and the labs are sprinting down the first route: Claude Code, multi-year enterprise contracts, digital employees embedded across a company's workflows. In his words, unless portability is explicitly built in, such a digital employee "effectively cannot be fired," with lock-in "that exceeds that of enterprise software at its worst." Put differently: even the bears' best theory agrees that the verdict turns on how fast lock-in forms, not on token prices.

What would prove this wrong. Flagship-tier pricing holding or moving up, while the lock-in metrics keep deepening: multi-year contracts, committed spend, product-layer revenue share. That's an observation window we've set ourselves — 12 months, already on our recheck list. Separately: if same-system comparisons keep showing open source behind on both quality and token efficiency for the hardest work, the "good enough" premise has to be narrowed.


The case that the numbers are inflated: demand isn't as hard as it looks

The third camp doesn't join the pricing-power argument at all. It questions the inputs.

What it claims. Two mechanisms. First, subsidized pricing: in Bill Gurley's version, every model company prices for market share, not cost, so the "demand exceeds supply" signal carries a subsidy component — when the capital window closes and repricing is forced, an air pocket appears in the demand curve (BG2). Second, unsettled ROI: Chamath Palihapitiya relays an internal measurement from his company's CTO — token spend (usage times price) is "doubling every 45 days," downstream productivity is up "maybe 5% max," and "we've effectively already asymptoted" (All-In). Sooner or later, every buyer faces the CFO's question: how much did this actually add to EPS?

Who's betting on it. Beyond Gurley, investor Gavin Baker has a cost-structure version: Google, running on its own silicon, is the lowest-cost token producer and can rationally run AI at negative 30% margins, subsidized by search cash flow — deliberately draining the economic oxygen out of the whole ecosystem.

How hard the evidence is. All-you-can-eat is genuinely dead: ChatGPT moved to $100 and $200 tiers with usage caps, and an OpenAI executive has said publicly that pricing will have to change substantially. An MIT survey from 2026 reportedly found 95% of enterprise AI pilots show no P&L impact (a secondhand relay, single source). The other half of the Uber story lives here: after the budget was gone, Uber's COO publicly questioned whether the money had been worth it (Fortune), and Coinbase's board has begun pushing back on token spend and tapping the brakes. The ROI interrogation is on the table — present tense, not prophecy.

Where it's weak. "Subsidy" has a rival explanation: Sourcegraph CEO Quinn Slack's rebuttal — single-source, not yet verified — is that low marginal cost plus heavy depreciation is simply what this business looks like, and asking when the subsidies end is like asking when an airline will stop selling economy class. The pricing record also runs against the subsidy story: it predicts prices heading down, but per-token list prices have risen, not fallen, over the past six to nine months. Most decisive is the broken transmission: the questioning is loud, but the money is still burning — the ROI interrogation has yet to show up in any lab's revenue curve.

What would prove this wrong. The capital window tightens and no air pocket appears. The two model labs' IPOs are a natural experiment: once public, the subsidy motive largely disappears. If the big buyers renew and expand their purchases anyway, the inflated-demand thesis gets downgraded.


Three cases, one table: margins are a race between two clocks

In the textbook, whether an industry keeps its profits is a question of pressure: rivals cutting prices, buyers gaining leverage, substitutes closing in. Lay July 20's facts over that frame and each of the three cases turns out to be right about one segment — and all three can be arranged on a single price ladder.

The lower rungs are already at war. Google cut consumer pricing 37.5%, OpenAI opened the $1/$6 low tier, Meta is fighting an API price war in the catch-up tier, Anthropic is running a mid-tier promotion. Commoditization is not a question of whether it arrives; it is already burning up the ladder's lower rungs. On this segment, the commoditization case is simply right.

The top rung is still rising. The two flagships sit at unbroken price parity, the tier above the flagship has moved up, and per-token list prices across the board have risen over the past half year. The top holds because three pillars stand at once: scarcity means a rival that wanted to buy share has no volume to buy it with — cutting prices buys nothing; the quality gap means the hardest work clears only at the frontier — substitutes can't get in; and the buyers' ROI interrogation hasn't yet frozen budgets — bargaining power hasn't formed. On this segment, the pricing-power case is simply right.

But all three pillars are moving. Scarcity now has a schedule — transformer lead times of three to five years put the relief point at 2027–28. At the same time, the quality gap has narrowed to a historic low, with open source pulling even on some axes. And the ROI interrogation is on the table, just not yet transmitted into revenue. So the right question is not "will commoditization happen." It is: can commoditization climb the ladder faster than the labs can lock customers into their workflows? That criterion isn't ours — the commoditization camp's own theory supplies it. The historical escapes from the commodity trap number exactly two: become functionally a software business, or achieve monopoly-level concentration. The labs plainly know it: Claude Code, multi-year committed spend, digital employees embedded across entire companies — all of it is escape engineering. Margin durability comes down to the relative speed of two things: how fast the escape engineering gets finished, and how fast the three pillars come down.

Where we land

The correct reading of "frontier margins won't be competed away" is layered. The per-token flip is real — official on the revenue side; on the margin side, an internal-and-analyst-consistent but unaudited improvement. High margin currently lives only in the flagship agentic tier, held up by three pillars — scarcity, the quality gap, and unsettled ROI — all three moving, none yet fallen. The verdict variables for the next 12 months (an observation window we set ourselves, already on the recheck list): First, whether the flagship tier sees an actual price cut — the nearest checkpoint is August 31, when Sonnet 5's promotional price either reverts to standard or doesn't, per Anthropic's own announcement; second, the compute delivery gap and the direction of reserved rental rates; third, whether open source pulls even on both quality and token efficiency for the hardest work in same-system comparisons; fourth, which ledger the audited margin in Anthropic's IPO filing lands on; fifth, whether token demand growth holds at 2x a year — the threshold set by Matsuoka's solvency corridor: if annual token demand growth falls below roughly 2x, the repayment math on existing compute investment stops working.


What this means for you

  • Investors allocating to AI: "The labs finally make money" and "token-selling is doomed to commoditize" are each half right. The actionable version is to value a lab's revenue in two layers: discount the API layer on commodity logic, and the product-and-lock-in layer — Claude Code, long-term enterprise contracts, digital employees — on software logic. When the prospectus lands, turn first to the margin-definition page, then to product-layer revenue share.
  • Enterprise CFOs: your bargaining power is growing fast on the ladder's lower rungs — the low tiers are undercutting each other, so route down everything that can be routed down. The hardest work will still command the top-tier premium for now. The Uber lesson isn't that you should stop using the tools. The lesson is that metered billing colliding with a culture of internal usage leaderboards makes the bill grow by itself. Install the cost discipline — route down what can be routed down — before the culture sets your prices for you. (Routing architecture and model selection are the product manager's call — see below.)
  • AI product managers: yours is the call on routing architecture and model selection — build dashboards for agentic usage, tier work by difficulty and route each tier to the matching model class, and watch whether your own product's pricing needs to move with the flagship tier. These are product decisions, not calls a CFO makes.
  • Clouds and the compute supply chain: the top rung's first pillar is scarcity, and scarcity's root cause lies in transformer and power delivery bottlenecks. Your delivery speed is the stopwatch on the labs' margins: the faster you deliver, the faster commoditization climbs the ladder. Track your own transformer and power lead times against the 2027–28 relief schedule this piece cites — if you are closing that gap faster than expected, you are the one accelerating commoditization. Those are two sides of the same coin.
  • Model-lab strategy watchers: two numbers suffice. First, the flagship list price — it measures whether the three pillars have fallen. Second, committed spend and product-layer revenue share — they measure whether the escape engineering is finished. Every other number is a footnote to these two.

Look at the tree once more. What this piece updates is how to read the fork: the three cases move from "pick one" to three segments of fact on a single price ladder — with five checkable verdict variables and two short-window checkpoints attached.

Infra-layer capture era (2023-25)
  (value captured by chips/power/
  memory; labs lost money selling
  tokens (2024 inference margin -94%))
 └ Per-token economics flip (2026H1)
    (agentic demand × token-cost
    collapse; inference margin turns
    positive, first profitable quarter
    in sight — but does it hold?)  ?
 ├ vs Path A: durable pricing power
    (SemiAnalysis): scarcity +
    quality gap → price to value;
    low-margin era is over
 ├ vs Path B: commoditization default
    (Evans/Narayanan): undifferentiated
    tokens + zero switching → margins
    pushed toward cost
 └ vs Path C: inflated demand
    (Gurley/Chamath): today's demand
    signal carries subsidies and
    unsettled ROI; a closing capital
    window would expose it

Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.


EN English edition|繁 中文版 Traditional Chinese →

Don't miss what's next. Subscribe to SecondSource:
Older → SecondSource Morning Brief · July 20, 2026 | An AI Says It Aced the Math Olympiad: the Proofs Are Checkable; the Contest Run Is Its Word Alone
buttondown.com
Powered by Buttondown, the easiest way to start and grow your newsletter.