SecondSource logo

SecondSource

Archives
Log in
Subscribe
July 20, 2026

SecondSource Morning Brief · July 20, 2026 | An AI Says It Aced the Math Olympiad: the Proofs Are Checkable; the Contest Run Is Its Word Alone

At a glance

  • An AI startup says its system solved all six problems at the International Mathematical Olympiad — the world's top competition for high-school mathematicians — for a perfect score. The published proofs can be machine-verified by anyone; the claim that it did so autonomously under contest conditions rests, for now, entirely on the company's own account. Four days on, officials and major media are silent. That confirmation gap is itself the signal.
  • Anthropic, the maker of Claude, officially announced a $65B raise at a $965B valuation, with annualized revenue past $47B. In the same round, three memory giants — Micron, Samsung, and SK hynix — invested. AI companies used to chase suppliers for capacity; now suppliers are using equity to lock in their biggest customer.
  • Our deep dive, finished overnight: selling tokens is profitable for AI model vendors for the first time, but the fat margins currently hold only at the priciest flagship tier, and all three pillars propping them up are shifting. The nearest checkpoint is September 1, when we learn whether Anthropic let Sonnet 5's promotional pricing revert.

About this issue's material: it draws on our research daily brief of July 20, 2026. Three genuinely new items today (two of them anchored to X posts from a single account — OpenAI President Greg Brockman's — that dependence flagged upfront), plus two verification updates on older events (event dates May 28 and June 15, labeled per item). The retrospective material comes mostly from re-reads of 2024–25 Dwarkesh Podcast interviews; they all trace to the same show as their only source and do not corroborate one another, though every one carries its original air date. Overnight we scanned 44 new pieces and separately collected 17 original X posts, yielding 21 linked receipts in this issue.

Today's leads

1. [This week] (events July 15–16; company published the proofs July 16) Six problems solved for a perfect 42/42, and the claimant is a startup, not one of the two big labs. The 67th International Mathematical Olympiad (IMO) ran July 15–16 in Shanghai. Afterward, Axiom Math, a startup specializing in AI theorem proving, published on GitHub complete formal proofs of all six contest problems by its system AxiomProver, claiming a full sweep and a perfect 42/42 (GitHub, Jul 16). The baseline for comparison: in 2025, OpenAI's and Google DeepMind's models publicly posted 35/42, at the gold-medal line. Note that the two results sit on different evidence grades: last year's score was openly published; this year's 42/42 has only the company's word behind it, so its standing depends on the verification below. If confirmed, competition math will have moved from "gold medal" to "perfect score" within a year, and a startup did it. But one detail complicates the picture: the company's self-reported per-problem times range from 24 minutes to over 14 hours, while human contestants get 4.5 hours a day for 3 problems. "Fourteen hours for one problem," if accurate, would fall outside the human contest's time limit. OpenAI President Greg Brockman posted on July 19: "Feels like a watershed moment for advancing mathematics. Scientific and medical advances which can really improve people's lives feeling very close now." The timing fits, but the post does not name the event; the connection is our inference (X / @gdb, Jul 19). Verification: the evidence structure here is asymmetric. The proofs are written in Lean 4, a formal proof language, so anyone can machine-verify that the proofs themselves are correct — no trust in the company required. But "within contest time limits, no human intervention, done autonomously by the AI" rests solely on the company's say-so. This is the only source we have seen; everything around it is reposts, which do not count as independent confirmation. Across the 51 tech and finance outlets this brief tracks, plus the wider wire-syndication landscape, four days on, not one independent confirmation has appeared. We book it conservatively as an unverified claim, with the upgrade condition stated: confirmation of the contest conditions by IMO officials or any first-tier outlet — either one suffices. Judgment update: the formal-proof lane already has a prior rung this year. First, what an agent is: an AI program that autonomously carries multi-step tasks through to completion, without a human issuing instructions step by step. In May, a research team's agent system autonomously solved 9 of the 353 Erdős open problems — the set of unsolved problems left by the Hungarian mathematician Paul Erdős — at a cost of a few hundred dollars per problem (arXiv, May 2026). Today's claim, if confirmed, is the next rung: from "solving open problems" to "a perfect contest score." Why this is worth watching: mathematical proofs can be machine-checked step by step, making math the cheapest domain in which to verify anything, so AI capability gains will surface here first. Brockman's line about scientific and medical advances feeling close is exactly this spillover — math's gains moving outward into other fields. If autonomous formal proving really is at this level, it is a directly usable capability jump for AI proof assistants and research-agent product lines, worth a re-rank of priorities. What to watch: the IMO's official record or any first-tier outlet confirming the contest conditions; until then, treat this strictly as the company's unverified, one-sided account.

2. [Evidence update] (event May 28; official release verified word for word today) Anthropic's $47B annualized revenue upgrades from analyst relay to official record; in the same round, the three memory giants take equity in their biggest buyer. Anthropic officially announced its Series H round on May 28: $65B raised at a $965B post-money valuation, with the press release stating that run-rate revenue crossed $47B earlier that month. Run-rate takes a single month's revenue and multiplies by 12 — an annualized gauge, not audited full-year revenue, and the difference has to be flagged (Anthropic press release, May 28). We finished checking that official release word for word today. The point is not the event itself — it happened in May — but the upgrade in provenance: the earlier "$44B+" figure came from SemiAnalysis, a semiconductor and AI-infrastructure research firm, whose same-methodology number for 2025 was $9B, roughly a 5× jump in half a year (SemiAnalysis, May 2026). The official number now lines up. Two structural signals sit inside the round: the $15B of hyperscaler investment previously committed has landed (including Amazon's $5B); and Micron, Samsung, and SK hynix joined as strategic investors. Verification: an official press release is the strongest possible evidence that "the company announced this"; the figures themselves remain company-reported and unaudited. The analyst gauge and the official gauge now agree; that is what earns this line its upgrade today. Judgment update: from 2023 to 2025, capital flowed one way — AI companies fighting upstream for capacity. Suppliers taking equity to bind their largest buyer is the first signal of capital flowing backward, from the infrastructure layer into the model layer. HBM — high-bandwidth memory, stacked DRAM — is the scarcest component in AI chips right now, and who gets priority in its next capacity allocation is worth watching against this new shareholder list. Memory capacity is finite; with suppliers now holding Anthropic equity, where other model vendors land in the queue for the next HBM allocation bears watching. For now this is an open question: there is no evidence that the round carries exclusivity terms or capacity guarantees. Read this together with the deep dive below: it upgrades the revenue side of that margin analysis; the margin side still runs on three separate ledgers — see that section.

3. [Evidence update] (product launched June 15; independent corroboration aligned today) AWS now lets content sites charge AI crawlers: third-party measurements split its two marketing numbers into one credible, one doubtful. AWS switched on "AI traffic monetization" in its web application firewall (WAF) service on June 15: content owners can charge AI crawlers per request. When a crawler hits a charging rule, the server replies "HTTP 402 Payment Required" along with a machine-readable price list; settlement runs in dollar stablecoins over the open x402 protocol; the service recognizes 650-plus AI crawlers; AWS takes no cut (AWS official blog, Jun 15). x402 is simply an open payment protocol built on top of the HTTP 402 "Payment Required" status code. The structural problem behind all this: traditional search crawlers index content and send traffic back; AI crawlers take the content, generate summaries, and send almost nothing back, so publishers carry the serving cost and get no readers in return. AWS offered two demand numbers to make its case, and we took each to third-party data. "AI crawler traffic up 300% year over year": Akamai, one of the world's largest content delivery network (CDN) operators, published the same figure in April from measurements on its own network, and has no financial stake in echoing AWS's numbers (Akamai press release, Apr 8); Cloudflare, Fastly, and Imperva (a web-security and bot-management vendor) each measure on different bases but point the same direction, some even higher; among peer measurements, AWS's number is actually the conservative one. "AI crawlers account for over half of traffic on many content sites": no independent source has measured this on the same basis; the closest is Imperva's finding that all bot traffic — not just AI — runs at 53%, with both numerator and denominator different, so it cannot serve as corroboration. Verification: the first number upgrades today to multi-party, independently measured; the second stays filed as vendor marketing. Judgment update: two directly usable lines for business leads at content platforms and publishers: quoting "AI crawler traffic tripled in a year" now has three-plus independent measurements behind it; and in negotiations, use the traffic share measured on your own site — not a denominator the vendor hands you.

4. [Today] OpenAI's president describes the key form factor of ChatGPT Work: agents run in the cloud, so the work continues with your laptop closed. ChatGPT Work is OpenAI's agent product for the workplace. Brockman wrote on July 19 that "one of the best features of ChatGPT Work is that it runs in the cloud, meaning that it works from mobile, with your laptop closed," adding that it is "kinda crazy how long the main way to get the magic of agents has been while leaving your laptop cracked open" (X / @gdb, Jul 19). Today's mainstream coding and office agents mostly tie long-running tasks to a local machine that stays on and online: if the agent runs for ten minutes, the laptop stays open for ten minutes. Move the run to the cloud, and hours-long background tasks can detach from the local machine for the first time and keep executing. Verification: a first-hand account from the product's own leadership, with the marketing incentive that implies; the post covers only this one feature, with full specs and pricing unseen. We log the direction. Judgment update: for enterprise IT and tool selection, a new procurement criterion is taking shape: does a long-running task require a human keeping a laptop open beside it.

Also notable

  • [Trend watch] (original June 2024; resurfaced July 19 via a short clip) The physical-security argument from Leopold Aschenbrenner, former OpenAI researcher and author of "Situational Awareness": a 100-gigawatt-class data center is a fixed, locatable, high-value target — as the podcast exchange put it, a single cruise missile would be enough to take one out. He now runs an AGI-themed hedge fund, so his stated stance aligns with the narrative he's promoting (Dwarkesh Podcast, Jun 2024)
  • [Trend watch] (original June 2024, same interview; relayed, unverified) Also from that conversation: DeepMind's safety framework grades security into levels 0–4, with 4 meaning resilience against state-level actors, and Aschenbrenner said DeepMind at the time self-assessed at level 0. This is a single person's relay; we have not yet checked the original document (Dwarkesh Podcast, Jun 2024)
  • [Trend watch] (original February 2025; we only recently read it) Microsoft CEO Satya Nadella on Muse, the company's gaming world model: it trains on gameplay data and generates playable worlds in real time. "What YouTube is perhaps to Google, gaming data is to Microsoft," he said — recasting years of gaming acquisitions as a training-data moat (Dwarkesh Podcast, Feb 2025)

Deep dive (finished overnight): Selling tokens finally makes money. Now what?

The core judgment. That Anthropic's token business has flipped from losing money to making it is fact, but read it in three layers: the revenue side has an official record (see today's lead 2); the margin side is an improvement that internal figures and analysts agree on, but nothing audited; and the "profitable quarter" is still guidance. The fat margin currently holds only at the priciest flagship agentic tier, propped up by three pillars: compute scarcity, the quality gap over open-weights models, and ROI scrutiny that buyers have not yet applied to their bills. All three are moving; none has fallen.

Why dig now. Our July 19 issue logged, in its first lead, four practitioners pointing, in the same week, to "the value in AI is moving away from the model itself." That same week, the opposite narrative — model vendors' margins flipping positive — was hardening too, and one of the two has to give. This dig asks exactly that: does the flip hold? A live contradiction still hangs over the market: one founder has said that over the past half year every model vendor has effectively raised prices, while media reporting says OpenAI is internally weighing deep price cuts. We keep both on the books and rule on neither yet.

Margins come in three ledgers — do not mix them. The Information, a subscription tech outlet known for reporting from internal documents, covers full-year 2025 on a full-cost basis: the margin projection was cut from 50% to about 40%, partly because cloud inference costs ran higher than expected (The Information). SemiAnalysis counts inference serving only: from 38% to around 70% (SemiAnalysis, May 2026). CNBC cites an investor gauge: compute cost per $1 of revenue down from 0.71 to 0.56 (CNBC, May 2026). Mind the sequence: the cut came first and covers 2025; the improvement came later and covers 2026. Pitting the 40% against the 70% is a basis mismatch. None of the three ledgers is audited; the IPO filings are the court of final appeal.

Infra capture era (2023-25): value eaten
  by chips/power/memory; token sales
  lose money (2024 inference margin -94%)
 └ Per-token economics flip (2026H1)
    agentic demand × cost crash; margin
    positive, profit qtr near, contested  ?
 ├ vs Path A: pricing power endures
    (SemiAnalysis): scarcity + quality gap
 ├ vs Path B: commoditization by default
    (Evans/Narayanan): undifferentiated
    tokens + zero switching → cost floor
 └ vs Path C: demand is watered down
    (Gurley/Chamath): subsidy + unsettled
    ROI; a closed capital window reveals it

What would prove this wrong. An actual price cut at the flagship tier flips the "durable pricing power" line. The nearest checkpoint is September 1: Anthropic's mid-tier model Sonnet 5 launched at a promotional $2/$10 per million tokens (input/output), with official notice that it reverts to the standard $3/$15 after August 31. If it does not revert, the promotion becomes a de facto price cut (launch-day pricing rundown: Simon Willison, Jun 30). The lower rungs of the price ladder are already at war: Google cut its consumer AI subscription from $7.99 to $4.99, which tech media called the starting gun of the AI subscription price wars (TechCrunch, Jun 9). One directly usable line for people building products on these APIs: if Sonnet 5 reverts to standard pricing after September 1, any unit economics modeled on the promo price need a rerun — and whatever traffic can route to mid- and low-tier models, route it down now.

Verdict dates. We have set ourselves a 12-month observation window. The nearest square on it: September 1, checking whether Sonnet 5 reverts. The heaviest: the audited margin in Anthropic's IPO filings — and which of the three ledgers it lands on.

Expert calls (retrospective)

This week's marquee statements (both Brockman's) are already in today's leads 1 and 4; this section carries one retrospective instead, in service of our ongoing work of nailing old predictions to the ledger.

  • Dylan Patel (founder of SemiAnalysis, front-line chip and data-center analyst), original October 2024. Multi-datacenter training splits a single training run across data centers hundreds of kilometers apart, to get around the problem of one campus not having enough power. His call at the time: this had been solved — "At least, Microsoft and OpenAI have figured it out." The evidence he offered was not technical documentation but capital footprints: Microsoft had "signed deals north of $10 billion with fiber companies," and dig permits had been filed along the routes. In the same session he predicted OpenAI would raise $50–100B and that single-site 1GW-class data centers would arrive in 2026 (Dwarkesh Podcast, Oct 2024). Seen from mid-2026, the direction has broadly borne out. The value to this brief: his claims and his evidence both carry timestamps, so 2026 can audit them line by line — pinning practitioners' original predictions to the ledger is exactly the raw material of our "who calls it right" accounting.

Research notes (academic and technical)

The 28 academic papers that arrived overnight are still being read; rather than force a pick today, this section carries two retrospectives, both directly relevant to today's leads, original dates labeled.

  • [Retrospective] (June 2024 interview) Chollet's shift-cipher probe, set beside today's perfect-score claim. François Chollet, creator of the Keras framework, once gave the claim "large models memorize rather than reason" a testable probe: models can solve a Caesar cipher — the beginner's encryption that shifts every letter forward by n places — but mostly only when n is 3 or 5, the values most common in online tutorials; move it to 9 and they fail. If a model were genuinely synthesizing the solution on the spot, the shift size would not affect difficulty at all (Dwarkesh Podcast, Jun 2024). The tension when set beside today's lead 1: the 2024 diagnosis was "no on-the-fly synthesis of algorithms," and 2026 has produced a perfect-score formal-proof claim. In fairness, the tasks differ in kind — the probe tests instant synthesis inside the model's parameters, while the IMO system is a multi-agent assembly wrapped around a Lean toolchain — and June 2024 predates the reasoning-model era. The probe's current status has no public retest; we log it as pending.
  • [Retrospective] (November 2024 interview) Gwern: agents are unreliable because agency was never a training target. Gwern, the anonymous independent researcher known for his early systematic case for the scaling hypothesis in 2020, diagnosed it this way: web corpora contain only descriptions of agent behavior, never the step-by-step decision traces — in his words, "All the agency there is is an accidental byproduct of somebody training on data" (Dwarkesh Podcast × Gwern, Nov 2024). That still has explanatory power for the 2026 reality of agents that demo brilliantly and wobble in production. It also echoes today's lead 4: the product form is evolving; the root problem of reliability is a separate thread. Time boundary: industry investment in reinforcement learning for agents has risen sharply across 2025–26; this is a diagnosis from a 2024 vantage point.

From the archive

"Three mutually corroborating pieces of evidence" — unpacked and verified, one and a half survive (our deep dive, July 4, 2026). On July 1, Meta was reported to be looking, SpaceX-style, at renting out surplus compute for cash; semiconductor stocks slid hard that day (TechCrunch, Jul 1, 2026). The popular reading strung three pieces of evidence into "compute oversupply": Oracle's thin GPU-cloud margins, three weeks of falling B200 chip rents, and Meta unloading compute. Our deep dive tested each for independence. The result: the Oracle item describes one company's structural disadvantage — leased facilities and concentrated customers — not industry-wide oversupply; the rent decline is a three-week point reading that traces back to the same day's same report as the Meta story; and the only genuinely independent supply-side item is SoftBank's announced 10GW-class compute service. Nominally three pieces, effectively about one and a half. The counter-evidence stays on the same page: SemiAnalysis argued against consensus in the same period that Meta's compute purchases are accelerating and that its rental tier is priced at a 3–4× premium over peers — genuine oversupply forces discounts, not premiums (SemiAnalysis, Jul 2, 2026). The takeaway: when you see three pieces of evidence corroborating one another, first ask whether they are three shadows cast by a single original point. (The actual figures for Oracle's GPU-cloud margin and the B200 three-week rent decline come from subscription sources and are omitted from the public edition.)

Sources and coverage

The past 24 hours. Overnight we scanned 44 new pieces: 28 academic papers (topic scan; separately, 47 tracked authors had no new work, and a few queries timed out unfinished — no guarantees beyond that scope), 13 podcast archive transcripts (11 read overnight under a deliberately conservative standard, 2 filtered out for carrying no new signal — saving you the time), 2 blogs, and 1 newsletter (filtered out). We also collected 17 original posts from among the 305 accounts we track on X (@gdb 2, @elonmusk 12, @LisaSu 1, @sundarpichai 1, @steipete 1); no full-network scan was run. Subscription newsletters and archives (Stratechery, Colossus, 7 Substack archives) had nothing new; 27 earnings sources had no new filings; no new official macroeconomic releases, so this issue carries no macro section. From all that we distilled 15 new facts overnight, 6 of them from re-reads of 2024 Dwarkesh interviews — all listed as retrospectives, none as news. Before dawn we finished one deep dive (see the deep dive section above); during the day we ran targeted same-day verification on 3 items (today's leads 2 and 3, and the revenue anchor in the deep dive). No sources were added today and no one-off backfills were run. Coverage statement: this issue can vouch only for signals within the scan scope above.

The sources we track. This brief's judgments rest on 529 named voices we currently track: 305 on X (Elon Musk, Andrej Karpathy, Greg Brockman, Nathan Lambert, and others), 90 podcast voices (Satya Nadella, Dario Amodei, Jensen Huang, and others), 51 news outlets, 48 personal bloggers (Simon Willison, Chris Olah, and others), 48 paper authors (Noam Shazeer, Percy Liang, Tri Dao, and others), 46 newsletters (Dylan Patel, Ben Thompson, Ethan Mollick, and others), 26 earnings and earnings-call sources, and 23 keynote series.

This brief is not a news digest: from each day's AI firehose we capture the insights that actually matter and the practitioner judgments worth tracking over time, and we show how every one of them was verified. The point is always which judgment got harder and who calls it right — not what happened today.

— SecondSource · generated by our research system · 19 sources · Reply to this email; it is the best feedback you can give us.

Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.


EN English edition|繁 中文版 Traditional Chinese →

Don't miss what's next. Subscribe to SecondSource:
← Newer SecondSource Deep Dive · 2026/7/20|Selling Tokens Finally Makes Money — Now Comes the Fight to Keep It Older → SecondSource Daily Brief · July 19, 2026 | Four Insiders Who've Never Met Just Pointed to the Same Conclusion: The Value in AI Is Moving Away From the Model Itself
buttondown.com
Powered by Buttondown, the easiest way to start and grow your newsletter.