SecondSource logo

SecondSource

Archives
Log in
Subscribe
August 31, 2026

SecondSource Morning Brief · August 31, 2026 | Money follows narrow tasks, not capability

1. Open-weight models — the kind whose parameters are published, so you can download and run them yourself — took 29% of the traffic at one

🌐 Read this issue on the web

At a glance

  1. Open-weight models — the kind whose parameters are published, so you can download and run them yourself — took 29% of the traffic at one model-routing platform and under 4% of the money. What blocks the money is not model capability.
  2. OpenAI's own chip produces 1.5 to 1.9 times more per unit of power than Nvidia's, and ties with it on total cost per token.
  3. One teardown, retold second-hand, lost the sentence about that tie — and the conclusion flipped to the opposite side.

This issue draws on our August 31 research round; the events run from June 2026 to August 30, 2026. Last night's sweep brought in 79 pieces, and 17 clickable outside receipts made it into this issue. This is the email edition; the full edition of this issue is the archive of record.

Today's main line

1. [Evidence update] (index published in July, covering June; technical report in August) The "open-weight only picks up the cheap work" explanation stops holding: on the one task frontier models cannot win, the money still has not moved

Vercel is a front-end deployment platform, and its AI Gateway is the middle layer enterprises use to route requests across several model providers — which leaves it holding real usage and spending data across vendors. Its official index, published in July on data from June 2026, carries two numbers pointing opposite ways: open-weight models took 29% of the platform's token volume and under 4% of the spending, at an average price per token roughly one tenth of the platform-wide average. The index reports those three separately; none is derived from the others, "under 4%" gives only a ceiling, and dividing it by 29% does not land exactly on one tenth. The same index gives the trend: in April open-weight was about a ninth of all usage, and by June it was close to a third (Vercel AI Gateway Production Index, July 2026).

Verification: the 29% is not new to us. We ran the same figure in our July 28 issue, where it came from the newsletter Exponential View computing it off the platform's public leaderboard. What is new today is the money column and the price column, plus the volume figure moving from a second-hand calculation to the company's own periodic index — a different grade of checkability. ⚠️ This is one platform's sample. Vercel's customers lean towards front-end and application development, so it cannot be stretched into the enterprise API market as a whole. The company sells the routing layer rather than models, which leaves it without an obvious stake in open versus closed — but the data is still self-reported from its own operations. ⚠️ One more figure, which we do not accept: someone cited the platform chief executive's live chart claiming open-weight set a record 62% share on August 22. That appears in no published index, it is a single day's reading, and the day was a Saturday, when weekend traffic skews experimental and does not compare with a monthly measure (the post citing that chart, 08-22).

The split itself is not a new observation. What changed today is the cause behind it. On July 22 the anonymous industry commentator @teortaxesTex drew the same distinction on X. Our roster has tracked that account for a long time, and he holds a consistent pro-China-model position. His distinction: to track competitive pressure from open weights, watch the spending displaced, not the share of tokens. His cause was the long tail — open weights pick up price-sensitive work that was never going to pay much, so volume climbs and money does not follow. That post carried no figures at all (the original post, 07-22).

The second piece of material we read today, this one with numbers in it, puts a hole in the long-tail explanation. Bridgewater, one of the world's largest hedge funds, published a technical report jointly with the AI lab Thinking Machines Lab. That lab's product, Tinker, offers enterprise fine-tuning — taking an existing model and training it a further round on your own data so it specialises in one job. The method: start from Alibaba's open-weight Qwen3-235B and fine-tune it on an expert-labelled task, deciding what counts as valuable financial information. Self-reported average accuracy came in at 84.7%, against 78.2% for the best of the frontier models they evaluated — the strongest general-purpose model from each lab. That works out to 29.8% fewer errors, at 13.8 times lower inference cost per task. The same report records that frontier models scored around 50% on that task without optimisation, rose to the mid-70s once experts wrote the prompts, and topped out below 80% (Thinking Machines × Bridgewater, August 2026). That task is not the long tail. Frontier models there are not more expensive; they are worse.

⚠️ The evaluation task, the expert labels and the comparison prompts given to frontier models were all designed by the two parties with a stake in the answer, and nobody has reproduced any of it. What that report raises is falsifiability, not credibility: concrete numbers give other people something to reproduce or overturn, and the finding itself gained no trustworthiness from being published. The 13.8 times is inference cost and leaves out the one-time expert labelling and fine-tuning, so it is not an end-to-end total. ⚠️ We learned of this fine-tuning case on June 30 through a third-party account, with no figures and no independent baseline, so we wrote you nothing about it. What today adds is the numbers.

Judgment update: if the money has not moved at scale even on work frontier models cannot win, then what blocks it is neither capability nor price sensitivity. What survives is narrowable, plus labellable: a workflow defined as a narrow task with clean boundaries and fixed inputs and outputs, in a customer that has experts who can mark what the right answer looks like. Both have to hold before volume and money move together; without both, open weights pick up only the cheap half. That moves the variable to watch off the calendar and onto structure — not "when does open source catch the frontier," but "what share of a customer's workload can be cut narrow." ⚠️ This is a working read, not a settled finding: each leg rests on a single source, and one of those sources is a party to the outcome reporting on itself. The evidence does not yet support a claim either way.

What would prove this wrong: three things, stated in full. An independent third party reproduces the Bridgewater task with a strong enough comparison and fails. Or the open-weight share of spending at the routing platform jumps from under 4% to above 15% within two months, with no change in customer task structure. Or an open-weight model arrives that matches the frontier on general capability. The window and the threshold in that middle condition are ours, not any participant's — two months, because the index updates roughly monthly, which leaves two fresh readings to compare. The verdict date is October 31, 2026, by which point the August and September editions should both be out and the spending share can be read directly for whether it has left the 4% range.

Investor note: the compression has already happened once, on one narrow task, while total spending share sits under 4%. Put those two readings side by side and the story that "the day open weights catch the frontier is the day model-layer margins compress" gets weaker: it puts the risk on timing, when the gap is opening somewhere else. The thing to measure instead is the share of a customer's workload that can be narrowed.

Three roles read three different actions out of this. Frontier labs need a new subject for the moat audit. The real competitor is not some open model that catches up; it is the data team inside the customer that can cut work narrow and afford expert labelling. Money spent defending the premium belongs in generality, service guarantees and outer orchestration, not in another point on a general benchmark — that fight has already been lost once on a narrow task. Cloud operators who size capacity off the revenue curve will run systematically late: 29% of the volume is worth 4% of the money, so pressure on low-priced capacity arrives first and the revenue signal may never arrive at all. Capacity and revenue belong on two separate sheets. Enterprise CIOs should replace the first question, "which model is strongest," with "can this workflow be defined as a narrow task, and do I have anyone who can label it." The threshold for building your own sits lower than most people assume, and what stalls you is the labelling staff, not the model. All three share one action: when you see a figure like "open weights are X% of the market," ask first whether that denominator is volume or money. Today those two numbers, from one platform in one month, differ by more than sevenfold.

2. [Evidence update] (teardown published August 27) OpenAI's own chip wins by 1.9x per unit of power and ties on total cost per token — the load-bearing half is the second one

Our August 26 issue noted that OpenAI's in-house inference chip Jalapeño had turned in its first results. What is new today is the measured figures — and the sentence that follows immediately after them in the same teardown.

Jalapeño is the first AI inference chip OpenAI designed itself, with Broadcom (inference being the half of the work a model does once it is live and answering questions, as opposed to training). SemiAnalysis, an independent semiconductor analysis firm founded by Dylan Patel, ran its own published InferenceX test and says it verified some of the runs on site at OpenAI's lab. The result: the 700-watt Jalapeño, against Nvidia's GB200 and GB300 in the 1,200 to 1,400-watt class — products from the Blackwell generation — produces 1.5 to 1.9 times more per unit of power and answers 1.7 to 3.6 times faster. In the firm's own words, "Jalapeño beats Blackwell on perf/W across almost all scenarios without being tuned for any specific point." The same teardown then judges the two roughly level on total cost of ownership per token — the money each token actually costs once chip price, yield, rack density, utilisation, electricity and depreciation are all counted in (SemiAnalysis, August 2026).

Two other figures are worth keeping. Design ran "from initial team hiring to manufacturing tape-out in ~16 months" — started mid-2024, finalised for manufacture in November 2025. And OpenAI reports that AI assistance in the design work cut the area of two computing blocks on the chip by 8% and 10%, and that its own code-generating model, Codex, wrote working and efficient low-level compute routines "without any of OpenAI's kernel engineering team intervening."

Verification: this rests on the SemiAnalysis teardown alone. OpenAI's own page returned 403 to us this round and we hold none of its text, so it does not count as a second independent source, and we treat "Jalapeño actually beats Nvidia" as a single-source claim. OpenAI chose the models tested and the machine configurations, and the on-site verification covered some runs rather than an independent rerun of everything. The 8% and 10% area reductions are OpenAI's own report; the teardown itself writes "they claim," and we do not upgrade that on its behalf. We have not checked the granularity of the wattage comparison against the primary page. ⚠️ Part of the 700-watt-against-1,400-watt efficiency advantage comes simply from the part being smaller, which is not the same as an architectural advantage at equal scale. ⚠️ The teardown separately compares Jalapeño's output per megawatt against Nvidia's next platform, Vera Rubin, with each side running a different mode — not one ruler.

Judgment update: 1.5 to 1.9 times reads like the in-house chip camp is about to eat into Nvidia's share, but the load-bearing sentence is the tie: share is decided by total cost of ownership, not by efficiency per watt. We have a live judgment running — that through the middle of 2027 Nvidia keeps more than seventy percent of the supply value in AI accelerators, the specialised chips that run AI computation, because the structure is sticky. Today's reading points towards strengthening that, not shaking it. What would loosen it, stated plainly: not a higher efficiency figure, but the first clear gap in total cost of ownership per token. One thing you can do today: write procurement and negotiation terms against total cost of ownership per token, not against an efficiency benchmark score.

Investor note: the prevailing story assumes that an in-house chip running on less power means Nvidia's share is about to loosen. This evidence separates those two. The power saving is real; the difference on the bill is zero. That makes the "in-house silicon erodes share" bet weaker, and the signal to wait for is not the next record on efficiency but the first real gap in total cost of ownership.

3. [This week] (newsletter published August 30) One teardown, two retellings, and what went missing is the sentence that cuts against the conclusion

The SemiAnalysis teardown we read today was also read the same week by the industry newsletter Exponential View, founded by Azeem Azhar, which linked to the original. It quoted the 1.5 to 1.9 times. It did not quote the tie on total cost of ownership. Word for word: the newsletter writes that Jalapeño "beats comparable Nvidia silicon by 1.5–1.9x on tokens per megawatt," and the string "total cost of ownership" appears nowhere in the piece (Exponential View #599, 08-30). The two conclusions then run in opposite directions. The newsletter's: "This differentiated hardware demand will expand the market even if it potentially reduces Nvidia's relative dominance." Ours: total cost of ownership is level, so this round is not evidence of share loosening.

Verification: ⚠️ this is not a case of getting something wrong. The newsletter linked to the primary document throughout and readers can click through; rounding 29.8% to "about 30%" and rendering 13.8 times as "a fourteenth of the cost" both sit within reasonable editorial licence. Exactly one thing here is checkable: put the two documents side by side, and the sentence that dropped out is the one in that teardown that cuts against the conclusion and carries the weight for a purchasing decision. We keep both conclusions and rule on neither: if total cost of ownership does open up a gap in the next generation, the newsletter's conclusion lands first.

Something the same author disclosed unprompted belongs alongside this. In that same piece he states that he is an investor in Fractile, a low-latency inference chip startup, and in Trainloop, a fine-tuning services company — two lines that map onto the direction of items 2 and 1 above. Voluntary disclosure counts in his favour, not against him. Our response is not to distrust him; it is that all three of today's facts go around him and read the primary pages directly.

Judgment update: for anyone tracking this industry through second-hand summaries, this item hands over an action you can take immediately: when a striking ratio catches your eye, go back to the original and find the "but" — it is usually there, and it is usually the sentence that decides whether you sign. Read this together with item 2: there, what got misread was the 1.9 times, which measures output per unit of power rather than cost per token; here, what got dropped is the sentence in the same document that bounds it. Two layers of one problem.

Investor note: several outlets retelling one document are not several sources. A good deal of today's optimism about in-house silicon traces back to the same ratio in the same teardown, one or two retellings removed. This evidence changes nothing about what that teardown says; it weakens the impression that several independent sources are saying it.

Also happened — not verified by us yet

  1. The philosopher Toby Ord published a paper arguing that the gate on recursive self-improvement — AI helping design and train the next version of itself, each generation stronger — is the length of a single improvement loop, and stacks three physical ceilings on top: the speed of light, how much information fits in a given volume, and the energy floor for irreversible computation. We read only a second-hand summary and have not read the paper itself (arXiv 2608.14426).
  2. Thomson Reuters, the legal and financial information supplier behind Westlaw and the Reuters news agency, is reported to have built its first in-house model on top of Alibaba's Qwen to bring costs down; we read only a relayed account and not the original reporting (Yahoo Finance). If it holds up, it would be a second enterprise case for the working read in item 1 above.
  3. The startup int21.ai says that swapping in a different outer orchestration system — the layer that breaks a problem up, retries and verifies — took GPT-5.6 from 13.3% to 100% on the public set of the abstract reasoning test ARC-AGI-3. ⚠️ The questions and the answers in a public set are both published, which is the same as seeing the exam paper beforehand, so that 100% carries almost no information; we have not read the vendor's page (int21.ai).
  4. A Reuters investigation says Meta considered cutting some teams by as much as 60% to become an "AI-native" organisation, and that the plan collapsed. ⚠️ This is a plan that was considered, not a thing that was done; we have not read the original piece (Reuters, 08-26).

Chips & semiconductors

[This week] (published August 30) Sixteen times the capacity per stack in flash memory — and dropping it straight in starves the GPU. Oxford proposes the compromise.

A research team at the University of Oxford published a paper: at comparable bandwidth, high-bandwidth flash memory holds 16 times the capacity per stack of the high-bandwidth memory in use today. That looks like the answer to inference workloads that cannot fit a large model in memory. Swapping it straight in drags performance down badly, because the tail latency of flash leaves the GPU's scheduler waiting. Their approach mixes the two kinds of memory and uses prediction to keep the slow half off the critical path (Semiconductor Engineering, 08-30; the paper, IEEE Computer Architecture Letters, August 2026). Why it is worth keeping: items 1 and 2 of the main line are both about the money in inference, and the physical floor under that cost is whether the model fits in memory. Capacity at a sixteenth of the cost is a big number, and what you pay for it is latency. Nobody has finished that calculation, and until someone does, every argument that memory is about to get cheaper is missing half its case. The signal that it has been finished is a published latency measurement of the mixed architecture under a real inference load, rather than one more capacity multiple.

[This week] (published August 30) On-chip high-speed memory shrinks 25% in area and writes 42% faster.

Georgia Tech and the electronic design software vendor Synopsys published a paper on the 6-transistor SRAM built on a 2-nanometre process. SRAM is the small, fast memory sitting next to the compute units, and one of the largest consumers of area on a chip. The approach stacks it vertically: the access transistors move up into the metal interconnect layers, paired with silicon nanosheets and buried power rails. The cell comes out 25% smaller in area than the high-performance silicon baseline, and sub-array write latency 42.2% lower (Semiconductor Engineering, 08-30; the paper, arXiv 2608.22741, August 2026). Why it is worth keeping: since SRAM is one of the biggest area consumers on a chip, 25% less of it pushes in the direction of that sentence in item 2 — total cost of ownership is what carries the weight. This one leaves nothing you can act on in purchasing or product decisions today, and we log it as technical tracking. ⚠️ It is a design proposal at the paper stage, not a production yield, with an entire process qualification in between.

Named commentary

[This week] (posted August 30) Simon Willison: OpenAI explains its new product in terms of what it is for rather than what it does — so he went and asked the product itself.

Simon Willison co-created the open-source web framework Django and has spent years publicly testing AI tools and writing down what he finds, item by item. He spent a while working through ChatGPT Work, which OpenAI released on July 9, and his conclusion is blunt. The official documentation defines the product by when you should reach for it: you want a deck, an analysis, a deliverable file. He judges that close to information-free, because ordinary ChatGPT already does those things. What he assembled instead is a list of capabilities: a code execution environment with network access, a browser with no interface, a file system shared across conversations, the ability to dispatch sub-tasks, and schedulable automation. After publishing, he did something more direct still: he opened a fresh session and asked the system to list every tool it holds, sort them into categories and explain each one, as a website. It listed 223 registered tools, 6 of them his own connections (Simon Willison, 08-30). The problem he names sits one layer down — OpenAI still does not publish its system prompts or tool descriptions. In his words: "If the ChatGPT Work documentation included the exact system prompt and tool descriptions used by the agent I wouldn't have needed to write this post." Why it is worth keeping: deciding whether an agent product can enter your own workflows takes a tool list and a permissions boundary, not marketing copy about use cases. Today that list exists because a user asked for it, not because the vendor supplied it. For procurement and security review that is an action already sitting on the table: you can ask it too.

Model watch

[This week] (published August 30) Purdue's GPU simulator matches real silicon at 99%: how a model behaves once it runs can now be measured without buying the machines.

What this simulator measures is how a model behaves on hardware, not how the model itself scores. A research team at Purdue University published a paper presenting a cycle-level simulation framework that covers three Nvidia generations: Ampere, Hopper and Blackwell. Cycle-level means computing the behaviour of every clock cycle on the chip rather than estimating an average. They report validation against physical silicon at a Pearson correlation of 99% on the H100 — the Pearson correlation measures how closely two sets of numbers track each other, where 1 is a perfect match (Semiconductor Engineering, 08-30; the paper, arXiv 2608.22602, August 2026). Why it is worth keeping: a public simulator that lines up with real silicon changes how much the next generation's benchmark readings can be trusted, and purchasing judgments ride on exactly those readings. It also moves "how should the next GPU be designed" out of a handful of companies' internal tooling and into reach of academics: studying asynchronous, distributed GPU architectures used to mean either buying the machines or settling for rough estimates. ⚠️ The 99% is a validation reading against the H100 and says nothing about accuracy on other generations or other workloads — the paper names the H100 as what it validated against.

Product moves

No product news this issue. No AI company's official blog carries a publication date inside the past 48 hours; the most recent one we hold went up on July 22, which is not today's news. We will not pad the column with a feature explainer for an existing product.

From the archive

No archive pick this issue. We have used up the older deep material worth reusing from our own back catalogue — the last pick ran on July 30, and this is the twelfth consecutive issue with the column empty. We would rather leave it blank than replay an item we have already run.

Sources & accounting

The past 24 hours. Last night's sweep put 79 pieces on our reading list: 64 academic papers, 8 podcast transcripts, 6 company and personal blog posts, and 1 industry newsletter. The names: the one newsletter is Exponential View #599; all 6 blog posts are paper summaries from Semiconductor Engineering, which is where both chips items and the model watch item come from. Those 8 podcast transcripts are backfill of January and February episodes; last night's genuinely new pull was 5 (Dwarkesh Podcast 1, BG2 4). Eight pieces were read to the end and judged today: that newsletter, the three primary documents it pointed us to — the SemiAnalysis teardown, Vercel's official index, and the Thinking Machines and Bridgewater technical report — three paper summaries, and Simon Willison's August 30 post. Nobody has read a single one of the 64 academic papers today. Material missing from this issue is mostly missing because nobody has read it, not because we filtered it out.

What you are not getting today. Three things. One, our largest tracked channel, X, would not accept a login today and we pulled no new social posts at all, so no main-line item rests on a social post. Two, the paid Stratechery subscription was equally unreadable, for the third day running. Three, 11 of our 12 podcast channels were shut out by the platform and aborted, and transcripts actually arrived from 2 shows. Those two figures add to more than the number of channels; we write them as the fetch record has them rather than forcing one accounting on both. What shut us out is a platform-side block, not a missing credential of ours. Separately, the OpenAI page that item 2 needed returned 403 to us, so every figure for that chip today came through a third-party teardown and we hold no text from the primary page.

Older material added by hand. None of last night's 8 podcast transcripts is cited in this issue. Nobody added a source by hand today.

Source concentration. Two denominators, two answers, and we give both. Measured by who led us to today's material, concentration is 100%: every new piece today came in through the same newsletter, Exponential View #599, far above our own one-third warning line. Measured by whose words we used as evidence, it splits three ways: the three facts in the main line rest on the SemiAnalysis teardown, Vercel's official index, and the Thinking Machines and Bridgewater technical report, and that newsletter carries the weight for none of them. The reason is that its author discloses positions on two of those lines, and we could go around him and read the primary pages ourselves, so we did.

The sources we track. After de-duplication the roster covers 529 sources (our count today). Six categories are countable today: podcasts 90, outlets and press rooms 51, paper authors 48, blogs 48, newsletters 46, earnings calls 26. X gives two different answers today: 171 sources whose main channel is X, and 388 that have an X account on the roster. Change the denominator and the answer more than doubles, so today we do not pick one as the number (our own version of that "ask what the denominator is" line in item 1). One person can appear on several channels, so the categories add to more than 529. Representative names: on X, Mark Zuckerberg and Arvind Narayanan; in newsletters, Zvi Mowshowitz and Nathan Lambert; on papers, Ion Stoica and Yann LeCun; on podcasts, Satya Nadella and Dario Amodei. The third ruler is the figure signed at the foot of the page: this issue uses 17 sources in the body, counting the receipts this piece actually cites that come from outside us, since our own back issues and platform home pages do not count. That population differs from the 529 on the roster, and the three ledgers do not add up.

I finished today's issue / I didn't finish

This is not a news digest: we hunt each day's AI firehose for the insights that actually matter and the practitioner judgments worth tracking over time, and we show how every item was verified — the point is always "which judgment got harder, and who's been right," never "what happened today."

— SecondSource · generated by our research system · 17 sources · Got a view? Reply and tell us

Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.


SecondSource publishes industry analysis, not investment advice. We do not evaluate, rate, or recommend any specific security, and nothing here should be treated as financial guidance — verify independently and use your own judgment.

EN English edition|繁 中文版 Traditional Chinese →

Don't miss what's next. Subscribe to SecondSource:
← Newer [Delayed: system issue] SecondSource Morning Brief · September 1, 2026 | That 14x measures routing, not the market Older → [Delayed: system issue] SecondSource Morning Brief · August 30, 2026 | Marvell's filing: warrants cut its own reven…
buttondown.com
Powered by Buttondown, the easiest way to start and grow your newsletter.