cost-per-token-is-set-outside-the-chip · conviction medium · status open · horizon 2027-2028 · as of 2026-08-04
Series available as data/cost-per-token-is-set-outside-the-chip.csv
Green/blue = model marks it as CAPTURING the rent (unbound and supplies the scarce good); faded = PAYING it (bound severe or moderate). Operating margin, live.
42% if the 3 gates are independent, 70% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: Elite Material Co., Ltd. (EMC) at 0.70 — M9-grade CCL with quartz-glass reinforcement is the substrate material for the Kyber midplane at ~1 square metre and ~78 layers, where yield compounds
31% if the 3 gates are independent, 60% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: Coherent Corp. at 0.60 — Co-packaged optics needs perimeter it does not control, so the optics opportunity is gated by a power-architecture decision made by the accelerator ve
48% if the 3 gates are independent, 75% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: Long-context inference (KV-cache growth) at 0.75 — Agentic workloads run 142k input tokens per turn over a median 65 turns, which is ~71x a chat turn on context alone.
39% if the 3 gates are independent, 70% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: Energy per token served (site level) at 0.70 — 0.000199 Wh per token at frontier scope, composite grade medium, every input now sourced: GB200 NVL72 rack power over GPU count, hyperscale PUE from o
29% if the 3 gates are independent, 60% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: NVIDIA GB200 (per-GPU in NVL72) at 0.60 — The accelerator is one term. Rack power per GPU is ~1,667W all-in against ~1,000W for the die alone, so even inside the hardware term the chip is 60%
Cost per token falls with each accelerator generation. Buy the compute leader and the deflation accrues to whoever serves the most tokens.
Accelerator performance is one of at least six terms. Package perimeter, memory architecture, power delivery, grid access and serving software each set the coefficient independently, and the binding term has changed twice in eighteen months. A view on token cost that only models the chip is modelling one variable in six.
Consensus watches FLOPS per dollar. This watches the product of six coefficients, each measurable, and asks which one binds. On 2026-08-04 the answer was serving software, not silicon.
Elite Material Co., Ltd. (EMC)0.70 moderateMaterials qualification gates the board before any chip ships.
CoWoS (Chip-on-Wafer-on-Substrate)0.85 strongHybrid Bonding0.70 moderateThe first term is physical assembly. Before an accelerator has a cost per token it has to exist as a package, and the package is gated by laminate qualification, interposer capacity and bonding capability — none of which the accelerator vendor controls end to end.
Package shoreline (perimeter and cross-section contention between power, memory and data)0.75 strongPower escape and optical I/O compete for the same millimetres.
Multilayer ceramic capacitor (high-capacitance, AI server)0.70 moderateThose capacitors occupy exactly the shoreline the SerDes need.
Coherent Corp.0.60 moderateThe second term is geometric and is usually invisible in cost models. Perimeter is fixed, power and data both need it, and the resolution — power down through the back, data up and out the front — is a packaging decision that determines how much bandwidth a given die can actually reach. Bandwidth per package is a cost-per-token input and it is set by millimetres, not by process node.
KV cache bytes per token, by attention architecture0.80 strongHBM40.80 strongLong-context inference (KV-cache growth)0.75 strongThe third term is architectural and it dominates the second. Compression bought 7.3x and the agentic workload took 71x, so cache pressure rose despite better attention. This is the term that has embarrassed the desk twice: it holds both the coefficient and the denominator and had never multiplied them.
Energy per token served (site level)0.70 moderateGrid interconnection queue position0.80 strongDisclosed token volume (operator-reported)0.70 moderateThe fourth term is where the chain leaves the building. Tokens become watt-hours become megawatts become a queue position, and the queue is the only term with a multi-year lead time. Efficiency relieves it in STEPS at hardware transitions rather than continuously, so between generations the pass-through is close to one for one.
Inference cost per token0.75 strongDecomposed: silicon perf/W 2-3x per generation, numerics ~4x cumulative, architecture and serving uncaptured and faster.
Inference serving stack (vLLM / TensorRT-LLM class)0.65 moderateTogether they contributed more cumulative reduction than process nodes.
NVIDIA GB200 (per-GPU in NVL72)0.60 moderateRack power per GPU is ~1,667W all-in against ~1,000W for the die alone, so even inside the hardware term the chip is 60% of the number.
THE INVESTABLE CLAIM. Consensus buys the compute leader because FLOPS per dollar improves. But the fastest deflation levers — architecture and serving — are published free and captured by nobody, which means a hyperscaler's cost curve improves substantially without buying anything. The terms that ARE capturable sit in materials, packaging, memory and power, which are priced as cyclicals rather than as coefficients on the token economy. THE FALSIFIER IS SHARP: if cost per token tracks accelerator generations more closely than it tracks architecture releases over the next four quarters, the chip is the term and this thesis is wrong. That is checkable against dated model releases and dated silicon launches.
Pre-filled skeleton: cost-per-token-is-set-outside-the-chip.md