← research bank

Cost Per Token Is Set Outside the Chip

cost-per-token-is-set-outside-the-chip · conviction medium · status open · horizon 2027-2028 · as of 2026-08-04

The market prices AI compute on accelerator performance. But the cost of a served token is set by a chain that runs from copper-clad laminate to the interconnection queue, and the accelerator is only one term in it. Every link is measurable, most are not priced, and the binding one moves between them.
Rests on filed figures, not on modelled shares. 15 premises (15 entity); no derived cell is involved, so the undisclosed supply weights that put a range on other pages in this bank cannot move this one.

Exhibits

Exhibit 1Relative performance, indexed to 100How the names in this thesis have traded against SOXX.
86051201000660.KS 5732383.TW 469SOXX 225TSM 176NVDA 12312mo, indexed to 100 at start · dashed = SOXX benchmark

Series available as data/cost-per-token-is-set-outside-the-chip.csv

Exhibit 2Who pays CoWoS advanced-packaging capacity, and who keeps the moneyCapturers average 47.1% operating margin against payers' 44.6% — the owners of the scarce thing capture the rent, as expected.
Taiwan Semiconductor Manufac56.1%Analog Devices, Inc.38.1%NVIDIA Corporation64.0%SK Hynix58.6%Broadcom Inc.44.2%Advanced Micro Devices11.8%

Green/blue = model marks it as CAPTURING the rent (unbound and supplies the scarce good); faded = PAYING it (bound severe or moderate). Operating margin, live.

Exhibit 3What the conviction is actually made ofEach premise and the number it composes to. A conjunction of plausible premises is far weaker than any of them.
Materials and packaging set a floor on how much compute fits in a package at allElite Material Co., Ltd. (EMC)70.0%CoWoS (Chip-on-Wafer-on-Substrate)85.0%Hybrid Bonding70.0%COMPOSED (and)41.6%

42% if the 3 gates are independent, 70% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: Elite Material Co., Ltd. (EMC) at 0.70 — M9-grade CCL with quartz-glass reinforcement is the substrate material for the Kyber midplane at ~1 square metre and ~78 layers, where yield compounds

The package perimeter forces a trade between power and bandwidth that neither side wins outrightPackage shoreline (perimeter and cross-sect…75.0%Multilayer ceramic capacitor (high-capacita…70.0%Coherent Corp.60.0%COMPOSED (and)31.5%

31% if the 3 gates are independent, 60% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: Coherent Corp. at 0.60 — Co-packaged optics needs perimeter it does not control, so the optics opportunity is gated by a power-architecture decision made by the accelerator ve

Memory architecture moves the coefficient more than a hardware generation doesKV cache bytes per token, by attention arch…80.0%HBM480.0%Long-context inference (KV-cache growth)75.0%COMPOSED (and)48.0%

48% if the 3 gates are independent, 75% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: Long-context inference (KV-cache growth) at 0.75 — Agentic workloads run 142k input tokens per turn over a median 65 turns, which is ~71x a chat turn on context alone.

Energy per token converts the whole chain into megawatts, and megawatts are queuedEnergy per token served (site level)70.0%Grid interconnection queue position80.0%Disclosed token volume (operator-reported)70.0%COMPOSED (and)39.2%

39% if the 3 gates are independent, 70% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: Energy per token served (site level) at 0.70 — 0.000199 Wh per token at frontier scope, composite grade medium, every input now sourced: GB200 NVL72 rack power over GPU count, hyperscale PUE from o

Therefore cost per token is a product of six coefficients, and the binding one movesInference cost per token75.0%Inference serving stack (vLLM / TensorRT-LL…65.0%NVIDIA GB200 (per-GPU in NVL72)60.0%COMPOSED (and)29.3%

29% if the 3 gates are independent, 60% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: NVIDIA GB200 (per-GPU in NVL72) at 0.60 — The accelerator is one term. Rack power per GPU is ~1,667W all-in against ~1,000W for the die alone, so even inside the hardware term the chip is 60%

The variant

Consensus

Cost per token falls with each accelerator generation. Buy the compute leader and the deflation accrues to whoever serves the most tokens.

Variant

Accelerator performance is one of at least six terms. Package perimeter, memory architecture, power delivery, grid access and serving software each set the coefficient independently, and the binding term has changed twice in eighteen months. A view on token cost that only models the chip is modelling one variable in six.

Differentiator

Consensus watches FLOPS per dollar. This watches the product of six coefficients, each measurable, and asks which one binds. On 2026-08-04 the answer was serving software, not silicon.

Open questions

Reasoning chain

Materials and packaging set a floor on how much compute fits in a package at all VALID
premises

The first term is physical assembly. Before an accelerator has a cost per token it has to exist as a package, and the package is gated by laminate qualification, interposer capacity and bonding capability — none of which the accelerator vendor controls end to end.

The package perimeter forces a trade between power and bandwidth that neither side wins outright VALID
premises

The second term is geometric and is usually invisible in cost models. Perimeter is fixed, power and data both need it, and the resolution — power down through the back, data up and out the front — is a packaging decision that determines how much bandwidth a given die can actually reach. Bandwidth per package is a cost-per-token input and it is set by millimetres, not by process node.

Memory architecture moves the coefficient more than a hardware generation does VALID
premises

The third term is architectural and it dominates the second. Compression bought 7.3x and the agentic workload took 71x, so cache pressure rose despite better attention. This is the term that has embarrassed the desk twice: it holds both the coefficient and the denominator and had never multiplied them.

Energy per token converts the whole chain into megawatts, and megawatts are queued VALID
premises

The fourth term is where the chain leaves the building. Tokens become watt-hours become megawatts become a queue position, and the queue is the only term with a multi-year lead time. Efficiency relieves it in STEPS at hardware transitions rather than continuously, so between generations the pass-through is close to one for one.

Therefore cost per token is a product of six coefficients, and the binding one moves VALID
premises

THE INVESTABLE CLAIM. Consensus buys the compute leader because FLOPS per dollar improves. But the fastest deflation levers — architecture and serving — are published free and captured by nobody, which means a hyperscaler's cost curve improves substantially without buying anything. The terms that ARE capturable sit in materials, packaging, memory and power, which are priced as cyclicals rather than as coefficients on the token economy. THE FALSIFIER IS SHARP: if cost per token tracks accelerator generations more closely than it tracks architecture releases over the next four quarters, the chip is the term and this thesis is wrong. That is checkable against dated model releases and dated silicon launches.

Sources

Write-up

Pre-filled skeleton: cost-per-token-is-set-outside-the-chip.md