A model's architecture is public the day it ships. Its procurement footprint is not public for quarters. This is the arithmetic between the two.
Colour is the grade of the input at that step. The argument gets weaker as it descends, and that is the honest shape of it — the physics is computed, the workload is assumed, the prices are estimates.
| market | physical | implied spend |
|---|---|---|
| HBM deployed | 236.9 PB | $2.8bn |
| Accelerators | 1.3m | $102.5bn |
| Facility capex | 976 MW | $9.8bn |
Coherence note. HBM content works out to $2,232 per accelerator, 2.8% of the unit price, against teardown estimates of 10–20% of BOM. The likely cause is a scope mismatch rather than a bad price — the accelerator figure is a system spend including networking and integration, while $/GB buys the memory alone. The two are different units of account and must not be summed, and the HBM figure should be read as a floor.
The HBM market split by the graph's own supply weights. A row marked assumed rests on the 0.5 default rather than a disclosed share — a revenue number resting on that is a guess wearing a company name.
| supplier | share | implied revenue | weight |
|---|---|---|---|
| sk-hynix | 55.0% | $1.6bn | from model |
| samsung | 25.0% | $0.7bn | from model |
| micron | 20.0% | $0.6bn | from model |
| input | value | grade | source |
|---|---|---|---|
| intensity — Computed from a published config or a measured series. The defensible layer. | |||
| kv_kb_per_token | 252 | medium | coefficients.py (Llama GQA, FP8) |
| wh_per_token | 0.000219 | medium | entity:energy-per-token |
| gflops_per_token | 810 | medium high | coefficients.py (2 x active params) |
| tokens_per_day_tn | 107 | low medium | entity:disclosed-token-volume (Google only — a FLOOR) |
| avg_context_tokens | 32,768 | asserted | VOLUME ASSUMPTION — not measured |
| concurrent_requests_k | 500 | asserted | VOLUME ASSUMPTION — not measured |
| workload — How the fleet is actually used. No provider publishes these — they are assumptions, and they are where the chain becomes an estimate. | |||
| mfu | 0.35 | asserted | Typical published serving MFU. NOT measured on this fleet |
| weight_fraction | 0.45 | asserted | Depends on model size, replication and parallelism strategy — a modell |
| price — Unit prices. Desk estimates from reported figures, not contract prints. | |||
| hbm_usd_per_gb | 12 | low | Desk estimate from reported HBM stack pricing — NOT a contract print |
| dc_usd_per_mw | 10.0m | asserted | Industry order-of-magnitude, carried from the megawatt thesis where it |
| accel_usd_each | 80,500 | low medium | Goldman build-out model Rubin baseline, cited in supplied analysis |
avg_context_tokens. A chain is no stronger than its weakest term, and 5 of 11 inputs are asserted: mfu, weight_fraction, dc_usd_per_mw, avg_context_tokens, concurrent_requests_k. Those are not the weak part by accident — context length, concurrency, MFU and weight fraction are precisely the quantities no provider publishes. The physics is the defensible part. The dollar conversion and the workload shape are where this becomes an estimate, and the output must not be quoted at the precision the physics has.