Token math
driver a named mechanism, not a conclusion
KV cache per token times context times concurrency sets whether a fleet is compute-bound or memory-bound. Compute binds by ~30x today; memory begins to bind at ~1M average context.
The path
| node | effect |
|---|---|
| KV cache bytes per token, by attention architecture | +0.35 |
| Memory (KV/state) per token | +0.35 |
| Long-context inference (KV-cache growth) | +0.35 |
| Memory capacity (per accelerator and per node) | +0.35 |
coverage
4 of 4 declared anchors resolve in the model. A driver is a named mechanism rather than a graded severity or a priced claim, so its effects are deliberately smaller than either.
Arguments about this mechanism
These name its subject — they are claims about this mechanism, and they are what the eligibility gates grade.
Arguments that run through it
These name a node on the path rather than the subject, so they pass through this mechanism without being about it. Shown separately and never graded as claims about it — hubs above the graph’s own 90th-percentile degree are excluded, or every argument touching NVIDIA would attach to every mechanism NVIDIA touches.