memory-demand-durability · conviction low · status open · horizon 2026-2028 (adoption is multi-quarter) · as of 2026-08-03
Series available as data/memory-demand-durability.csv
52% if the 2 gates are independent, 65% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: Ai Capex Cycle — signal 2026-08-04 at 0.65 — Wells Fargo NVL72 BOM: memory (HBM+LPDDR) rises from 26.6% to 30.8% of rack cost between GB300 and Vera Rubin, growing +128% against compute's +116%.
14% if the 3 gates are independent, 30% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: KV cache bytes per token, by attention architecture at 0.30 — Computed from published configs: GQA at 252 KB/token against MLA at 34 KB/token, a 7.3x spread. Held at 0.80 because the arithmetic is trivially check
† 3 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.
One gating premise, so the conclusion is exactly as strong as it. The figure is an ordering device, not a calibrated probability — see how the numbers are made.
Weakest link: Hybrid Linear Attention (constant-state KV) at 0.85 — Hybrid-linear attention exists and is implemented, not a paper proposal. The doubt is adoption breadth, scored separately.
† 1 premise marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.
85% if the 2 gates are independent, 85% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.
Weakest link: Kimi K3 (Moonshot) at 0.85 — Kimi K3 ships and is documented. Existence is not the doubt.
† 11 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.
No gating premise. Every premise here is supporting evidence, so this conclusion states no necessary condition — it is asserted from cited data rather than derived from a claim that could fail. Read it as a summary, not as a falsifiable call.
Weakest link: Frontier hybrid-linear (constant-state) adoption — value single-lab at 0.60 — SINGLE-LAB is the load-bearing premise and the weakest. One lab shipping constant-state attention at frontier quality is proof of possibility, not of
Long-context/agentic AI drives structurally rising memory demand; the KV-cache-per-token requirement is a given, so HBM/DRAM shortage persists — validated by SK Hynix Q2 records, $33B 2026 capex, HBM4 mass production, and KLA raising 2026 WFE to ~$150B. Any efficiency gain is eaten by Jevon's paradox.
The consensus overweights a premise that just became contestable. Hybrid-linear attention (KDA) removes the linear-KV mechanism at ~73% memory saving, and Kimi K3 proved it frontier-viable AND open-sourced it — so adoption is cheap and likely (as MLA diffused post-DeepSeek-V3). Jevon's is NOT guaranteed to outrun the saving, because denser scaling laws mean smaller models per unit intelligence. So memory-demand durability is lower than a +257% tape implies — a small demand-growth deceleration could trigger game-theoretic profit-taking in a crowded trade.
The desk does not take a naked short (the bull spine has overwhelming current confirmation). It SEPARATES the scoreable from the narrative: adoption count and HBM-guidance trajectory are tracked, falsifiable measurements; the 'game-theoretic crash' mechanism is flagged unfalsifiable-in-advance and carries no weight. The edge is holding a scored, updating probability on a premise the momentum crowd treats as settled.
Ai Capex Cycle — signal 2026-08-040.65 moderateMemory is the fastest-growing major line except storage.
HBM40.80 strongEvery prior leg of this thesis argued from DEMAND — tokens, context length, cache pressure. This is the first evidence from the COST side, and it is independent: a bill of materials does not care why memory is needed. Memory growing faster than compute inside the same package means the accelerator vendor is choosing to spend marginal package cost on memory rather than on logic, which is a revealed preference from the party with the best information about the bottleneck. HELD AT 0.65 because it is a sell-side BOM estimate rather than a disclosed teardown, and this desk refused a sell-side memory TAM this month for failing an arithmetic check. AMD points the same way from a different direction: MI455X ships 432GB HBM4 per GPU against Vera Rubin's 288GB, a 1.5x capacity lead chosen as a competitive axis.
AI total addressable market (revenue, forward)0.60 moderateAI-attributable memory is ~70% of DRAM, so ~$1.18tn in 2028 against an interpolated ~$1.1-1.3tn AI TAM that year — roughly 90-100% of the revenue pool, which is not survivable for an input. Scope is what makes or breaks this check: the $1.69tn DRAM+NAND total covers phones, PCs and autos and cannot be set against an AI-only denominator. Held at 0.60 because the AI TAM itself is a supplied buy-side build with no public link, so this bounds the memory forecast only as well as the denominator is known.
KV cache bytes per token, by attention architecture0.30 weakHeld at 0.80 because the arithmetic is trivially checkable but production serving quantises, evicts and shares prefixes, so the figure is an upper bound.
Long-context inference (KV-cache growth)0.75 strongLARGELY REFUTED. Compression does not stop context growth from driving HBM demand, because the workload grew faster than the compression did. Kimi K3 ships the most aggressive compression available — Kimi Delta Attention linear layers plus MLA at a 3:1 ratio — and still exhausts a B300 node at EIGHT concurrent users: 3.25M tokens of cache budget against a workload of 142k input tokens per turn over a median 65 turns, with hit rate collapsing from a theoretical 95% to under 10% past that point (SemiAnalysis, free portion, 2026-08). THE ARITHMETIC: compression bought 7.3x; the agentic workload took ~71x on context per turn versus a chat turn. Compression lost by an order of magnitude. The general form of the error is holding two terms and multiplying neither — tokens-per-agentic-task and kv-bytes-per-token were both in the model and nothing joined them, the same way tokens and megawatts sat unconnected. THE SURVIVING CLAIM IS NARROW: the leg is contingent on attention architecture rather than on the calendar, and at 1M context a single GQA request wants more HBM than an eight-accelerator node holds while the MLA equivalent fits. WHAT WOULD SETTLE IT is the SHARE OF SERVED TOKENS by attention scheme, which nobody publishes and which is estimable from open-weight download counts and provider model catalogues.
Hybrid Linear Attention (constant-state KV)0.85 strongThe doubt is adoption breadth, scored separately.
Hybrid Linear Attention (constant-state KV) — kv cache scaling constant0.80 strongHeld at 0.80: constant in the attention layer does not mean constant end-to-end once weights and activations are counted.
Hybrid Linear Attention (constant-state KV) — memory saving vs full attention 0.730.70 moderateMemory (KV/state) per token — value linear-with-context0.75 strongThe bull spine assumes memory/token grows with context (linear KV cache). A constant-size recurrent state removes that growth at ~73% saving, so the mechanism the demand curve is priced on is not a law — it is an architecture choice that can change.
Kimi K3 (Moonshot)0.85 strongExistence is not the doubt.
Kimi K3 (Moonshot) — frontier proof confirmed0.75 strongHeld at 0.75 because 'frontier' is a comparative judgement against a moving benchmark, not a measured threshold.
Hybrid Linear Attention (constant-state KV) supplies Kimi K3 (Moonshot)1.00 strongHybrid-linear was dismissed as sub-frontier until K3 delivered near-full recall + 2.5x train efficiency at 1T+ scale and open-sourced it. Cheap-to-copy frontier proof is exactly how MLA diffused after DeepSeek V3 — so the demand-dampening is a live adoption vector, which is what makes the premise contestable NOW rather than someday.
Memory Super-Cycle0.80 strongNames the theme this conclusion argues with; the conclusion does not fail if the theme is renamed.
SK Hynix0.85 strongThe name most exposed to the outcome, with primary-sourced financials.
Hybrid Linear Attention (constant-state KV)0.80 strongFrontier hybrid-linear (constant-state) adoption — value single-lab0.60 moderateOne lab shipping constant-state attention at frontier quality is proof of possibility, not of adoption — and the thesis is a bet on adoption. It is why this conclusion is scored low rather than refuted.
HBM supply tightness — value severe0.80 strongTightness is what makes the question worth asking; it is not a condition of the adoption claim.
China HBM market entry — value in-development0.65 moderateA second supply source in development bears on durability but does not gate the adoption bet, and 'in-development' is a status with no date attached.
Incentive × Capacity — the indigenization / margin-compression generator0.70 moderateThe prior that a firm with both incentive and capacity to enter eventually does.
Tokens consumed per agentic task — value rising-fast0.80 strongThis premise is why compression losing is the base case.
Aggregate inference token demand — value rising-fast0.85 strongInference cost per token — value falling-fast0.75 strongFalling unit cost is the mechanism by which demand keeps rising; it does not gate whether hybrid-linear attention gets adopted.
Semis/tech sector valuation (FMP sector P/E) — value elevated0.70 moderateElevated valuation sets what a re-rating costs if the thesis is wrong. It is a risk statement, not a condition.
The variant resolves on falsifiable measurements (frontier-lab hybrid-linear adoption count; HBM demand-guidance trajectory; memory-per-token trend) that the momentum crowd is not tracking. The desk holds a probability on those, fades the 'premise is settled' consensus, and explicitly excludes the unfalsifiable game-theoretic-crash mechanism from the weighted case.
Pre-filled skeleton: memory-demand-durability.md