← research bank

Memory-demand durability — does algorithmic efficiency cap the memory super-cycle?

memory-demand-durability · conviction low · status open · horizon 2026-2028 (adoption is multi-quarter) · as of 2026-08-03

The memory re-rating (SK Hynix +257% rev, HBM/DRAM 10-30x) prices long-context AI needing linearly-growing KV cache. Kimi K3 proved hybrid-linear (constant-state) attention works at frontier scale with near-full recall and ~73% memory saving — the net-new fact that makes the load-bearing premise contestable. This thesis holds the DEBATE, not a naked short: the variant is that the market underweights algorithmic efficiency; the falsifiable gate is frontier-lab ADOPTION, and Jevon's + deployment inefficiency are real offsets.
Rests on filed figures, not on modelled shares. 23 premises (11 field, 10 entity, 1 signal, 1 edge); no derived cell is involved, so the undisclosed supply weights that put a range on other pages in this bank cannot move this one.

Exhibits

Exhibit 1Relative performance, indexed to 100How the names in this thesis have traded against SOXX.
116061201MU 739000660.KS 573005930.KS 331SOXX 22512mo, indexed to 100 at start · dashed = SOXX benchmark

Series available as data/memory-demand-durability.csv

Exhibit 2What the conviction is actually made ofEach premise and the number it composes to. A conjunction of plausible premises is far weaker than any of them.
Memory is taking share of the bill of materials, which is cost-side evidence the demand thesis previously lackedAi Capex Cycle — signal 2026-08-0465.0%HBM480.0%COMPOSED (and)52.0%

52% if the 2 gates are independent, 65% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: Ai Capex Cycle — signal 2026-08-04 at 0.65 — Wells Fargo NVL72 BOM: memory (HBM+LPDDR) rises from 26.6% to 30.8% of rack cost between GB300 and Vera Rubin, growing +128% against compute's +116%.

The long-context leg of the memory thesis is architecture-dependent, not time-dependentAI total addressable market (revenue, forwa…60.0%KV cache bytes per token, by attention arch…30.0%Long-context inference (KV-cache growth)75.0%COMPOSED (and)13.5%

14% if the 3 gates are independent, 30% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: KV cache bytes per token, by attention architecture at 0.30 — Computed from published configs: GQA at 252 KB/token against MLA at 34 KB/token, a 7.3x spread. Held at 0.80 because the arithmetic is trivially check

Hybrid-linear (constant-state) attention breaks the linear-KV-cache premise the memory super-cycle rests onHybrid Linear Attention (constant-state KV)85.0%Hybrid Linear At… — kv cache scaling consta… †80.0%Hybrid Linear Attention (constant-state KV)… †70.0%Memory (KV/state… — value linear-with-conte… †75.0%COMPOSED (and)85.0%

† 3 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.

One gating premise, so the conclusion is exactly as strong as it. The figure is an ordering device, not a calibrated probability — see how the numbers are made.

Weakest link: Hybrid Linear Attention (constant-state KV) at 0.85 — Hybrid-linear attention exists and is implemented, not a paper proposal. The doubt is adoption breadth, scored separately.

Kimi K3 is the net-new fact: hybrid-linear now works at frontier scale, so industry adoption is credible, not hypotheticalKimi K3 (Moonshot)85.0%Kimi K3 (Moonshot… — frontier proof confirm… †75.0%Hybrid Linear Attention (constant-state KV)…100.0%COMPOSED (and)85.0%

† 1 premise marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.

85% if the 2 gates are independent, 85% if they move together. They are claims about one industry, so the truth is between and nobody can say where. Treat this as an ordering device rather than a calibrated probability — the ranking of premises is the information, not the level.

Weakest link: Kimi K3 (Moonshot) at 0.85 — Kimi K3 ships and is documented. Existence is not the doubt.

Therefore memory-demand durability is a scored bet on ADOPTION, not a settled given — and the tradable claim is measurableMemory Super-Cycle †80.0%SK Hynix †85.0%Hybrid Linear Attention (constant-state KV) †80.0%Frontier hybrid-linear (c… — value single-l… †60.0%HBM supply tightness — value severe †80.0%China HBM market entr… — value in-developme… †65.0%Incentive × Capacity — the indigenization /… †70.0%Tokens consumed per agen… — value rising-fa… †80.0%Aggregate inference toke… — value rising-fa… †85.0%Inference cost per toke… — value falling-fa… †75.0%Semis/tech sector valuation… — value elevat… †70.0%

† 11 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.

No gating premise. Every premise here is supporting evidence, so this conclusion states no necessary condition — it is asserted from cited data rather than derived from a claim that could fail. Read it as a summary, not as a falsifiable call.

Weakest link: Frontier hybrid-linear (constant-state) adoption — value single-lab at 0.60 — SINGLE-LAB is the load-bearing premise and the weakest. One lab shipping constant-state attention at frontier quality is proof of possibility, not of

The variant

Consensus

Long-context/agentic AI drives structurally rising memory demand; the KV-cache-per-token requirement is a given, so HBM/DRAM shortage persists — validated by SK Hynix Q2 records, $33B 2026 capex, HBM4 mass production, and KLA raising 2026 WFE to ~$150B. Any efficiency gain is eaten by Jevon's paradox.

Variant

The consensus overweights a premise that just became contestable. Hybrid-linear attention (KDA) removes the linear-KV mechanism at ~73% memory saving, and Kimi K3 proved it frontier-viable AND open-sourced it — so adoption is cheap and likely (as MLA diffused post-DeepSeek-V3). Jevon's is NOT guaranteed to outrun the saving, because denser scaling laws mean smaller models per unit intelligence. So memory-demand durability is lower than a +257% tape implies — a small demand-growth deceleration could trigger game-theoretic profit-taking in a crowded trade.

Differentiator

The desk does not take a naked short (the bull spine has overwhelming current confirmation). It SEPARATES the scoreable from the narrative: adoption count and HBM-guidance trajectory are tracked, falsifiable measurements; the 'game-theoretic crash' mechanism is flagged unfalsifiable-in-advance and carries no weight. The edge is holding a scored, updating probability on a premise the momentum crowd treats as settled.

Falsifiers

Open questions

Reasoning chain

Memory is taking share of the bill of materials, which is cost-side evidence the demand thesis previously lacked VALID
premises

Every prior leg of this thesis argued from DEMAND — tokens, context length, cache pressure. This is the first evidence from the COST side, and it is independent: a bill of materials does not care why memory is needed. Memory growing faster than compute inside the same package means the accelerator vendor is choosing to spend marginal package cost on memory rather than on logic, which is a revealed preference from the party with the best information about the bottleneck. HELD AT 0.65 because it is a sell-side BOM estimate rather than a disclosed teardown, and this desk refused a sell-side memory TAM this month for failing an arithmetic check. AMD points the same way from a different direction: MI455X ships 432GB HBM4 per GPU against Vera Rubin's 288GB, a 1.5x capacity lead chosen as a competitive axis.

The long-context leg of the memory thesis is architecture-dependent, not time-dependent VALID
premises

LARGELY REFUTED. Compression does not stop context growth from driving HBM demand, because the workload grew faster than the compression did. Kimi K3 ships the most aggressive compression available — Kimi Delta Attention linear layers plus MLA at a 3:1 ratio — and still exhausts a B300 node at EIGHT concurrent users: 3.25M tokens of cache budget against a workload of 142k input tokens per turn over a median 65 turns, with hit rate collapsing from a theoretical 95% to under 10% past that point (SemiAnalysis, free portion, 2026-08). THE ARITHMETIC: compression bought 7.3x; the agentic workload took ~71x on context per turn versus a chat turn. Compression lost by an order of magnitude. The general form of the error is holding two terms and multiplying neither — tokens-per-agentic-task and kv-bytes-per-token were both in the model and nothing joined them, the same way tokens and megawatts sat unconnected. THE SURVIVING CLAIM IS NARROW: the leg is contingent on attention architecture rather than on the calendar, and at 1M context a single GQA request wants more HBM than an eight-accelerator node holds while the MLA equivalent fits. WHAT WOULD SETTLE IT is the SHARE OF SERVED TOKENS by attention scheme, which nobody publishes and which is estimable from open-weight download counts and provider model catalogues.

Hybrid-linear (constant-state) attention breaks the linear-KV-cache premise the memory super-cycle rests on VALID
premises

The bull spine assumes memory/token grows with context (linear KV cache). A constant-size recurrent state removes that growth at ~73% saving, so the mechanism the demand curve is priced on is not a law — it is an architecture choice that can change.

Kimi K3 is the net-new fact: hybrid-linear now works at frontier scale, so industry adoption is credible, not hypothetical VALID
premises

Hybrid-linear was dismissed as sub-frontier until K3 delivered near-full recall + 2.5x train efficiency at 1T+ scale and open-sourced it. Cheap-to-copy frontier proof is exactly how MLA diffused after DeepSeek V3 — so the demand-dampening is a live adoption vector, which is what makes the premise contestable NOW rather than someday.

Therefore memory-demand durability is a scored bet on ADOPTION, not a settled given — and the tradable claim is measurable VALID
premises

The variant resolves on falsifiable measurements (frontier-lab hybrid-linear adoption count; HBM demand-guidance trajectory; memory-per-token trend) that the momentum crowd is not tracking. The desk holds a probability on those, fades the 'premise is settled' consensus, and explicitly excludes the unfalsifiable game-theoretic-crash mechanism from the weighted case.

Sources

Write-up

Pre-filled skeleton: memory-demand-durability.md