← research bank

MoE inference: converted miners capture memory scarcity, hyperscalers leak it

gen-moe-inference · conviction — · status open · horizon — · as of 2026-08-06

MoE shifts the binding constraint from compute to memory capacity and interconnect, which hyperscalers provision to a compute-bound baseline and cannot rebalance mid-cycle. Converted miners with greenfield builds can size for memory-bound workloads from inception, capturing utilization the hyperscalers leak. The margin spread reflects legacy cost structure, not workload fit—today's losers may be tomorrow's only profi
Rests on filed figures, not on modelled shares. 12 premises (6 field, 5 edge, 1 entity); no derived cell is involved, so the undisclosed supply weights that put a range on other pages in this bank cannot move this one.

Exhibits

Exhibit 1Relative performance, indexed to 100How the names in this thesis have traded against SOXX.
86185285SOXX 22512mo, indexed to 100 at start · dashed = SOXX benchmark

Series available as data/gen-moe-inference.csv

Exhibit 2What the conviction is actually made ofEach premise and the number it composes to. A conjunction of plausible premises is far weaker than any of them.
MoE inference requires memory capacity and interconnect bandwidth more than compute throughput, inverting the resource profile thaMemory capacity (per accelerator and per no…100.0%Interconnect bandwidth (scale-up and scale-…100.0%Memory bandwidth (HBM-class) supplies Mixtu… †78.0%Mixture-of-experts inference †92.0%

† 2 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.

No uncertain claim. Every gating premise here is a fact the model verifies on each rebuild — a graph edge, a recorded weight, a computed cell. Those are preconditions, not risks, so there is nothing left to put a probability on. That makes this a derivation from current data rather than a forecast, and no composed figure is shown.

Weakest link: Memory capacity (per accelerator and per node) supplies Mixture-of-experts inference at 1.00 — Edge weight is 'strong'; all experts resident is architectural, but activation ratio variance limits precision

Hyperscalers with high operating margins provisioned fleets for compute-bound workloads and face utilization penalties rebalancingMicrosoft Corporati… — operating margin 45.… †94.0%Alphabet Inc. — operating margin 32.7% †94.0%Meta Platforms, Inc… — operating margin 41.… †94.0%Mixture-of-experts inference supplies Micro…100.0%

† 3 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.

No uncertain claim. Every gating premise here is a fact the model verifies on each rebuild — a graph edge, a recorded weight, a computed cell. Those are preconditions, not risks, so there is nothing left to put a probability on. That makes this a derivation from current data rather than a forecast, and no composed figure is shown.

Weakest link: Mixture-of-experts inference supplies Microsoft Corporation at 1.00 — Exposure exists but intensity unknown; hyperscaler status implies compute-optimized baseline is legacy burden

Converted miners with negative margins and greenfield builds can dimension for MoE from inception, avoiding the hyperscaler utilizTeraWulf Inc. — operating margin -127.8% †93.0%Core Scientific, I… — operating margin -45.… †93.0%Nebius Group N.V. — revenue growth $684 †91.0%Mixture-of-experts inference supplies Core…100.0%

† 3 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.

No uncertain claim. Every gating premise here is a fact the model verifies on each rebuild — a graph edge, a recorded weight, a computed cell. Those are preconditions, not risks, so there is nothing left to put a probability on. That makes this a derivation from current data rather than a forecast, and no composed figure is shown.

Weakest link: Mixture-of-experts inference supplies Core Scientific, Inc. at 1.00 — Exposure exists; converted miner status means no legacy ratio lock-in, though business model transition risk remains

The variant

Consensus

MoE inference is a pure efficiency win—better quality per FLOP. Hyperscalers with massive capital budgets and operating leverage (Microsoft 45% margin, Alphabet 33%, Meta 41%) should monetize architectural progress better than subscale converted bitcoin miners running negative margins. The supply chain sells into both, so the technology itself is margin-neutral.

Variant

MoE shifts the binding constraint from compute to memory capacity and interconnect, which hyperscalers provision to a compute-bound baseline and cannot rebalance mid-cycle. Converted miners with greenfield builds can size for memory-bound workloads from inception, capturing utilization the hyperscalers leak. The margin spread reflects legacy cost structure, not workload fit—today's losers may be tomorrow's only profitable hosts for MoE serving at scale.

Differentiator

Consensus reads margins as skill; supply-chain structure reveals they reflect architectural lock-in. Hyperscaler fleets were dimensioned for dense inference and training. MoE requires the *opposite* resource ratio, so installed-base leaders cannot deploy it economically without stranding compute—exactly the cost a greenfield operator avoids.

Falsifiers

Open questions

Reasoning chain

MoE inference requires memory capacity and interconnect bandwidth more than compute throughput, inverting the resource profile that defines existing GPU deployments VALID
premises

Strong enablement by capacity and interconnect vs. moderate by bandwidth confirms memory-bound, not FLOP-bound, profile

Hyperscalers with high operating margins provisioned fleets for compute-bound workloads and face utilization penalties rebalancing to memory-bound MoE VALID
premises

High margins reflect past optimization; MoE exposure on legacy infrastructure creates stranded-compute risk hyperscalers cannot escape

Converted miners with negative margins and greenfield builds can dimension for MoE from inception, avoiding the hyperscaler utilization trap VALID
premises

Negative margins mask structural advantage: no stranded compute from legacy provisioning, freedom to size memory and interconnect to MoE

Write-up

Pre-filled skeleton: gen-moe-inference.md