MoE inference: converted miners capture memory scarcity, hyperscalers leak it
gen-moe-inference · conviction — · status open · horizon — · as of 2026-08-06
MoE shifts the binding constraint from compute to memory capacity and interconnect, which hyperscalers provision to a compute-bound baseline and cannot rebalance mid-cycle. Converted miners with greenfield builds can size for memory-bound workloads from inception, capturing utilization the hyperscalers leak. The margin spread reflects legacy cost structure, not workload fit—today's losers may be tomorrow's only profi
Rests on filed figures, not on modelled shares. 12 premises (6 field, 5 edge, 1 entity); no derived cell is involved, so the undisclosed supply weights that put a range on other pages in this bank cannot move this one.
Exhibits
Exhibit 1Relative performance, indexed to 100How the names in this thesis have traded against SOXX.
Series available as data/gen-moe-inference.csv
Exhibit 2What the conviction is actually made ofEach premise and the number it composes to. A conjunction of plausible premises is far weaker than any of them.
MoE inference requires memory capacity and interconnect bandwidth more than compute throughput, inverting the resource profile tha
† 2 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.
No uncertain claim. Every gating premise here is a fact the model verifies on each rebuild — a graph edge, a recorded weight, a computed cell. Those are preconditions, not risks, so there is nothing left to put a probability on. That makes this a derivation from current data rather than a forecast, and no composed figure is shown.
Weakest link: Memory capacity (per accelerator and per node) supplies Mixture-of-experts inference at 1.00 — Edge weight is 'strong'; all experts resident is architectural, but activation ratio variance limits precision
Hyperscalers with high operating margins provisioned fleets for compute-bound workloads and face utilization penalties rebalancing
† 3 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.
No uncertain claim. Every gating premise here is a fact the model verifies on each rebuild — a graph edge, a recorded weight, a computed cell. Those are preconditions, not risks, so there is nothing left to put a probability on. That makes this a derivation from current data rather than a forecast, and no composed figure is shown.
Weakest link: Mixture-of-experts inference supplies Microsoft Corporation at 1.00 — Exposure exists but intensity unknown; hyperscaler status implies compute-optimized baseline is legacy burden
Converted miners with negative margins and greenfield builds can dimension for MoE from inception, avoiding the hyperscaler utiliz
† 3 premises marked supporting — shown and arguable, but the conclusion does not depend on them, so they are not multiplied into the composed figure. Citing a filed figure should not cost conviction.
No uncertain claim. Every gating premise here is a fact the model verifies on each rebuild — a graph edge, a recorded weight, a computed cell. Those are preconditions, not risks, so there is nothing left to put a probability on. That makes this a derivation from current data rather than a forecast, and no composed figure is shown.
Weakest link: Mixture-of-experts inference supplies Core Scientific, Inc. at 1.00 — Exposure exists; converted miner status means no legacy ratio lock-in, though business model transition risk remains
The variant
Consensus
MoE inference is a pure efficiency win—better quality per FLOP. Hyperscalers with massive capital budgets and operating leverage (Microsoft 45% margin, Alphabet 33%, Meta 41%) should monetize architectural progress better than subscale converted bitcoin miners running negative margins. The supply chain sells into both, so the technology itself is margin-neutral.
Variant
MoE shifts the binding constraint from compute to memory capacity and interconnect, which hyperscalers provision to a compute-bound baseline and cannot rebalance mid-cycle. Converted miners with greenfield builds can size for memory-bound workloads from inception, capturing utilization the hyperscalers leak. The margin spread reflects legacy cost structure, not workload fit—today's losers may be tomorrow's only profitable hosts for MoE serving at scale.
Differentiator
Consensus reads margins as skill; supply-chain structure reveals they reflect architectural lock-in. Hyperscaler fleets were dimensioned for dense inference and training. MoE requires the *opposite* resource ratio, so installed-base leaders cannot deploy it economically without stranding compute—exactly the cost a greenfield operator avoids.
Falsifiers
claim: Hyperscalers report MoE utilization penalties · criterion: Microsoft, Alphabet, or Meta disclose lower-than-baseline utilization or margin compression attributed to MoE inference in earnings or cloud segment commentary · horizon: 2027-07-31 · settles: confirmed
claim: Converted miners reach positive operating margins on MoE-serving revenue · criterion: TeraWulf, Core Scientific, or IREN report positive operating margin in a quarter with disclosed AI/inference revenue contribution >50% of total · horizon: 2027-10-31 · settles: confirmed
claim: MoE models dominate frontier inference mix · criterion: OpenAI, Anthropic, or xAI publicly state MoE architectures represent >60% of served inference tokens across their product lines · horizon: 2027-03-31 · settles: confirmed
Open questions
What is the actual memory-to-FLOP ratio hyperscalers deployed in 2023–2025 vs. the ratio MoE models require at scale?
Can hyperscalers retrofit interconnect or must they strand compute to serve MoE economically?
At what MoE adoption rate does the converted-miner cost advantage offset hyperscaler distribution and customer lock-in?
Reasoning chain
MoE inference requires memory capacity and interconnect bandwidth more than compute throughput, inverting the resource profile that defines existing GPU deploymentsVALID
premises
Memory capacity (per accelerator and per node) supplies Mixture-of-experts inference1.00 strong
Edge weight is 'strong'; all experts resident is architectural, but activation ratio variance limits precision
Interconnect bandwidth (scale-up and scale-out) supplies Mixture-of-experts inference1.00 strong
Edge weight 'strong'; routing is collective, but scale-up vs scale-out binding unclear from model
Edge weight 'moderate'; matters but secondary to capacity given sparse activation pattern
Mixture-of-experts inference0.92 strong
Summary states MoE 'shifts binding resource away from compute'—architectural claim, high confidence
Strong enablement by capacity and interconnect vs. moderate by bandwidth confirms memory-bound, not FLOP-bound, profile
Hyperscalers with high operating margins provisioned fleets for compute-bound workloads and face utilization penalties rebalancing to memory-bound MoEVALID
premises
Microsoft Corporation — operating margin 45.2%0.94 strong
Reported figure; margin embeds historical provisioning decisions, not forward workload fit
Alphabet Inc. — operating margin 32.7%0.94 strong
Reported figure; high margin suggests optimized for legacy workload mix
Meta Platforms, Inc. — operating margin 41.4%0.94 strong
Reported figure; fleet already deployed, MoE requires different ratio
Mixture-of-experts inference supplies Microsoft Corporation1.00 strong
Exposure exists but intensity unknown; hyperscaler status implies compute-optimized baseline is legacy burden
High margins reflect past optimization; MoE exposure on legacy infrastructure creates stranded-compute risk hyperscalers cannot escape
Converted miners with negative margins and greenfield builds can dimension for MoE from inception, avoiding the hyperscaler utilization trapVALID
premises
TeraWulf Inc. — operating margin -127.8%0.93 strong
Reported; negative margin confirms greenfield/ramp, not legacy optimization
Core Scientific, Inc. — operating margin -45.8%0.93 strong
Reported; unprofitable today but unconstrained by installed base
Nebius Group N.V. — revenue growth $6840.91 strong
Reported growth; ramp speed implies recent provisioning decisions align with current workload needs