newsroom
AI Infrastructure · Open-source token share gains shift margin from model layer to infrastructure layer
now playing · AI Infrastructure
Open-source token explosion is 'dark matter' driving net compute demand acceleration
Open-source models (GLM 5.2, Kimmy K3, Neotron) and inference clouds (Fireworks, Together, Modal, Base10) are accelerating token volumes. Public markets miss this 'dark matter.' Open-source…
Contracted compute rolling to spot drives hyperscaler cash flow acceleration, funding buildout without debt
Hyperscalers have massive contracted compute bases at prices far below current spot rates. As contracts roll off, revenue reprices higher, accelerating operating cash flow from ~28% to 35%+…
Contracted compute repricing to spot drives hyperscaler cash flow acceleration
Hyperscalers' contracted GPU base trades at a massive discount to spot. As contracts roll off, compute reprices higher, accelerating operating cash flow from 28% to 35% YoY. Consensus model…
Open-source token share gains shift margin from model layer to infrastructure layer
Open-source models (GLM 5.2, Kimi K3, Nemotron) are accelerating and taking token share from frontier models. This shifts margin dollars from the 90%-margin frontier layer to the 30%-margin…
Memory & Storagetailwindscore 9/10gavin baker
Memory LTAs create game-theory lock-in; vendors gain strategic leverage over hyperscalers
Hyperscalers are signing long-term supply agreements with memory vendors (Micron, SK Hynix, Samsung) with prepay and floor/ceiling pricing. Breaking an LTA risks permanent allocation loss w…
Inference clouds (Fireworks, Together, Modal, Base10) growing at frontier-lab pace with minimal cash burn
A new layer of capital-efficient inference clouds is emerging, monetizing open-source models via router/RL fine-tuning products. They grow nearly as fast as frontier labs but burn little ca…
Credit market stress (widening CDS, rising real yields) is manageable if compute repricing continues; otherwise flops become scarcer and more valuable
Rising real yields and widening CDS spreads signal credit market concern about AI capex funding. However, if contracted compute reprices higher as modeled, hyperscaler cash flows accelerate…
AI Agentstailwindscore 8/10gavin baker
Agentic AI adoption at 500K users today; scaling to 1% of global population implies 100M+ users and massive compute demand
Only ~500K people use agentic AI today. Scaling to 1% of 8B people (80M) or 10% (800M) creates exponential compute demand. AI natives already spend 20-50% of comp budgets on tokens and grow…
Acute compute shortage: only 500K agentic AI users vs 8B population implies massive latent demand
Approximately 500,000 people globally use agentic AI today. Scaling to 1% (80M) or 10% (800M) of population creates exponential compute demand. AI natives already spend 20-50% of compensati…
Memory & Storagetailwindscore 8/10gavin baker
Memory LTAs create game-theoretic lock-in; hyperscalers cannot break without risking future HBM allocations
Memory vendors (Micron, SK Hynix, Samsung) are shifting to long-term supply agreements with floor/ceiling pricing. With 4+ major buyers and supply-constrained HBM, breaking an LTA risks per…
Nvidia credit-wrapper model de-risks financing and increases revenue per gigawatt
Nvidia's new business model provides a credit wrapper (equity investment + revenue share above GPU price floor) for GPU buyers. This is not vendor financing—third parties provide debt—but N…
Regulatory risk is largest threat; industry PR failure allows 'data centers raise prices, take water, take jobs' narrative to spread unchecked
Political narrative (data centers raise electricity prices, consume water, eliminate jobs) is factually wrong — data centers lower local power prices via behind-the-meter deals, build commu…
Japanese capacitor stocks completed full 3-year cycle in 6 weeks; market speed exceeds fundamentals
Semiconductor supply-chain stocks (capacitors) saw vertical moves and round-trips in 6 weeks that historically took 3 years. Claude-driven homogeneous interpretation of news creates Walter…
Power/land energizing accelerating; turbines, diesel gens, regulatory improving; data centers lower local power prices
Gigawatt energization is the bottleneck, but capitalism is solving it: reconditioned aircraft turbines, diesel generators, behind-the-meter deals lowering community power prices, and regula…
China DUV breakthrough is real phase transition but 25 years behind; market overreacted
China's alleged DUV lithography capability is a phase transition (prop plane to jet turbine) but remains ~25 years behind EUV. Learning-by-doing cannot be accelerated; market overreacted to…
China DUV breakthrough real but 25 years behind EUV; decoupling self-reinforcing; ASML impact 5+ years out
China's DUV capability is a phase transition (prop plane vs jet turbine) but EUV remains 25 years ahead. Learning-by-doing cycles cannot be accelerated. Market overreacted near-term; any AS…
Inference clouds (Fireworks, Together, Modal, Baseten) growing at frontier-lab pace with superior cash efficiency
Open-source inference clouds are growing nearly as fast as frontier labs did in early days but burning minimal cash. They enable AI natives to customize open models via RL and routing, capt…
SRAM accelerators (Etched) could disaggregate inference into prefill/attention/feed-forward, boosting ROI on installed base
Disaggregating inference — prefill on one chip, attention on HBM-heavy chip, feed-forward on SRAM accelerators — optimizes each stage. SRAM beats HBM for feed-forward networks regardless of…
Semiconductorstailwindscore 7/10gavin baker
SRAM-based inference accelerators enable disaggregated prefill/attention/FFN for higher ROI
Disaggregating inference into prefill (SRAM chip), attention (HBM chip), and feed-forward network (SRAM chip) improves ROI on installed compute. SRAM excels at feed-forward networks; worklo…
Orbital compute becoming tangible: SpaceX Starship, StarCloud, Starlink lasers
SpaceX's Starship progress, Benchmark's StarCloud investment, and Starlink's laser interconnect technology make orbital compute a credible long-term option. SpaceX's internal launch cost ad…
Continual/sample-efficient learning breakthrough could reduce training compute but increase inference demand
If labs solve continual learning (training on 10T tokens vs 300T, then sample-efficient real-world learning), training compute demand could drop. However, Baker argues inference demand woul…
Continual/sample-efficient learning (SSI, new labs) could disrupt pre-training compute demand but grow inference
If labs like SSI solve continual learning (train on 10T tokens, then learn sample-efficiently in wild), massive pre-training compute demand could face a discontinuity. However, this would l…
Orbital compute gaining credibility; Benchmark-funded StarCloud partners with SpaceX Starlink lasers
Orbital compute (StarCloud, SpaceX Starship) is becoming real. Benchmark's investment without SpaceX's launch cost advantage validates the thesis. Starship landing success and Starlink lase…