newsroom
AI Infrastructure · Inference becomes new AI moat as demand shifts 80% to serving, agents drive exponential token burn
now playing · AI Infrastructure
AI Infrastructuretailwindscore 9/10ejaaz
Inference becomes new AI moat as demand shifts 80% to serving, agents drive exponential token burn
Inference compute demand has flipped from 1/3 to 2/3 of total AI workloads and heads toward 80%; autonomous agents running 24/7 will burn tokens at unprecedented scale, making tokens-per-se…
Semiconductorstailwindscore 8/10ejaaz
Purpose-built inference ASICs disrupt GPU dominance with 10-50x efficiency via transformer hardcoding
Specialized ASICs (Etched, Cerebras, Groq, OpenAI Jalapeno) achieve 80-90% utilization vs GPU's 30-40% by hardcoding transformer computation graphs into silicon, halving voltage for 75% pow…
Performance-per-watt becomes primary metric as data center power constraints bind
Voltage halving delivers 75% power reduction for same intelligence output; with data center power as the scarcest resource, chips maximizing tokens-per-second-per-watt (Etched, Cerebras, cu…
Hardcoding transformer into silicon yields 10-50x gains but creates existential architecture lock-in risk
Etched's chips only run transformer-based models; if a post-transformer architecture emerges (e.g., from Karpathy), the silicon becomes obsolete and requires full redesign. Nvidia GPUs and…