TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 30:58 · $GROQ · Groq and Cerebras seen as accelerators for hybrid inference, not standalone replacements
episode briefing
Invest Like The Best

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

2026-08-25 · 16 company · 20 thematic
sentiment
6 bull1 bear9 neu
speakers
neil

Neil is the co-founder and CTO of K2 Space, leading engineering on the Mega-class satellite platform. He oversees development of 20kW Hall thrusters, deployable solar arrays, high-voltage power systems, and radiation-hardened electronics for operation across LEO, MEO, GEO, and deep space.

quote
“Crisis and Grock and maybe a couple others, you should think of them as accelerators. What they are really good at is being used in conjunction with an more traditional GPU like device that critically has this offchip memory built in.”
—neil
now playing · $GROQ
$SAILbullish· high· posneil
Ex-NVIDIA engineer building 'token factory' targeting 1000x cheaper inference via heterogeneous chips, distributed micro-data-centers, and intermittent renewable power
Sail Research aims to become the lowest-cost token provider by optimizing across the full stack: kernel-level GPU efficiency, heterogeneous chip orchestration (Nvidia, AMD, TPU, T…
$CBRS···neutral· mediumneil
Neil predicts hybrid role for Cerebras as MLP accelerator paired with GPUs
Cerebras' wafer-scale SRAM architecture excels at compute-bound MLP layers but lacks capacity for dynamic KV cache; it will serve as a specialized accelerator paired with GPUs tha…
$GROQneutral· lowneil
Neil groups Groq with Cerebras as specialized SRAM accelerators
Groq's SRAM-focused architecture similarly suits it for accelerator roles in hybrid systems rather than standalone inference.
$GROQneutral· mediumneil
Groq and Cerebras seen as accelerators for hybrid inference, not standalone replacements
Low-latency specialists like Groq and Cerebras will serve as accelerators paired with traditional GPUs that provide off-chip memory capacity for KV cache, creating a heterogeneous…
$CBRS···neutral· mediumneil
Cerebras wafer-scale SRAM excels at MLP weights but KV cache capacity forces hybrid GPU pairing
Cerebras' massive on-chip SRAM (21 PB/s bandwidth) is ideal for compute-bound MLP layers, but dynamic KV cache growth requires off-chip DRAM capacity, so optimal architecture pair…
$AMD···bullish· high· posneil
Buying AMD aggressively as market sleeps on programming difficulty
AMD chips offer great compute per dollar but are underutilized because vendors invest less in kernel optimization; Sail Research builds software to unlock that alpha and is buying…
$DEEPSEEKbullish· mediumneil
DeepSeek publishing order-of-magnitude KV cache compression advances yearly
DeepSeek's open research on KV cache compression demonstrates rapid algorithmic progress (10x+ per year), signaling massive remaining headroom to reduce memory bandwidth pressure…
$TSM···neutral· mediumneil
TSMC over-tightens process corners; Sail Research would accept wider variance for lower cost
TSMC spends heavily to minimize die-to-die variation, but inference workloads can tolerate lower-performing chips; relaxing process controls could increase usable die supply and r…
$OPENAIneutral· mediumneil
Closed-source labs' 3-6 month lead may not sustain premium as enterprise adoption lags
Frontier labs pay huge premiums to stay months ahead, but enterprises adopt slowly (still on older models), and open-source diffusion via distillation and AI-generated code is ine…
$ANTHROPICneutral· mediumneil
Anthropic's closed-source lead faces same diffusion dynamics as OpenAI
Same structural forces apply: enterprise inertia, inevitable distillation via AI-generated artifacts, and open-source catch-up will compress the value of a temporary model lead.
$SAILbullish· high· posneil
Sail Research builds 'token factory' targeting 1000x cheaper intelligence via heterogeneous chips, distributed data centers, and scavenged power
By optimizing software for throughput (not latency), buying undervalued chips (AMD, TPU, Trainium, custom), deploying 1MW distributed data centers with 95% uptime tolerance, and p…
$NVDA···neutral· mediumneil
Ex-Nvidia engineer sees slowing perf-per-watt gains, geopolitical risk overblown
Nvidia remains best-in-class for both latency and throughput today, but performance-per-watt scaling across process nodes (TSMC 5nm to 2nm) is minimal, and western fabs like Intel…
$INTC···bullish· mediumneil
Intel's western leading-edge process only ~2x behind TSMC, reducing geopolitical risk
If TSMC access is lost, Intel's best process nodes are at worst 2x worse performance-per-watt — a far smaller gap than chip-industry dialogue suggests, making supply shock managea…
$MU···bullish· highneil
HBM supply bottleneck structural; memory makers burned by cycles will not add capex fast
HBM capacity is the hardest bottleneck to expand — Micron, SK Hynix, Samsung have been burned by cyclical downturns and resist massive capex, so shortage persists and forces syste…
$AAPL···bearish· mediumneil
Neil predicts iPhone memory cuts and price hikes due to HBM shortage
HBM shortage will force Apple to reduce memory in iPhones and raise prices, as memory manufacturers prioritize AI data center demand.
$CRWV···neutral· mediumneil
Nvidia strategically builds neocloud ecosystem; CoreWeave a key beneficiary
Jensen Huang deliberately fosters a diverse set of neocloud buyers (like CoreWeave) to maintain competitive tension and avoid vertical integration that would alienate customers.