TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 29:00 · $DEEPSEEK · DeepSeek's fine-grained MoE and sparse attention are architectural breakthroughs
episode briefing
Dwarkesh Patel

How GPT, Claude, and Gemini are actually trained and served – Reiner Pope

2026-04-29 · 4 company · 5 thematic
sentiment
3 bull0 bear1 neu
speakers
dwarkesh patel

Host of the Dwarkesh Podcast, known for deeply researched long-form conversations on artificial intelligence, science, economics and history.

reiner pope

Reiner Pope is the CEO of MatX, a stealth AI chip startup. He previously worked on TPU architecture at Google. He delivers a blackboard-style lecture on transformer inference economics, covering roofline analysis, batch size optimization, MoE parallelism, scale-up networking, memory hierarchy, and training/inference compute tradeoffs.

quote
“DeepSeek has published a sparse attention mechanism. I'll just put a plug in that some of the DeepSeek papers that have published sparse attention end up putting a square root in this term. DeepSeek's mixture of experts was a big change in…”
—reiner pope

episode shorts · 1

Neural Networks Are Cryptography in Reverse - Reiner Pope

now playing · $DEEPSEEK
$MATXneutral· low· posdwarkesh patel
MatX building custom AI chips targeting memory bandwidth bottleneck
Reiner Pope left Google TPU architecture to found MatX, a chip startup addressing the memory bandwidth constraints he analyzes. The host is an angel investor. The technical lectur…
$DEEPSEEKbullish· highreiner pope
DeepSeek's fine-grained MoE and sparse attention are architectural breakthroughs
DeepSeek V3's 256 experts with 32 activated (8x sparsity) and fine-grained expert design fundamentally changes the compute-memory tradeoff, enabling efficient inference at scale.…
$NVDA···bullish· highreiner pope
Nvidia's rack-scale architecture drives AI inference economics
Nvidia's progression from Hopper (8-GPU) to Blackwell (72-GPU NVL72) to Rubin (500+ GPU) scale-up domains directly solves the memory bandwidth bottleneck for sparse MoE models, en…
$GOOGL···bullish· mediumreiner pope
Google's early large scale-up domains gave Gemini inference advantage
Google deployed very large scale-up domains (TPU pods) long before Nvidia's rack-scale NVLink, allowing Gemini to train and serve larger sparse models with higher memory bandwidth…