TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 36:12 · $GOOGL · TPU systolic arrays are the most efficient known circuit for matrix multiply; 128x128 arrays in older generations
episode briefing
Dwarkesh Patel

Chip design from the bottom up – Reiner Pope

2026-05-22 · 3 company · 5 thematic
sentiment
2 bull0 bear1 neu
speakers
dwarkesh patel

Host of the Dwarkesh Podcast, known for deeply researched long-form conversations on artificial intelligence, science, economics and history.

reiner pope

Reiner Pope is the CEO of MatX, a stealth AI chip startup. He previously worked on TPU architecture at Google. He delivers a blackboard-style lecture on transformer inference economics, covering roofline analysis, batch size optimization, MoE parallelism, scale-up networking, memory hierarchy, and training/inference compute tradeoffs.

quote
“Older TPUs were described as 128x128 of this circuit shown here. This ends up being the most efficient known circuit for implementing a matrix multiply. ... From a very high-level point of view, the GPU has a lot of tiny TPUs tiled across…”
—reiner pope
now playing · $GOOGL
$MATXbullish· high· posdwarkesh patel
Dwarkesh discloses angel investment in MatX; CEO reveals splittable systolic array architecture
Dwarkesh Patel is an angel investor in MatX. CEO Reiner Pope describes their 'splittable systolic array' design that can function as both large and small systolic arrays, amortizi…
$NVDA···neutral· mediumreiner pope
Nvidia's precision scaling shifts from 2x to 3x FP4/FP8 ratio in B300+ acknowledging quadratic area scaling
Reiner Pope explains that historically Nvidia doubled FLOPs when halving precision (2x ratio), but due to quadratic area scaling with bit-width, the true advantage is ~4x. Nvidia'…
$GOOGL···bullish· mediumreiner pope
TPU systolic arrays are the most efficient known circuit for matrix multiply; 128x128 arrays in older generations
Reiner Pope, drawing on TPU architecture knowledge, states that Google's TPU systolic arrays (128x128 in older generations) represent the most efficient known circuit for matrix m…