TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 3:17 · AI Hardware & Chip Architecture · Disaggregated inference drives specialized chip architectures like LPU-GPU co-packaging
episode briefing
SemiAnalysis

Bryan Shan x Cameron Quilici | Researcher Conversations at GTC

2026-04-07 · 3 company · 3 thematic
sentiment
1 bull0 bear2 neu
speakers
bryan shan

Co-host of SemiAnalysis podcast and co-developer of InferenceX, an open-source AI inference benchmarking suite. Attended GTC 2025 and works directly with Nvidia and AMD engineers on benchmark optimization.

cameron quilici

Co-host of SemiAnalysis podcast and co-developer of InferenceX. Focuses on agentic benchmark design, KV cache offloading strategies, and real-world inference performance modeling.

now playing · AI Hardware & Chip Architecture
AI Hardware & Chip Architecturetailwindscore 8/10bryan shan
Disaggregated inference drives specialized chip architectures like LPU-GPU co-packaging
Inference workloads are splitting into specialized components (prefill, decode, interactivity), creating demand for heterogeneous chip architectures where LPUs handle latency-sensitive task…
AI Infrastructuretailwindscore 7/10cameron quilici
Open-source InferenceX benchmarks expose real-world Pareto frontier for chip evaluation
Independent, open-source benchmarking with random-data baselines reveals worst-case chip performance, while upcoming agentic multi-turn benchmarks with prefix caching will better reflect pr…
Memory & Storagetailwindscore 7/10bryan shan
KV cache offloading becomes critical bottleneck for long-context inference at scale
Million-token context windows (e.g., Opus 4.6) exceed HBM capacity, forcing KV cache offloading to slower memory tiers; the frontier labs are likely six months ahead of open-source solution…