Co-host of SemiAnalysis podcast and co-developer of InferenceX, an open-source AI inference benchmarking suite. Attended GTC 2025 and works directly with Nvidia and AMD engineers on benchmark optimization.
Independent open-source InferenceX benchmarks show Nvidia chips achieve highest throughput and lowest cost per token across the board, with Nvidia's customer intimacy enabling optimized disaggregated inference solutions like LPU integration.
Groq's LPU architecture surpasses GPUs for specific disaggregated inference tasks like fast interactivity, leading to co-packaging with Nvidia's Rubin system, though the claim that Nvidia acquired Groq appears to be a misunderstanding.