Co-host of SemiAnalysis podcast and co-developer of InferenceX. Focuses on agentic benchmark design, KV cache offloading strategies, and real-world inference performance modeling.
no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY
AMD's MI355X offers 1.5x HBM capacity versus Nvidia's B200, which could enable better prefix caching performance in multi-turn workloads, but current InferenceX benchmarks show Nvidia surpassing AMD across almost all aspects.