newsroom
AI Infrastructure · Open-source InferenceX benchmarks expose real-world Pareto frontier for chip evaluation
now playing · AI Infrastructure
Disaggregated inference drives specialized chip architectures like LPU-GPU co-packaging
Inference workloads are splitting into specialized components (prefill, decode, interactivity), creating demand for heterogeneous chip architectures where LPUs handle latency-sensitive task…
Open-source InferenceX benchmarks expose real-world Pareto frontier for chip evaluation
Independent, open-source benchmarking with random-data baselines reveals worst-case chip performance, while upcoming agentic multi-turn benchmarks with prefix caching will better reflect pr…
Memory & Storagetailwindscore 7/10bryan shan
KV cache offloading becomes critical bottleneck for long-context inference at scale
Million-token context windows (e.g., Opus 4.6) exceed HBM capacity, forcing KV cache offloading to slower memory tiers; the frontier labs are likely six months ahead of open-source solution…