Disaggregated inference drives specialized chip architectures like LPU-GPU co-packaging
Inference workloads are splitting into specialized components (prefill, decode, interactivity), creating demand for heterogeneous chip architectures where LPUs handle latency-sensitive task…