newsroom
AI Infrastructure · Inference complexity accelerating from scale, diversity, and agents
now playing · AI Infrastructure
AI Infrastructuretailwindscore 9/10simon mo
VLM emerging as universal inference layer abstracting all AI hardware
VLM is becoming the de facto standard inference engine running on 400k-500k+ GPUs globally, with a community of 2000+ contributors and participation from every major chip vendor (Nvidia, AM…
AI Infrastructuretailwindscore 8/10wuk kuan
Inference complexity accelerating from scale, diversity, and agents
Three structural forces make inference harder: models scaling to multi-trillion parameters requiring massive sharding; explosion of model architectures (sparse attention, linear attention)…
Open Source AItailwindscore 8/10wuk kuan
Open source will win AI infrastructure due to model and hardware diversity
The complexity of matching diverse model architectures (sparse attention, linear attention, varying context lengths) to diverse hardware (H100, B200, TPU, etc.) for each use case makes a si…
Semiconductorstailwindscore 7/10wuk kuan
Model-hardware co-design creates new abstraction layer opportunity
Each Nvidia chip generation (H100, B200, GB200) demands different model architectures for optimal performance, while TPUs and other accelerators diverge further; this vertical integration c…