Mistral CTO praises Blackwell GB200/GB300 and NVFP4 for 2.5x training speedup on MoE models
NVIDIA's Blackwell architecture delivers a 2.5x out-of-the-box improvement for training large sparse mixture-of-experts models, and NVFP4 quantization enables efficient inference…