Leads product marketing for AI infrastructure at AWS, discussing the company's 15-year Nvidia partnership, 2M GPU fleet expansion, Trainium 3 custom silicon roadmap, and Cerebras disaggregated inference partnership.
Amazon's custom Trainium 3 chips offer 2-3x performance over prior generation and 30-40% better price-performance vs alternative accelerators, with plans to deploy 1M+ chips in 2026, reducing dependence on Nvidia and lowering customer inference costs.
AWS is expanding its Nvidia GPU fleet by 50% in a single year, signaling massive sustained demand for Nvidia hardware and a deepening strategic partnership that includes upcoming Rubin architecture.
AWS and Cerebras are collaborating on a disaggregated inference architecture separating prefill and decode stages to drive down dollar-per-token costs for customers deploying GenAI at scale.