Inference costs projected to drop 90% in 1-2 years via specialized chips and models
The shift from training to inference workloads, combined with diverse purpose-built chips and smaller specialized models routed via agentic frameworks, will radically reduce compute costs a…