Local model inference and massive token spend signal compute shift to inference-time scaling
Holtz runs local TTS (Parakeet) on 128GB RAM Mac, spends $22k/month on API tokens, and anticipates agents running 10x longer in cloud environments unconstrained by local hardware, highlight…