newsroom
AI Infrastructure · Local model inference and massive token spend signal compute shift to inference-time scaling
now playing · AI Infrastructure
AI coding shifts from writing code to orchestrating agent fleets with prompts as core IP
Charlie Holtz describes a workflow where he manages multiple concurrent AI agents like a conductor, reviewing PRs rather than writing code; he argues code becomes 'sawdust' while prompts be…
Local model inference and massive token spend signal compute shift to inference-time scaling
Holtz runs local TTS (Parakeet) on 128GB RAM Mac, spends $22k/month on API tokens, and anticipates agents running 10x longer in cloud environments unconstrained by local hardware, highlight…