TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 6:00 · AI Agents · Background agents with self-administered token budgets will consume 90% of inference compute within years
episode briefing
Invest Like The Best

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

2026-08-25 · 16 company · 20 thematic
sentiment
6 bull1 bear9 neu
speakers
neil

Neil is the co-founder and CTO of K2 Space, leading engineering on the Mega-class satellite platform. He oversees development of 20kW Hall thrusters, deployable solar arrays, high-voltage power systems, and radiation-hardened electronics for operation across LEO, MEO, GEO, and deep space.

now playing · AI Agents
Open Source AItailwindscore 7/10neil
Open source model sovereignty drives robust inference market for customized models
Enterprises increasingly demand ownership and control over model weights, creating sustained demand for serving open-source and customized models rather than relying solely on closed APIs.
AI Agentstailwindscore 9/10neil
Background agents with self-administered token budgets will consume 90% of inference compute within years
Human-in-the-loop chatbots saturate human attention; proactive agents running for hours/days on verifiable tasks (coding, research, security) have unbounded token demand and will shift infe…
AI Infrastructuretailwindscore 9/10neil
Inference shifting from latency-optimized chatbots to throughput-optimized background agents
As agents run hours-long tasks autonomously, the GPU's throughput-oriented happy path (large batches, high utilization) becomes dominant over latency-optimized interactive serving, unlockin…
AI Hardware & Chip Architecturetailwindscore 9/10neil
Shift to background agents favors throughput over latency, enabling diverse chip architectures
As AI workloads shift from interactive chatbots to long-horizon background agents, the GPU trade-off favors throughput-optimized serving, allowing use of chips without high-speed interconne…
AI Hardware & Chip Architecturetailwindscore 9/10neil
Throughput-optimized inference stacks will replace latency-optimized chatbot stacks as agents move to background
The shift from interactive chatbots to long-horizon background agents eliminates the latency-throughput tradeoff, enabling rack-scale batch processing that maximizes GPU utilization and low…
Networking & Optical Infrastructuretailwindscore 8/10neil
NVLink essential for low-latency but irrelevant for throughput-oriented background inference
High-speed GPU interconnect (NVLink) is only critical for latency-sensitive serving; throughput-optimized background inference can use alternative parallelism schemes (expert, pipeline) on…
Semiconductorstailwindscore 9/10neil
Heterogeneous chip fleets and software-defined hardware arbitrage replace Nvidia monoculture
No single chip wins all workloads; the optimal inference stack mixes Nvidia (NVLink/latency), AMD (compute/$), custom ASICs (TPU/Trainium), and accelerators (Cerebras/Groq) via a software l…
AI Bubble / Capex Debatetailwindscore 8/10neil
Inference spend is non-speculative and monotonically increasing, unlike training capex
Investor fears of semiconductor capex bubble (semis at 20% of S&P) are misplaced because current spend is on inference with immediate ROI, not speculative training; inference demand will gr…
AI Bubble / Capex Debatetailwindscore 8/10neil
Inference spend is non-speculative and monotonically increasing, unlike training capex cycles
Token consumption is driven by immediate economic value (enterprise caps on cloud coding spend), not speculative training runs; this makes current semiconductor demand fundamentally differe…
Data Center Infrastructuretailwindscore 8/10neil
Distributed 1MW liquid-cooled data centers enable inference at 95% uptime, unlocking stranded power
Training requires massive clustered GPUs with high-reliability infrastructure; inference can tolerate single-digit MW sites with minimal redundancy (single power feed, one fiber, no generat…
Data Center Infrastructuretailwindscore 9/10neil
Distributed 1MW data centers on intermittent renewables viable for background inference
Background agent workloads can tolerate 95% uptime, enabling use of small (1MW) data centers powered by intermittent solar/wind without expensive redundancy, unlocking stranded power and la…
Energy & Power Generationtailwindscore 8/10neil
Intermittent solar/wind power becomes viable for compute via workload migration and idle-chip economics
If chips are cheap enough, tolerating days-long renewable outages by migrating workloads elsewhere is economical; this unlocks vast stranded renewable capacity that traditional data centers…
Memory & Storagetailwindscore 8/10neil
KV cache compression is the biggest software efficiency bottleneck, with order-of-magnitude gains possible
The KV cache remains largely uncompressed, consuming kilobytes per token; DeepSeek's research shows annual order-of-magnitude compression improvements, signaling massive remaining headroom…
Memory & Storagetailwindscore 9/10neil
KV cache compression is the next order-of-magnitude efficiency frontier; current storage is uncompressed by 10-100x
The KV cache grows dynamically with context length and currently stores many kilobytes per token at far above its information entropy; DeepSeek and others are achieving order-of-magnitude c…
Memory & Storagetailwindscore 8/10neil
KV cache compression and flash offload are the next 10-100x inference efficiency frontiers
KV cache currently stores uncompressed, high-entropy representations; algorithmic compression (DeepSeek) and architectural offload to cheap flash memory can reduce memory bandwidth pressure…
Advanced Manufacturingtailwindscore 7/10neil
Accepting wider process corners at TSMC could increase effective yield for inference workloads
TSMC's tight process corner controls add cost; inference workloads can tolerate wider chip-to-chip variance, enabling use of "worst" chips that would otherwise be rejected, lowering effecti…
Open Source AItailwindscore 8/10neil
Open-source models inevitable due to AI-generated code diffusion and enterprise sovereignty demands
Distillation is unstoppable because AI-generated artifacts (GitHub repos, etc.) already permeate training data; enterprises want weight ownership and deployment control, creating structural…
AI Economics & Business Modelstailwindscore 9/10neil
Token cost on track for 1000x decline, enabling trillion-token-per-day per-user workloads
Current $5M/day for 1T tokens (OpenAI pricing) falls to ~$5k via hardware arbitrage, software efficiency, distributed infra, and scavenged power — making 'abundant intelligence' economicall…
AI Economics & Business Modelstailwindscore 8/10neil
Token cost dropping 3-6 orders of magnitude will unleash trillion-token daily demand per user
Driving token cost from millions to thousands/hundreds of dollars per trillion tokens will trigger Jevons paradox, making previously uneconomical verifiable tasks (coding, math, science) vi…
Semiconductorsheadwindscore 8/10neil
HBM capacity is the binding constraint on AI scaling; memory vendors' capex trauma creates structural shortage
Micron, SK Hynix, and Samsung refuse large fab investments after past cyclical burns, making HBM supply inelastic; this bottleneck will raise costs across the stack (including consumer devi…