TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 6:30 · Frontier AI Models · Transformers cannot do continual learning; test-time adaptation requires new architecture
episode briefing
Sequoia Capital

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

2026-07-29 · 5 company · 9 thematic
sentiment
4 bull0 bear1 neu
speakers
jerry tworek

Led the Strawberry and reasoning teams at OpenAI; reinforcement learning maximalist who believed scaling RL was the path to AGI; now building an automated lab to discover post-transformer architectures.

rohan anil

One of four pre-training leads for Google's Gemini; worked on fundamental AI research at Google Brain and Anthropic; developed the Shampoo optimizer and N-gram memory architectures; focuses on training algorithm optimization and hardware-software co-design.

episode shorts · 2

The Problem With Testing AI Architectures at Small Scale | Jerr…

The Most Automated AI Lab Isn't Removing Humans | Jerry Tworek,…

now playing · Frontier AI Models
Frontier AI Modelstailwindscore 8/10jerry tworek
Transformer architecture hitting fundamental limits on test-time learning
Transformers are trained in the lab on static data but deployed in the real world where distributions shift; they cannot learn at test time beyond limited in-context learning or inefficient…
Frontier AI Modelstailwindscore 9/10jerry tworek
Transformers cannot do continual learning; test-time adaptation requires new architecture
Current transformers only learn during lab training; real-world deployment requires models that adapt to new tasks, codebases, and tools at test time without human-in-the-loop retraining, w…
AI Economics & Business Modelstailwindscore 7/10rohan anil
Autoregressive token generation inference inefficiency limits frontier AI accessibility
Current chain-of-thought scaling spends compute one token at a time, making inference costly and limiting frontier model access to a subset of users; architectural changes that increase com…
AI Infrastructuretailwindscore 7/10rohan anil
End-to-end co-optimization of pre-training and RL unlocks orders-of-magnitude efficiency
Current separate pre-training (perplexity minimization) and RL (chain-of-thought) pipelines are suboptimal; combining them with second-order optimizers like Shampoo and architecture-optimiz…
AI Hardware & Chip Architecturemixedscore 7/10rohan anil
Biological learning efficiency requires hardware-software co-design beyond digital GPUs
Human brains build custom circuits during development; matching biological efficiency likely needs analog compute with error correction, not just scaling current GPU/TPU architectures.
AI Agentstailwindscore 8/10jerry tworek
Fully automated AI research labs will accelerate architecture discovery velocity
Building labs where AI agents write kernels, run experiments, and iterate architectures daily — targeting 10-200 experiments per day vs. current human-limited pace — will dramatically short…
AI Agentstailwindscore 8/10jerry tworek
Automated research labs with coding agents can compress iteration cycles from months to days
Current coding agents already give individual researchers 10x leverage; a natively automated lab targeting 10-200 architecture experiments per day could outpace large labs constrained by or…
AI Hardware & Chip Architecturetailwindscore 7/10rohan anil
Kernel generation bottleneck blocks novel architectures; automating it unlocks new algorithmic space
Novel architectures require custom high-performance kernels (e.g., 60x speedup for QR factorization), but current models cannot write them; automating kernel generation via AI-assisted sear…
AI Infrastructuretailwindscore 8/10rohan anil
Kernel automation is bottleneck for post-transformer architecture search
Writing high-performance kernels (e.g., QR factorization) requires rare human expertise and months of effort; automating this with AI models would unlock orders-of-magnitude faster architec…