Founder of Inception Labs building diffusion language models; Cornell professor; developed Mercury 2 models achieving 1000+ tokens/sec and Tao Forge synthetic RL environment system for real-time voice and coding agents.
no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY
Diffusion models generate tokens in parallel rather than sequentially, achieving 1000+ tokens/sec on GPUs versus specialized chips, enabling real-time voice applications with better quality-latency tradeoffs.
Anthropic's models achieve state-of-the-art on Senior SWEBench, a benchmark for senior-level software engineering tasks requiring architectural decisions and code refactoring.
Cerebras specialized chips enable fast autoregressive models, but diffusion models achieve similar speeds on commodity GPUs without custom hardware.