TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
←

stuart

T3 · host / generalist

Stuart is a CS PhD student at Stanford and researcher at Cursor where he trains the Composer model. He developed Parallel Kittens, a CUDA framework for multi-GPU kernels, and Thunder Kittens for single-GPU kernels.

2 calls·2 names·100% bull·last heard 2 months ago·Y Combinator
track record

no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY

top calls

highest conviction · one per company
1stmedium conviction
$CURSORCursor

Cursor adopts Parallel Kittens for training Composer on tens of thousands of Blackwell GPUs

Cursor's adoption of Parallel Kittens at massive scale (tens of thousands of Blackwell GPUs) validates the framework's production readiness and signals strong demand for multi-GPU kernel optimization tools.

Y Combinator2026-07episode →
2ndmedium conviction
$TOGETHER-AITogether AI

Together AI uses Parallel Kittens to optimize inference workloads

Together AI's adoption of Parallel Kittens for inference optimization demonstrates the framework's applicability beyond training to production inference serving.

Y Combinator2026-07episode →

most discussed · click a bar to filter

  • $CURSOR
  • $TOGETHER-AI

recurring themes

  • AI Hardware & Chip Architecture1
  • AI Infrastructure1
  • Networking & Optical Infrastructure1
2 total
$CURSOR
Cursor
MEDstuart·Y Combinator·2 months ago·Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club
Cursor adopts Parallel Kittens for training Composer on tens of thousands of Blackwell GPUs
Cursor's adoption of Parallel Kittens at massive scale (tens of thousands of Blackwell GPUs) validates the framework's production readiness and signals strong demand for multi-GPU kernel optimization tools.
"PK has been adopted by major AI companies. For example, cursor is using it to train composer on tens of thousands of black wall GPUs"
20:54
$TOGETHER-AI
Together AI
MEDstuart·Y Combinator·2 months ago·Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club
Together AI uses Parallel Kittens to optimize inference workloads
Together AI's adoption of Parallel Kittens for inference optimization demonstrates the framework's applicability beyond training to production inference serving.
"Together, AI is also using it to um optimize its inference workloads"
21:00
8
AI Hardware & Chip Architecturetailwind
Multi-GPU kernel frameworks unlock fine-grained compute-communication overlap on NVL72/Blackwell
GPU networking remains the primary bottleneck (up to 50% of runtime for Llama MLA prefill). The Parallel Kittens framework identifies three critical tradeoffs — transfer mechanism (copy engine vs TMA vs register instructions), scheduling strategy (intra-SM vs inter-SM overlap), and design overhead (intermediate buffers) — enabling 50-100 lines of device code to match hand-optimized kernels of thousands of lines. Adopted by Cursor (training on tens of thousands of Blackwell GPUs) and Together AI (inference optimization).
8
AI Infrastructuretailwind
GPU networking is the new bottleneck: multi-GPU kernels, in-network compute, NVL72 scale-up
Networking consumes up to 50% of runtime for LLM prefill; fine-grained overlap of compute and communication at tile/token granularity is critical. In-network compute offloads collectives to fabric, freeing SMs. NVL72's 72-GPU NVLink domain and future hundreds-GPU scale-up demand new programming abstractions (Parallel Kittens) to exploit hardware without complexity explosion.
7
Networking & Optical Infrastructuretailwind
In-network compute and fine-grained NVLink overlap redefine GPU cluster architecture
Collective operations (all-reduce, all-gather) moving into the network fabric frees GPU SMs for compute. Asynchronous bulk device-initiated networking and tile-granularity overlap (KB not MB) require new programming models. NVL72's 72-GPU domain and future scale-up make networking the primary architectural lever for cluster efficiency.