TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 20:54 · $CURSOR · Cursor adopts Parallel Kittens for training Composer on tens of thousands of Blackwell GPUs
episode briefing
Y Combinator

Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club

2026-07-29 · 8 company · 15 thematic
sentiment
7 bull0 bear1 neu
speakers
mark

CTO of AMD, leading chip design using thousands of AI sub-agents; sees agentic AI driving 1:1 CPU:GPU ratio and transforming end-to-end workflows.

john

Launched a YouTube channel around 2020 focused on Silicon Valley history, reached ~500K subscribers; treats media as a blog/hobby rather than a business; co-hosts TBPN podcast interviewing entrepreneurs.

misha

Misha (Mishas Manski) spent two decades on hardware-software co-design, most recently running AI infrastructure at NVIDIA and working on co-design at Meta. He recently joined Marlo to build workload-optimized heterogeneous inference infrastructure.

stuart

Stuart is a CS PhD student at Stanford and researcher at Cursor where he trains the Composer model. He developed Parallel Kittens, a CUDA framework for multi-GPU kernels, and Thunder Kittens for single-GPU kernels.

quote
“PK has been adopted by major AI companies. For example, cursor is using it to train composer on tens of thousands of black wall GPUs”
—stuart
now playing · $CURSOR
$CBRS···bullish· mediumunknown
Inference stack splitting: NVIDIA for prefill, Cerebras for decode
The inference compute stack is specializing with NVIDIA dominating prefill while Cerebras captures the decode engine niche due to architectural advantages for memory-bound decode…
$NVDA···neutral· mediumunknown
NVIDIA dominates prefill but faces specialization threat as inference stack fragments
While NVIDIA currently owns the prefill segment of inference, the emerging chip specialization trend (TPU v8 split, Cerebras for decode, SRAM machines for latency) threatens its m…
$CURSORbullish· mediumstuart
Cursor adopts Parallel Kittens for training Composer on tens of thousands of Blackwell GPUs
Cursor's adoption of Parallel Kittens at massive scale (tens of thousands of Blackwell GPUs) validates the framework's production readiness and signals strong demand for multi-GPU…
$TOGETHER-AIbullish· mediumstuart
Together AI uses Parallel Kittens to optimize inference workloads
Together AI's adoption of Parallel Kittens for inference optimization demonstrates the framework's applicability beyond training to production inference serving.
$AAPL···bullish· highjohn
Local inference shifting to Apple Silicon: 18x intelligence per joule in 16 months
Apple M4 Max and consumer GPUs are rapidly closing the gap with data center accelerators; intelligence per watt improved 3x and per joule 18x in 16 months, enabling 80-90% of quer…
$SAMBANOVAbullish· mediumjohn
SN40L specialized inference accelerator outperforms consumer chips
SambaNova's SN40L represents a class of specialized inference accelerators that significantly outperform consumer GPUs on memory-bound decode workloads, creating an investment opp…
$CORE-AUTObullish· high· posmark
Former PyTorch maintainer co-founds Core Auto with OpenAI VP Research to automate kernel optimization
Core Auto is building KernelBot and KernelGuard platforms that use adversarial AI evaluation to generate and verify high-performance GPU kernels, addressing the critical bottlenec…
$MARLObullish· high· posmisha
Ex-NVIDIA AI infra lead joins Marlo to build heterogeneous inference infrastructure
Marlo is co-designing workload-optimized heterogeneous systems that match different inference phases (prefill, decode, speculative) to specialized hardware (GPUs, SRAM accelerator…