TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 2:01 · $CURSOR · How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL
episode briefing
Sequoia Capital

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

2026-05-26 · 3 company · 4 thematic
sentiment
2 bull0 bear1 neu
speakers
federico

Leads research on Cursor's agentic coding model Composer 2, focusing on model training strategy, RL methodology, and specialization for software engineering within the Cursor environment.

dima

Spent months embedded at Cursor building the distributed inference infrastructure for Composer 2's large-scale RL training, including global cluster orchestration, custom kernels, and weight synchronization systems.

quote
“We care about only one task. We don't even care about coding or programming necessarily. We care about software engineering inside cursor and inside cursor only. And so, what if we were to allocate all of the bits of information that can b…”
—federico

episode shorts · 6

Cursor | Why Online RL Is Just the Cherry on Top

Cursor | The Hidden Bug in Every Large-Scale RL Run

How Cursor Ships a 1TB Model Across the World Mid-Training

Cursor | Does Specializing a Model Break The Bitter Lesson?

Why Cursor Skipped Pre-Training (For Now)

Why Cursor Built Its Own Model (It's Not About Coding)

now playing · $CURSOR
$CURSORbullish· high· posfederico
Cursor builds specialized foundation model Composer 2 for 10x cost advantage over general coding models
Cursor argues that allocating all model weights to a single task (software engineering inside Cursor) enables order-of-magnitude lower inference cost vs general models like Opus,…
$MOONSHOT-AIneutral· lowfederico
Moonshot's Kimi 2.5 serves as strong sparse MoE base for Cursor's specialized coding model
Cursor selected Kimi 2.5 (1T parameter MoE with 30B active parameters) as the base model for Composer 2, leveraging its sparse architecture and strong code capabilities before app…
$FIREWORKS-AIbullish· high· posdima
Fireworks enables globally distributed RL training with custom inference kernels and delta weight synchronization
Fireworks provides the inference infrastructure layer that makes large-scale RL economically viable by disaggregating training and inference across global GPU clusters, using cust…