TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
←

dima

T3 · host / generalist

Spent months embedded at Cursor building the distributed inference infrastructure for Composer 2's large-scale RL training, including global cluster orchestration, custom kernels, and weight synchronization systems.

1 call·1 name·100% bull·last heard 4 months ago·Sequoia Capital
track record

no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY

top calls

highest conviction · one per company
1sthigh conviction
$FIREWORKS-AIFireworks AIposition

Fireworks enables globally distributed RL training with custom inference kernels and delta weight synchronization

Fireworks provides the inference infrastructure layer that makes large-scale RL economically viable by disaggregating training and inference across global GPU clusters, using custom kernels (FP4, router replay for MoE) and delta weight compression to synchronize 1TB model snapshots across continents in under a minute.

Sequoia Capital2026-05episode →

most discussed · click a bar to filter

  • $FIREWORKS-AI

recurring themes

  • AI Infrastructure1
  • Semiconductors1
1 total
$FIREWORKS-AI
Fireworks AI
HIGHdima·Sequoia Capital·4 months ago·How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL· position
Fireworks enables globally distributed RL training with custom inference kernels and delta weight synchronization
Fireworks provides the inference infrastructure layer that makes large-scale RL economically viable by disaggregating training and inference across global GPU clusters, using custom kernels (FP4, router replay for MoE) and delta weight compression to synchronize 1TB model snapshots across continents in under a minute.
"We can globally distribute that across small clusters all over the world. So, I think for the composer to run, we used the four clusters in total that were all over the world, ver…"
16:43
9
AI Infrastructuretailwind
Distributed RL training across heterogeneous global clusters reduces need for massive contiguous GPU clusters
Fireworks' architecture disaggregates training (needs high-bandwidth interconnect) from inference (can run on smaller, heterogeneous, cheaper clusters worldwide), using delta weight synchronization and custom inference kernels to achieve near-contiguous-cluster efficiency at fraction of the cost and cluster availability constraints.
7
Semiconductorstailwind
Custom GPU kernels and numerical determinism are critical for MoE model RL training stability
MoE architectures amplify floating-point non-determinism during asynchronous RL (expert routing divergence), requiring kernel-level fixes like router replay and deterministic accumulation order to align inference and training passes — a systems-algorithm co-design challenge that favors vertically integrated infrastructure providers.