TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
←
▶ 7:25 · $DEEPSEEK · Dwarkesh Patel: DeepSeek v3 requires 2400+ batch size for inference efficiency
episode briefing
Dwarkesh Patel

8 Predictions for the Era of Continual Learning

2026-08-07 · 2 company · 9 thematic
sentiment
0 bull0 bear2 neu
speakers
dwarkesh patel

Host of the Dwarkesh Podcast, known for deeply researched long-form conversations on artificial intelligence, science, economics and history.

quote
“Back-of-the-envelope math suggests that the optimal inference batch size for a sparse model like, say, DeepSeek v3 is more than 2400 concurrent sequences being generated at once. If you don't do this, then you're underutilizing your comput…”
—dwarkesh patel
now playing · $DEEPSEEK
$ANTHROPICneutral· mediumdwarkesh patel
Dwarkesh Patel: Continual learning will eliminate internal deployment gaps like Anthropic's four-month Mythos delay
Continual learning will force AI labs to deploy their smartest models immediately because competitors who ship earlier will accumulate real-world experience and surpass them, elim…
$DEEPSEEKneutral· mediumdwarkesh patel
Dwarkesh Patel: DeepSeek v3 requires 2400+ batch size for inference efficiency
Sparse models like DeepSeek v3 require batch sizes exceeding 2400 concurrent sequences for compute efficiency, creating massive economies of scale that favor large organizations s…