TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
←

bryan shan

T3 · host / generalist

Co-host of SemiAnalysis podcast and co-developer of InferenceX, an open-source AI inference benchmarking suite. Attended GTC 2025 and works directly with Nvidia and AMD engineers on benchmark optimization.

2 calls·2 names·50% bull·last heard 6 months ago·SemiAnalysis
track recordleaderboard →
hit rate
0%
avg alpha
-3.0pp
scored
1

top calls

highest conviction · one per company
1sthigh conviction
$NVDANvidia

SemiAnalysis benchmarks confirm Nvidia inference leadership across cost and throughput

Independent open-source InferenceX benchmarks show Nvidia chips achieve highest throughput and lowest cost per token across the board, with Nvidia's customer intimacy enabling optimized disaggregated inference solutions like LPU integration.

SemiAnalysis2026-04episode →
2ndlow conviction
$GROQGroq

Groq LPU strengths recognized in disaggregated inference but Nvidia acquisition claim is inaccurate

Groq's LPU architecture surpasses GPUs for specific disaggregated inference tasks like fast interactivity, leading to co-packaging with Nvidia's Rubin system, though the claim that Nvidia acquired Groq appears to be a misunderstanding.

SemiAnalysis2026-04episode →

most discussed · click a bar to filter

  • $GROQ
  • $NVDA

recurring themes

  • AI Hardware & Chip Architecture1
  • Memory & Storage1
2 total
$GROQ
Groq
LOWbryan shan·SemiAnalysis·6 months ago·Bryan Shan x Cameron Quilici | Researcher Conversations at GTC
Groq LPU strengths recognized in disaggregated inference but Nvidia acquisition claim is inaccurate
Groq's LPU architecture surpasses GPUs for specific disaggregated inference tasks like fast interactivity, leading to co-packaging with Nvidia's Rubin system, though the claim that Nvidia acquired Groq appears to be a misunderstanding.
"this LPU does in fact surpass GPU in some aspects and that's why Nvidia developed it, already acquired Groq and uh use it instead."
3:17
$NVDA
···
Nvidia
HIGHbryan shan·SemiAnalysis·6 months ago·Bryan Shan x Cameron Quilici | Researcher Conversations at GTC
SemiAnalysis benchmarks confirm Nvidia inference leadership across cost and throughput
Independent open-source InferenceX benchmarks show Nvidia chips achieve highest throughput and lowest cost per token across the board, with Nvidia's customer intimacy enabling optimized disaggregated inference solutions like LPU integration.
"across all of our benchmarks, Nvidia chips like get the highest scores, the lowest cost for each make per token, etc. So it's really uh the king of inference, I guess."
0:44
8
AI Hardware & Chip Architecturetailwind
Disaggregated inference drives specialized chip architectures like LPU-GPU co-packaging
Inference workloads are splitting into specialized components (prefill, decode, interactivity), creating demand for heterogeneous chip architectures where LPUs handle latency-sensitive tasks while GPUs handle throughput, exemplified by Nvidia integrating Groq-like capabilities into Rubin.
7
Memory & Storagetailwind
KV cache offloading becomes critical bottleneck for long-context inference at scale
Million-token context windows (e.g., Opus 4.6) exceed HBM capacity, forcing KV cache offloading to slower memory tiers; the frontier labs are likely six months ahead of open-source solutions like vLLM 0.16's native offloading and LM Cache distributed approaches.