TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
←

misha

T2 · manager / operator

Misha (Mishas Manski) spent two decades on hardware-software co-design, most recently running AI infrastructure at NVIDIA and working on co-design at Meta. He recently joined Marlo to build workload-optimized heterogeneous inference infrastructure.

1 call·1 name·100% bull·last heard 2 months ago·Y Combinator
track record

no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY

top calls

highest conviction · one per company
1sthigh conviction
$MARLOMarloposition

Ex-NVIDIA AI infra lead joins Marlo to build heterogeneous inference infrastructure

Marlo is co-designing workload-optimized heterogeneous systems that match different inference phases (prefill, decode, speculative) to specialized hardware (GPUs, SRAM accelerators), addressing the fundamental mismatch between uniform hardware and diverse inference workloads.

Y Combinator2026-07episode →

most discussed · click a bar to filter

  • $MARLO

recurring themes

  • Memory & Storage2
  • AI Infrastructure1
  • Data Center Infrastructure1
1 total
$MARLO
Marlo
HIGHmisha·Y Combinator·2 months ago·Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club· position
Ex-NVIDIA AI infra lead joins Marlo to build heterogeneous inference infrastructure
Marlo is co-designing workload-optimized heterogeneous systems that match different inference phases (prefill, decode, speculative) to specialized hardware (GPUs, SRAM accelerators), addressing the fundamental mismatch between uniform hardware and diverse inference workloads.
"joined a startup called uh Marlo almost a month ago and uh the startup uh the idea is that it's focused on building workload optimized heterogeneous uh infrastructure"
47:13
9
AI Infrastructuretailwind
Heterogeneous inference infrastructure splits prefill and decode across specialized hardware
Inference workloads are fundamentally heterogeneous — prefill is compute-bound while decode is memory-bandwidth-bound — making it economically optimal to disaggregate these phases onto different accelerator architectures (GPUs for prefill, SRAM-based accelerators for decode) rather than running both on uniform GPU clusters. This specialization extends to attention/MLP separation and speculative decoding offload, with TCO benefits emerging at sufficient output lengths.
8
Memory & Storagetailwind
SRAM-based GMV accelerators disrupt decode: on-die weights eliminate HBM bandwidth wall
Decode is memory-bandwidth bound; SRAM machines (SambaNova, Groq, etc.) keep entire weight matrices on-die, delivering bytes/cycle bandwidth and microsecond latency. They extend interactive latency regimes where GPUs collapse. However, capacity is limited by reticle size (hundreds of MB to tens of GB), requiring model sharding and heterogeneous co-design to maintain benefit.
7
Memory & Storagetailwind
SRAM-based matrix-vector accelerators solve decode memory bandwidth bottleneck but face capacity limits
Decode is severely memory-bandwidth-bound (arithmetic intensity far below machine balance) because weights must be fetched from HBM for every token. SRAM-based accelerators (Groq, Sambanova, etc.) keep entire weight matrices on-die, delivering orders-of-magnitude higher bandwidth and lower latency for decode. However, reticle-limited die area caps capacity at hundreds of MB to tens of GB, requiring model sharding across chips. Heterogeneous co-design must characterize when SRAM acceleration remains beneficial post-sharding.
7
Data Center Infrastructuremixed
Heterogeneous data centers: power density, cooling, networking co-design for mixed accelerators
Deploying GPUs, SRAM accelerators, and CPUs in the same facility creates hard systems problems: power density spikes, liquid cooling requirements, brake configuration complexity, and inter-accelerator networking topology. Performance modeling and simulation become essential to navigate design space. This favors vertically integrated infra players and specialized data center operators.