TickerTain
TickerTain
NewsroomShortsPortfolioConvergence
NewsroomShortsPortfolioConvergence
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
$ANTHROPIC·$MA····$INTC····$BLUE-ORIGIN·$SPCX····$CRWV····$CRM····$MSFT····$NVDA····$ORCL····$CURSOR·$AAPL····$OPENAI·$AMZN····$UBER····$GOOGL····$META····$TSLA····$DATABRICKS·$PERPLEXITY·$LYFT····$NBIS····$TSM····$LITE····$ANDURIL·
←

gabe pereyra

T3 · host / generalist

Co-founder and president of Harvey, an AI platform for legal and professional services. Previously a research scientist at DeepMind and a machine-learning engineer at Meta.

23 calls·13 names·83% bull·last heard 2 months ago·Sequoia Capital+1
track record

no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY

top calls

highest conviction · one per company
1sthigh conviction
$HARVEYHarveyposition

Harvey reaches roughly $300M ARR as product velocity compounds

Harvey is about four years old, near 900-960 employees, around 2,000 customers and roughly $300M ARR. The key signal is that product build cycles are compressing while the company expands into a broader legal AI platform.

Sourcery VC2026-06episode →
2ndhigh conviction
$CODACoda

Harvey differentiates from Coda by focusing on organizational vs individual productivity in legal vertical

Coda's horizontal, individual-focused productivity tools (Coda Work) cannot address the organizational workflow needs of large law firms — coordinating 20-30 person teams across 6-month projects, resource allocation across thousands of matters, and multi-firm enterprise coordination — which is Harvey's vertical moat.

Sequoia Capital2026-08episode →
3rdhigh conviction
$FIREWORKS-AIFireworks AI

Harvey partners with Fireworks for post-training and open-source model serving infrastructure

Fireworks provides critical infrastructure for post-training open-source models and serving them in production, enabling Harvey to train models like GLM 5.1 with advisor models and integrate open-source models into their serving stack alongside closed-source providers.

Sequoia Capital2026-08episode →

most discussed · click a bar to filter

  • $HARVEY
  • $CURSOR
  • $CODA
  • $SNORKEL
  • $LANGCHAIN

recurring themes

  • AI Infrastructure2
  • Open Source AI2
  • Enterprise AI Adoption2
  • AI Applications1
  • Synthetic Data & AI Training1
23 total
$CURSOR
Cursor
MEDgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey draws product inspiration from Cursor's Composer for building integrated post-trained models
Cursor's Composer product demonstrates the end-state Harvey aims for: packaging synthetic data, Neo lab partnerships, and post-training into a served model that runs alongside closed-source models in production.
"And inspired by Cursor, the goal of these efforts is for us to build our version of Composer one. How do we package all of the work we've done with synthetic data, scaling it with…"
8:29
$N-GRAM
N-gram
LOWgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey collaborates with N-gram on enterprise search and firm knowledge management
N-gram is working with Harvey on enterprise search and firm knowledge capabilities, part of the multi-lab partnership strategy to build domain-specific AI infrastructure.
"N gram, who I think is here, we're doing interesting work on enterprise search and firm knowledge."
7:01
$CODA
Coda
HIGHgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey differentiates from Coda by focusing on organizational vs individual productivity in legal vertical
Coda's horizontal, individual-focused productivity tools (Coda Work) cannot address the organizational workflow needs of large law firms — coordinating 20-30 person teams across 6-month projects, resource allocation across thousands of matters, and multi-firm enterprise coordination — which is Harvey's vertical moat.
"I think the big shift that we're thinking about is like our original product and then things like Coda's Coda Work are very individual-focused products, and so they're about indiv…"
27:26
$TRAJECTORY
Trajectory
LOWgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey partners with Trajectory to train NeMo-Megatron models for legal AI
Trajectory provides post-training expertise and infrastructure for training NVIDIA's NeMo-Megatron open-source models, contributing to Harvey's strategy of leveraging multiple Neo labs for different model architectures.
"Trajectory, um we worked with them to train NeMo-Megatron models."
7:01
$APPLIED-COMPUTE
Applied Compute
LOWgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey works with Applied Compute on Vault product for secure model deployment
Applied Compute is collaborating with Harvey on the Vault product, likely addressing secure deployment of post-trained models for enterprise customers with sensitive data requirements.
"Applied Compute, we're doing some interesting work on our Vault product."
7:01
$SNORKEL
Snorkel AI
MEDgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey leverages Snorkel AI for scaling synthetic data generation in legal domain
Snorkel's platform helps Harvey scale the synthetic data generation process initiated by domain experts, building larger training datasets for post-training legal AI models.
"And so, we work with companies like Mercor and Snorkel, who let you scale up this process and build larger sets particularly for training."
4:23
$LANGCHAIN
LangChain
MEDgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey works with LangChain to build efficient RL evaluation environments for large datasets
LangChain's tooling helps Harvey make expensive RL evaluation environments efficient, critical for running thousands of LLM-as-judge unit tests on large diligence datasets during post-training.
"And so, there's a lot of work, and here's some we did with LangChain, of making these very efficient."
5:01
$FIREWORKS-AI
Fireworks AI
HIGHgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey partners with Fireworks for post-training and open-source model serving infrastructure
Fireworks provides critical infrastructure for post-training open-source models and serving them in production, enabling Harvey to train models like GLM 5.1 with advisor models and integrate open-source models into their serving stack alongside closed-source providers.
"Fireworks, we got some very interesting results training GLM 5.1 to use Fable or maybe Opus 4.8 as an advisor model. Um Baseten, we did some interesting work on KB compaction. N g…"
7:01
$MERCOR
Mercor
HIGHgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey uses Mercor to scale synthetic data generation for model training
Mercor enables Harvey to scale domain-expert-guided synthetic data generation, solving the data privacy bottleneck for vertical AI by allowing lawyers to create realistic training datasets without using confidential client data.
"And so, we work with companies like Mercor and Snorkel, who let you scale up this process and build larger sets particularly for training."
4:23
$HARVEY
Harvey
HIGHgabe pereyra·Sequoia Capital·2 months ago·Building Frontier AI at the Application Layer: Harvey's Playbook | Gabe Pereyra
Harvey demonstrates application-layer playbook: synthetic data, Neo labs, and model serving to compete with frontier labs
Harvey's playbook proves application-layer companies can achieve frontier-level intelligence on specific domains by combining domain-expert synthetic data (via Mercor/Snorkel), multi-lab post-training partnerships (Fireworks, Baseten, Trajectory), and production-grade model serving infrastructure with routing/fallbacks — all without frontier-lab-scale capital.
"This talk is going to be our high-level playbook for doing this. And I'm going to talk about how we build benchmarks and training data, how we work with the Neo Labs to do post tr…"
1:07
9
AI Applicationstailwind
Application-layer AI companies can rival frontier labs by orchestrating the frontier ecosystem
Gabe argues that vertical application companies no longer need to build everything in-house; by leveraging open-source base models, NeMo lab partnerships, synthetic data pipelines, and model-serving infrastructure, they can achieve frontier-level performance on domain-specific tasks at a fraction of the cost.
8
Synthetic Data & AI Trainingtailwind
Domain-expert-guided synthetic data solves the vertical AI data privacy bottleneck
Vertical AI companies blocked from training on sensitive customer data can use domain experts (lawyers, doctors) to guide synthetic data generation via coding models, then scale with platforms like Mercor and Snorkel. This creates realistic, privacy-safe training datasets that enable post-training without data access.
8
AI Infrastructuretailwind
Domain-expert-guided synthetic data generation unlocks training for data-sensitive verticals
In domains like legal where customer data is privileged, using domain experts to guide synthetic data creation (via tools like Mercor, Snorkel) solves the cold-start problem and enables rigorous benchmarking and RL training without touching private data.
8
Open Source AItailwind
Post-training open-source models has become viable and cost-effective for domain-specific frontier intelligence
With strong open-source base models (GLM, Kimi, NeMo-Megatron) and accessible post-training APIs (Fireworks, Baseten, Tinker), application companies can now build specialized models that compete with closed-source frontier models on targeted tasks, changing the economics of vertical AI.
8
AI Infrastructuretailwind
Production model-serving infrastructure (routing, fallbacks, evals) is a prerequisite for effective post-training
Before investing in post-training, companies must build robust model-serving infrastructure including multi-model routing, automated/human evaluation gates, A/B testing, and production monitoring—this infrastructure works for both closed and open-source models and de-risks post-training investments.
8
Enterprise AI Adoptiontailwind
Vertical AI wins by going hyper-vertical into organizational workflows, not individual productivity
Horizontal tools (Coda, Notion) optimize individual productivity; vertical AI wins by orchestrating organizational workflows — multi-person, multi-month projects, resource allocation across thousands of matters, and multi-firm enterprise coordination — creating deep domain moats that horizontal platforms cannot replicate.
7
Open Source AItailwind
Open-source base models now competitive enough for domain-specific post-training to reach frontier performance
Models like Kimi 3, GLM 5.2, NeMo-Megatron, and Inkling have reached sufficient base capability that post-training on domain-specific data can yield frontier-level performance on narrow tasks (legal diligence, contracting), making open-source a viable alternative to closed-source APIs for vertical AI.
7
AI Economics & Business Modelsmixed
Continual learning on private customer data—without training on it—is the endgame for vertical AI
The ultimate product is not a single best model but enabling each enterprise to continuously customize models on their private work streams while preserving data privacy, requiring breakthroughs in continual learning, data isolation, and operational tooling.