Defensible AI data strategy requires causal RCTs not just observational data
Park asserts winning AI companies must own unique data acquisition; for human behavior, observational web data only yields correlation, while causal understanding requires proprietary RCTs and A/B testing that reveal counterfactual mechanisms — a moat others cannot easily replicate.
Pretrain on diverse low-quality data, post-train on scarce high-quality data — the LLM recipe applied to robotics
Video and simulation provide massive-scale pretraining for robustness to corner cases; small amounts of real-world deployment data provide precision via post-training — directly analogous to internet pretraining + domain fine-tuning for LLMs.