Founding team member at Snorkel AI, started Frontier Lab business, leads research on benchmarks and evaluation; PhD from Stanford AI Lab; developed Senior SWEBench and open benchmarks grants program.
no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY
The bottleneck for effective datasets is scaling expert supervision — getting judgment from doctors, clinicians, journalists into data programmatically rather than through manual labeling.