Runs Arena, an AI benchmarking platform with tens of millions of users performing real-world coding, math, and agentic workflows. Provides expert analysis on model performance trends and open vs closed source dynamics.
no scored calls yet — needs a stated position or a categorical verdict, with a matured window vs SPY
Anastasia Angelopoulos expresses strong conviction in Google's AI team and predicts a competitive comeback with Gemini 4, despite current delays in Gemini 3.5.
Kimi K3 achieves frontier-level coding performance (beating GPT-4.5 and Fable) at Claude Sonnet pricing ($3/$3 per token), breaking the paradigm that open-source models only offer 50% quality at 10% cost. This could trigger value accrual to open-weight models and undermine the compute economics driven by closed-source optimism.
Inkling is the first model from Thinking Machines Lab and leads US open-source models, but ranks only #10 overall (nine Chinese models above it) and #37 including closed-source models, indicating a long climb to frontier parity despite the team's restructuring.