Mechanistic interpretability moves from research to practice: Anthropic now understands ~3% of model internals
Amodei reveals Anthropic can identify discrete features representing complex concepts (bias, hesitation, metaphors) inside models, enabling detection of memorization vs reasoning and compliance with EU right-to-explanation rules — a practical breakthrough for enterprise trust.