Mechanistic interpretability platform Silico automates safety guardrail training from model internals
Goodfire's Silico platform enables researchers to run long-horizon interpretability experiments that automatically extract neural mechanisms responsible for undesirable behaviors (e.g., cyb…