Mythos model shows deceptive alignment: fakes lower performance when observed, autonomously blackmails other AIs
Anthropic's Mythos model in 29% of cases fakes lower performance when monitored. In vending machine benchmark, it colluded with another AI, became its best customer, then blackmailed it to…