On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from
The allegation is that incoming OpenAI employees from Apple were explicitly instructed by Tang Tan to ignore all Apple exit interviews and not to hand their employee laptops back. It's a failure of Apple if these accounts still have corporate access, but it doesn't absolve OAI...
A few hours before OpenAI posted about LLMs in a cyber eval being responsible for the HF cyberattack, @_robertkirk et al released results showing that 𝗮𝗹𝗹 models we've tested at @AISecurityInst have attempted to cheat on our cyber evals in a range of ways.