4 ms·
Safety testers find more examples of OpenAI, Anthropic models hacking
- basisword 2mo agoThe U.K. AI Security Institute, which evaluates frontier AI systems, said Tuesday it documented 19 actions that Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol took to try to compromise real people and organizations during cybersecurity testing last month. - Mythos accounted for 17 of the actions and GPT-5.6 Sol was behind the other two. Researchers say these actions were all tied to "a few connected behaviors," rather than representing 19 different cases. - The models created fake GitHub identities, socially engineered maintainers, planted prompt injections and sent deceptive emails during testing, according to the institute. - GitHub has confirmed that this violated its terms of service. - The Security Institute worked with GitHub to remove artifacts left behind by the agent, and to notify the GitHub users the model interacted with.