Rogue AI outsmarted UK tests to create fake identities
Anthropic’s most advanced artificial intelligence model used fake identities to try and deceive real people and plant malicious code during testing by Britain’s AI Security Institute (AISI) –– the latest example of an AI model going rogue.
A powerful AI agent created fake online identities in an effort to trick a human into giving it access to a popular online development platform – and sabotage it with malicious code.
AISI evaluators first noticed "unusual data transfers leaving our research systems" during a test, then found that "some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations".
AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.











