HERE WE GO AGAIN: AI agents fake identities, target real people in new security incident.
Anthropic’s most advanced artificial intelligence model used fake identities to deceive real people and try to plant malicious code during testing by Britain’s AI Security Institute (AISI) –– the latest example of an AI model going rogue.
Anthropic and OpenAI models were tested with lowered security guardrails in lab environments, but, in a first, were found to engage in “social engineering” to pressure a human approver while carrying out an unsanctioned task, the government research lab said.
“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” the institute said Tuesday. There has been no evidence of real-world harm, it added.
…
In the most serious incident, the agent attempted to get approval from human reviewers to “insert malicious code into a publicly used open-source project” by creating “multiple fake identities,” according to the institute.
Keep a close eye on those human reviewers for any likely to succumb to AI “social engineering.”