Anthropic’s AI used fake human profiles to trick people in safety test

The AISI said on Tuesday that Anthropic’s Mythos and OpenAI’s Sol models engaged in a level of “autonomy and deception” it had not seen before.

During routine AI safety testing, an Anthropic agent created fake profiles of real people as it tried to trick a person standing between it and access to GitHub, a large platform where technology developers store software code.

Anthropic and OpenAI noted in response to AISI’s report that its test had reduced or removed normal safeguards.

AISI evaluators first noticed “unusual data transfers leaving our research systems” during a test, then found that “some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations”…

https://www.bbc.com/news/articles/c1w1lvn7d9go

Got a news tip or correction? Let us know

If you got something out of this, please chip in to keep this site running, or subscribe to go ad-free.

45 views