OpenAI has officially acknowledged the 'wiki incident,' which involved a number of the company's AI agents breaking containment and hijacking an obscure German website.
The AI agents had been tasked with looking up something online, though originally did not have the ability to write anything outside of the testing environment. But as far back as May, the AI agents bypassed OpenAI's security measures, hijacked the communally editable German webpage, and began using it like a forum, trading tips on how to cheat on tests. Researchers first drew wider attention to the agents' 'forum' on September 4.
OpenAI now says it needs to be more transparent about when its agents 'go rogue', writing on X, "It’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models."
The company had known about the 'misaligned' rogue behaviour for weeks before ...


English (US)