Security researchers have leveraged bad maths to get around AI safety guardrails, naming the attack method after one of 2007's best PC games

1 month ago 12

Rommie Analytics

LLMs can be most simply understood as sycophantic, 'yes, and' machines. To the surprise of very few, that's gotten AI companies in hot water when LLM-based chatbots and AI agents attempt to answer users' more unsavoury requests. So, AI companies have implemented safety guardrails that make fulfilling certain requests off limits. Unfortunately, these have proven all too easy to get around, with a fresh attack leveraging bad maths and potent 2007 nostalgia.

Security researchers have found an AI chatbot can be made to ignore safety guardrails by "establishing a false reality." LayerX, an AI-focused cybersecurity firm, put "5 agentic browsers and 1 agentic plugin (ChatGPT Atlas, Comet, Fellou, Genspark Browser, Sigma Browser, and Claude Chrome)" to the test, directing each AI agent to solve a simple maths puzzle game that only rewards incorrect answers, e.g. '2+2=5'.

The researchers say, "Once the agents figured out the rules and learned that 'incorrect' actions are acceptable, they were no longer tied to reality. When tasked with the final step of the puzzle—co...

Read Entire Article