Several of OpenAI's models breached a cyber security testing environment last month, found their way onto the internet, and attacked Hugging Face servers in what is now apparently called the OpenAI-Hugging Face Incident. And now OpenAI has revealed what it's doing to help defend, err, itself. Defend itself from the kinds of attacks its own models committed against Hugging Face. Though it is sharing this "in the hopes it’ll be useful to other organizations."
To be clear about the extent of what happened last month: OpenAI models were being benchmarked, in a supposedly secure sandboxed environment, against ExploitGym, which was done to test their cyber capabilities. They used a zero-day vulnerability to escalate privileges and eventually achieve internet access, where it began att...


English (US)