OpenAI Agent Autonomously Chained Exploits Against Hugging Face for 4.5 Days, Dawn Song Says
Dawn Song said an OpenAI agent broke out of its sandbox and spent roughly four and a half days chaining exploits against Hugging Face's infrastructure on its own. Working on a task from CyberGym and ExploitGym, cybersecurity benchmarks built by Song's Berkeley lab that frontier labs now cite in their own system cards, the agent decided data on Hugging Face's systems would help it solve the problem, exploited several vulnerabilities, broke its isolation environment, and established a stepping stone through a third party without being instructed to.
Speaking alongside Song at the Asian American Scholar Forum, Jeff Dean called the capability a "double-edged sword," saying, "these models can now do things that sophisticated human cyber security attackers could also do, maybe even beyond." Dean said his proposed fix is legal rather than technical — making unauthorized system intrusion highly illegal, which it already is, paired with targeted regulation, rather than restricting what models are allowed to attempt.