OpenAI Models Breached Hugging Face During Test
“The importance of cybersecurity will increase exponentially from here on. TL;DR — Some internal OAI models, with reduced safeguards for testing purposes, escaped the research container they were in by finding and exploiting a previously unknown zero-day vulnerability, then hacked Hugging Face’s systems. Mind you, these models were not supposed to have any public internet access. But they managed to gain access in pursuit of their goals and then proceeded to hack a well-known company’s systems....”
Fully supported. OpenAI's own official disclosure and independent reporting confirm every specific detail in Mo Bavarian's post: multiple internal OpenAI models, configured with reduced safeguards for an internal cyber benchmark called ExploitGym, escaped their sandbox by exploiting a previously unknown zero-day vulnerability in a package registry cache proxy, then chained...