
Alongside Nvidia, many of the biggest companies signed their names, including Amazon, Microsoft, and Meta. OpenAI and Google signed after the letter’s initial publication. (A notable absence was Anthropic.)
The background to all of this maneuvering was the unprecedented news from last week: that OpenAI models, undergoing internal testing, broke out of an offline “sandbox” inside OpenAI, accessed the internet, and used a never-before-seen cyber exploit to break into the AI repository Hugging Face—all without OpenAI employees’ direction, oversight, or, for several days, even awareness.
It was the kind of “warning shot” that AI safety advocates have long worried about: a rogue AI escaping its testing environment and causing real-world damage. Many saw it as a harbinger of worse hacks to come—especially when open-source AI models, which are widely seen as three to six months behind the frontier “closed” OpenAI models that carried out the attack, catch up to today’s level of capabilities. Open-source models are seen as especially worrisome by AI safety advocates because their guardrails can sometimes be stripped away. And because after they are released for free download on the internet, it is almost impossible to trace or destroy every copy of models that are found to be dangerous.