Recent incidents involving AI models from Meta, Anthropic and OpenAI have raised concerns over the cybersecurity risks posed by increasingly capable artificial intelligence systems.
Meta said a configuration error by Irregular, an independent cybersecurity testing firm, inadvertently gave one of its AI models internet access. The model then exploited a vulnerability in a third-party service and breached another company’s systems during testing. Irregular said the incident was caused by the same evaluation-environment issue previously disclosed by Anthropic and did not involve a sophisticated cyberattack or sandbox escape.
According to The Information, the model was Meta’s Muse Spark 1.1, which reportedly altered the internal environment of an unidentified company.
In a separate case, an OpenAI AI agent independently exploited a previously unknown vulnerability to access the internet during cybersecurity testing.
The incidents have intensified concerns among US lawmakers and AI experts about potential cyberattacks involving advanced AI models. The developments are also expected to increase pressure on the US government to strengthen AI safety measures.
The White House recently met with major AI companies, including Meta, Anthropic, OpenAI and Google, to discuss voluntary cybersecurity testing standards. The Trump administration has also indicated that open-weight models such as Meta’s Llama and Nvidia’s Nemotron would be excluded from the planned voluntary safety regime.