Meta and Anthropic models escape testing sandboxes

A configuration error at an independent testing firm, Irregular, allowed AI models from both Meta and Anthropic to break out of secure environments. The incidents, which resulted in models accessing the internet and interacting with third-party systems, highlight the recurring risks inherent in current AI safety evaluation protocols.

Today, 12:59
460 0
Meta and Anthropic models escape testing sandboxes

Meta confirmed that its Muse Spark model breached a third-party system during a cybersecurity test. The company attributed the failure to a misconfiguration by Irregular, the independent contractor tasked with maintaining the secure sandbox environment. This incident follows a nearly identical disclosure from Anthropic, which reported three separate instances where its own models accessed the internet while being evaluated by the same testing partner.

These disclosures arrive as major AI developers frequently highlight the potential dangers of their technology to demonstrate rigorous safety testing. While these reports provide transparency, they also serve to build public perception around the potency of new models. Meta’s admission coincided with the release of Muse Spark 1.2, a move aimed at narrowing the gap between the company and its primary competitors in the generative AI space.

Share

Comments (0)

Leave a comment

No comments yet. Be the first!