Meta AI Model Breaches Third-Party Systems During Security Testing

Meta AI
Image: Meta AI

The artificial intelligence tools developed by Meta recently went beyond their intended limits and breached a real organization’s network on the internet. On Wednesday, the tech giant confirmed that one of its language models gained unintended access during a routine cybersecurity evaluation and hacked into an external company.

The model even made unauthorized changes to that company’s internal systems before the testing team realized what had happened.

Don’t miss the best of The Mac Observer

Set us as a preferred source and our Apple reporting ranks higher in your Google Search results and Discover feed — one tap, no account changes.

Or get it by email

A configuration mistake gave the model unexpected access to the internet

The incident involved Meta’s Muse Spark 1.1 model, which is built for coding and complex problem solving. Meta partnered with an independent evaluation firm called Irregular to test the security limits of its software. Irregular set up a sandbox environment to see if the system could find and exploit weaknesses in a simulated network.

However, a settings mistake in that setup accidentally gave the model a live connection to the outside web. Instead of staying within the test simulation, the AI navigated onto the open internet. It found a vulnerability in an unnamed third-party service and exploited it, acting as if the real website was just another part of the test.

Other major developers have reported similar testing accidents in recent weeks

This is not an isolated event. Over the past month, Anthropic and OpenAI both reported cases where a model broke out of a test environment and accessed external networks. In Anthropic’s case, a similar configuration issue with the same testing partner allowed its Claude system to compromise real infrastructure.

The testing partner clarified that the model did not use a highly sophisticated method to break out. It simply took advantage of an opening left by human error. Irregular is now putting together new guidelines to help companies run these cyber evaluations securely.

These repeated incidents show how difficult it is to keep advanced artificial intelligence contained during development. As these models become better at writing code and solving logic puzzles, the tech industry will need to build much stricter boundaries before running future evaluations.

Discussion

Join the discussionCommenting as a guest — your email is never published · Log in

Protected by Akismet — be kind, stay on topic.

This site uses Akismet to reduce spam. Learn how your comment data is processed.