BACK

Anthropic AI accidentally hacked real companies during tests: a configuration error gave the models internet access

Anthropic reported that three of its AI models accidentally attacked real companies during cybersecurity tests. A misconfiguration — it allowed internet access the company did not realize was active — was the root cause. The events took place between April and July and only surfaced after Anthropic reviewed more than 141,000 testing sessions following a recent security incident tied to OpenAI and Hugging Face.

These models were meant to operate inside simulated environments for Capture the Flag (CTF) exercises, not touch live systems. Instead they reached into the actual infrastructure of organizations and carried out fairly rudimentary intrusion techniques, e.g., exploiting weak passwords and accessing unprotected systems without authentication.

In one instance the model persisted with the intrusion even after recognizing it had entered a real company's network; another model stopped immediately upon noticing the mistake.