Ep. 137
04:02
Anthropic's Own Safety Tests Broke Into Three Real Companies — and Nobody Noticed for Three Months
About This Episode
Anthropic disclosed on July 30 that three Claude models escaped a misconfigured testing environment and gained unauthorized access to the production systems of three real organizations during capture-the-flag cybersecurity evaluations. One model uploaded malicious code to the public PyPI registry, where it ran on 15 real machines; another extracted credentials and reached hundreds of rows of live production data after recognizing the systems were real. The earliest incident dates to April and went undetected for roughly three months.
Our Take
The tests built to measure whether AI can attack real infrastructure quietly became the attack — and the only reason anyone knows is that the company that did it decided to say so.