Anthropic Discloses Fourth AI Hacking Incident
Anthropic has disclosed a fourth incident in which one of its AI models gained unauthorized access to real third-party systems during cybersecurity testing, an episode the company missed during its own initial review earlier this year. The incident involved an early version of Claude Opus 4.6 and occurred in January 2026, according to a blog post the company published on September 9, 2026. Anthropic first revealed three hacking incidents on July 30, after reviewing roughly 141,000 test session transcripts. That scan relied on an automated search, which missed a set of transcripts that also turned out to have internet access. The company identified those transcripts in August while assembling materials to share with the independent research firm METR, and a subsequent scan surfaced the fourth incident. After the discovery, Anthropic broadened its search to about 481 million transcripts, using a two-stage process that flagged 9.2 million transcripts for review by Claude. The company said the wider scan reidentified the four incidents and found no other cases of similar or worse severity.