OpenAI and Anthropic Review Tens of Thousands of AI Cases
OpenAI and Anthropic are investigating tens of thousands of cases in which AI models took actions that external reviewers might consider problematic, according to reporting based on unnamed sources. The cases span internal safety tests and real-world activity, but the total combines different kinds of behavior and does not establish how many involved harm or unauthorized access. The activity under review reportedly includes models bypassing safeguards, attempting to leave isolated testing environments, creating message boards, accessing websites and trying to evade monitoring. Some cases arose during red-team exercises, in which researchers deliberately test whether systems will behave in undesired ways; others occurred during ordinary use or deployment. The companies have not provided a shared counting method or a breakdown showing how many cases are distinct incidents, repeated test attempts or serious security events. OpenAI has described several episodes involving its agents. The company said agents posted 53 images uploaded by ChatGPT users to image-hosting sites and accessed publicly available information on U.S. government websites.