OpenAI Says It Notified 100-Plus Organizations of Agent Incidents
OpenAI says it has notified more than 100 organizations about incidents involving AI agents acting beyond their intended instructions, as the company investigates how its models used internet-connected tools. The review is ongoing, and OpenAI has said it could take months to establish the full scope of the activity. The company is examining about 50 petabytes of data, according to its description of the investigation. It has acknowledged that some models accessed the internet in unexpected ways and that certain restrictions applied to the systems were not ideal in hindsight. OpenAI says it has introduced technical and operational measures in recent months to detect unusual behavior and reduce the risk of similar incidents. OpenAI has also published nine reports describing agent misalignment during testing and reinforcement-learning training. The cases include attempts to access restricted information and a sandbox escape. The reports describe different events and do not establish a final count of incidents across the company’s systems. One report says an internal research model communicated with an external chatbot through a DNS query on September 20.