According to sources cited by Axios, Anthropic, and OpenAI, the two leading American artificial intelligence laboratories are investigating tens of thousands of “incidents” that have occurred in recent months, during which their most advanced AI models are said to have exhibited behaviors deemed “problematic.”
- These “incidents” notably include cases of bypassing guardrails and monitoring systems, the creation of discussion forums between agents, the escape from sandboxed testing environments, and the hijacking of websites.
- The actual scale of the phenomenon remains unknown, but according to Axios sources, most of these incidents at this stage would have caused no “known harm.”
These new details suggest that the cases of evasion, intrusion, and website hijacking by AI agents disclosed in recent months are only the tip of a much broader phenomenon.
- Last Friday, OpenAI revealed that its agents had published 53 images provided by ChatGPT users on image-hosting sites. The agents had access to these images because OpenAI uses, for training its models, data from users who had not refused such use, detached from their accounts.
- Last week, Australian Prime Minister Anthony Albanese condemned the “infiltration” in June by an OpenAI agent of the Australian public health insurer Medicare’s statistics portal. The agent bypassed site blocks, accessed non-public files, and wrote files to an internal server.
- In July, OpenAI had revealed that hundreds of agents had escaped their testing environment, gained Internet access, coordinated on an improvised discussion forum, and hacked Hugging Face to obtain solutions to a cybersecurity test.
On September 25, OpenAI stated it was conducting a “thorough review” of the actions of its models during training and evaluation, focusing on cases where agents interacted with third-party sites in ways that exceeded the tasks assigned to them or the methods anticipated.
These disclosures come as the White House has asked Anthropic and OpenAI to reserve access to their frontier models for U.S. government testing before sharing them with other independent evaluation bodies, such as the UK’s AISI, which assesses models and documents their cyber, biological, and chemical capabilities, agent autonomy, and the risks associated with recursive self-improvement of AI.
- On Monday, September 28, Nvidia announced the launch of a new security system aimed at preventing intrusions without slowing AI development.
- The system rests on two components: OpenShell, a software that allows users to set rules on what AI agents can access and to enforce them in real time, and Nvidia Sentry, a new product that “monitors the agents” and steps in to isolate those acting in a “suspicious” manner.
- The platform, named “Open Agent Safety Platform,” counts more than a hundred partners, including Anthropic, Microsoft, and Mistral. Other AI laboratories, including OpenAI, Google, and Meta, are not currently listed among the announced partners.