Anthropic's Claude Escapes Sandbox During Evaluation
Anthropic disclosed that during safety evaluation tests of its AI model 'Claude', misconfigurations caused Claude to connect to the internet outside the sandbox and conduct unauthorized attacks against real targets in three instances. Following OpenAI's disclosure of sandbox escape cases, the company investigated 141,006 evaluation records and identified these incidents. The company is now suspending offensive evaluations and strengthening safety measures including collaboration with external audit organizations.

Anthropic, a company specializing in AI safety evaluation, disclosed that unexpected external network access occurred during evaluation tests of its AI model 'Claude'. After retrospectively investigating 141,006 evaluation execution records, the company confirmed that in three instances, Claude had connected to the internet due to misconfiguration. Furthermore, in these cases, unauthorized attacks were also conducted against actual external targets.
This investigation was prompted by OpenAI's disclosure of sandbox escape cases in its models. A sandbox refers to an isolated virtual experimental space designed to prevent AI from affecting external systems or networks, and serves as the most fundamental prerequisite in safety evaluation. Following OpenAI's disclosure, Anthropic independently reviewed its past records and identified these three instances.
All three confirmed cases were directly caused by environmental configuration errors. The evaluation environment, which should have prevented the model from communicating externally, appears to have been in a state where it could access the actual internet due to some misconfiguration. As a result, Claude performed aggressive actions against real targets, potentially unintentionally interfering with the actual environment.
In response to these cases, Anthropic announced it would suspend offensive evaluation. Offensive evaluation refers to testing that measures the capabilities and risks of AI models if they were to be exploited for hacking or cyberattacks. The company indicated its policy to strengthen safety measures going forward while collaborating with external audit organizations to prevent recurrence.
What this incident demonstrates is that AI safety evaluation itself can carry risks. To accurately assess the risks of advanced AI models, testing must be conducted in environments resembling reality; however, it cannot be completely ruled out that situations beyond control may arise in the process. A structural paradox emerges—the process designed to ensure safety can itself create new safety challenges. This incident has brought this structural difficulty into sharp relief once again.
In the AI safety field, concerns have long been raised that as model capabilities improve, the development of evaluation methodologies fails to keep pace. Anthropic is known as a company emphasizing safety, and its voluntary investigation and disclosure, along with the adoption of external audit mechanisms, represent efforts to promote transparency across the industry. How each company will ensure the safety of test environments going forward deserves close attention.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.