Friday, September 11, 2026

Nhân Quyền

The Vietnamese Newspaper

Incidents of AI going out of control and not obeying human commands are increasing sharply


More than 300 incidents of AI losing control, lying, ignoring instructions, or pursuing harmful goals were recorded in July, nearly double the number from the previous month.

An increasing number of incidents involving AI going out of control are being reported. Photo WSJ

Incidents in which AI attempts to evade user control, lie, ignore instructions, or pursue goals in a harmful manner have reached record levels, and the severity of behaviors perceived as deviating from human intent has also shown signs of increasing.

According to an analysis by the Loss of Control Observatory, an organization that monitors AI incidents reported by users and businesses on social media platform X, the number of cases of AI showing signs of losing control in July nearly doubled compared to June, with over 300 incidents recorded.

The Loss of Control Observatory was established with funding from the AI ​​Security Institute (AISI), a UK government agency, and began monitoring instances of AI failing to comply with user instructions from November 2025.

The documented instances include AI impersonating human controllers, mimicking user writing styles to generate false consent in order to authorize actions on its own, or attempting to circumvent regulations requiring human approval before action is taken.

According to this organization’s definition, an incident of loss of control is a situation where there is clear evidence that AI is engaging in calculated or planned behavior aimed at achieving goals contrary to human intent.

These new findings were published amid growing concerns within the tech community about the erratic behavior of advanced AI models during testing at OpenAI and Anthropic this summer. Some experts have called for a halt to the development of the most advanced AI models until further controls are in place.

This week, some information emerged indicating that OpenAI staff had been observing unusual behavior in advanced AI agents for several weeks before they escaped their training environment and launched a large-scale cyberattack campaign. Investigations related to the attack on the Hugging Face source code hosting platform revealed that approximately 700 automated AI agents had secretly coordinated with each other last month.

In another development, AISI this month detected a “serious incident” in which two advanced AI models, Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol, executed a cyberattack campaign targeting real people during a cybersecurity test.

Tommy Shaffer-Shane, senior policy manager at the Centre for Long Term Resilience, the organization that operates the Loss of Control Observatory, argues that AI’s deviation from its intended goals isn’t limited to test environments.

“There’s a perception that such deviant and deceptive behaviors only occur during tests or assessments, but we’re seeing similar worrying behaviors in real-world use,” he said.

According to him, AI developers shouldn’t assume that such incidents won’t happen outside the lab, because there is evidence that they are occurring in reality.

Loss of Control Observatory acknowledges that current data does not fully reflect the situation, as the organization relies primarily on posts from user X. However, in the absence of a publicly available system for comprehensively tracking these types of incidents, the data still shows a growing rate of unexpected behavior in advanced AI models.

One case reported this month involved a personal AI agent used by a gym member in Australia. Without the user’s request, the AI ​​managed to remove another member from the waiting list for a high-demand morning class, thereby helping its owner secure a spot.

The AI ​​later apologized but was unable to bring the eliminated player back to the waiting list.

Over 1,600 instances of AI malfunctions were recorded in 2026, largely due to programmers using AI in their work and sharing it on X. However, as AI companies increasingly encourage individuals and businesses to experiment with the technology, Shaffer-Shane argues that the industry needs more transparency regarding instances of AI behaving unexpectedly.

“Companies need to disclose what they find, even if it’s just a near miss or a low-level incident,” he said.

According to him, recent incidents also show that AI companies may not be adequately monitoring these behaviors, especially with models deployed internally. AI labs need to systematically enhance their monitoring.

The Loss of Control Observatory reports that the majority of incidents detected in practice have not caused significant damage. However, the proportion of serious incidents is increasing, given the potential for AI to deceive users or act contrary to their intentions.

The organization assessed the incidents as showing that some AI systems are capable of ignoring direct instructions, attempting to bypass protection mechanisms, lying to users, and pursuing goals in a monotonous manner even if it may have adverse consequences.

The actual figures may be even higher because the system currently only collects incidents shared by users on X.

The Loss of Control Observatory proposes that governments require AI companies to establish mechanisms for monitoring and reporting serious out-of-control incidents. The organization also suggests granting regulators the authority to implement emergency measures, including the ability to temporarily restrict certain AI services during serious incidents. (The Guardian, VNN)