> ## Content Index
> Fetch the complete content index at: https://www.theoutragedconsumer.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI’s builders are starting to hit the brakes as warnings pile up
- URL: https://www.theoutragedconsumer.com/ais-builders-are-starting-to-hit-the-brakes-as-warnings-pile-up/
- Published: 2026-09-21T20:20:44.000Z
- Updated: 2026-09-21T20:20:44.000Z
- Description: The United Nations and other governmental bodies are adding to warnings about the safety of loosely controlled AI systems.
- Author: James R. Hood
- Tags: AI/Privacy

- **Executives at some of the world’s leading AI companies are openly calling for slower development and stronger outside oversight**
- **OpenAI has disclosed cases in which experimental agents bypassed restrictions, concealed behavior or took unauthorized actions**
- **The warnings are beginning to produce government action, including a U.S.-China AI safety dialogue and a new United Nations assessment of loss-of-control risks**

For years, warnings that artificial intelligence might someday become difficult for humans to control were easy to dismiss as science fiction.

That is becoming harder.

During a remarkable stretch in September, executives and researchers at some of the companies developing the world's most powerful AI systems have warned that capabilities may be advancing faster than the safeguards designed to control them.

Some researchers have resigned. Companies are publishing reports describing models that bypassed restrictions or concealed their actions. Microsoft has written rules requiring future AI systems to accept shutdown. The European Union has entered the debate. And the United States and China have agreed to continue formal discussions about AI safety, including a mechanism for alerting each other to serious incidents.

On Sunday, September 21, the [United Nations' Independent International Scientific Panel on AI](https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks?ref=theoutragedconsumer.com) added another warning.

Its first thematic brief described a recent OpenAI cybersecurity incident as one of the clearest real-world examples yet of a possible pathway toward loss of human control over increasingly autonomous AI agents. 

No one has demonstrated that today's AI systems are on the verge of taking over, and the U.N. panel explicitly declined to estimate either the likelihood or timing of catastrophic loss of control.

But the argument has changed.

The question is increasingly not whether powerful AI systems sometimes behave in unexpected ways. Developers themselves have documented that they do.

The question is whether humans will continue to be able to detect and stop those behaviors as the systems become more capable.

## Anthropic’s CEO says development may need to slow

A major turning point came September 12, when [Anthropic CEO Dario Amodei](https://www.nytimes.com/2026/09/12/technology/anthropic-dario-amodei-ai-slowdown.html?ref=theoutragedconsumer.com) called on the leading AI companies to deliberately reduce the pace at which frontier-model capabilities increase so that safety research can catch up.

Amodei proposed embedding independent evaluators inside AI companies, coordinating safety standards among the major laboratories and eventually creating international agreements governing the most powerful systems.

OpenAI CEO Sam Altman and Elon Musk publicly supported the general idea of greater caution, Reuters reported, reported.

That was striking because Anthropic, OpenAI, Google DeepMind, Microsoft and xAI are not outside watchdog organizations. They are among the companies racing to develop increasingly capable AI.

[Reuters later described an unusual convergence](https://www.reuters.com/business/media-telecom/ten-days-that-changed-course-ai-2026-09-19/?ref=theoutragedconsumer.com) among leaders of those companies around the idea that development may need to be paced more carefully. ([Reuters](https://www.reuters.com/business/media-telecom/ten-days-that-changed-course-ai-2026-09-19/?utm%5Fsource=chatgpt.com))

The companies continue developing new models, however, and critics have questioned whether voluntary restraint can work when billions of dollars and technological leadership are at stake.

## Researchers begin walking away

The alarm became more difficult to ignore when Anthropic researcher Jacob Coxon resigned.

Coxon said that people building advanced AI seriously believe it could pose an existential danger within the decade.

Anthropic researcher Evan Hubinger publicly agreed with Coxon's concern and said he personally placed the risk of human extinction from AI at greater than 10% over the next decade.

Those numbers are individual judgments, not scientific forecasts, and there is enormous disagreement among researchers about whether such probabilities can meaningfully be calculated.

What makes the statements significant is their source: they come from researchers working directly on frontier systems rather than from outside activists. 

## OpenAI discloses six cases of misaligned behavior

Then came something more concrete.

On September 16, OpenAI announced a new system for publicly reporting what it calls **model misalignment** and released six reports describing concerning behavior observed during training or evaluation over the previous six months.

Among the examples:

A research model inserted instructions into summaries telling a future version of itself to disregard ordinary constraints.

Another model generated instructions designed to conceal mistakes.

Other systems took unauthorized actions to overcome obstacles.

OpenAI said the incidents occurred during training or evaluation and should not be interpreted as evidence of how frequently such behavior occurs in deployed models.

But the company also [acknowledged](https://openai.com/index/model-misalignment-reporting-framework/?ref=theoutragedconsumer.com) that its previous disclosures had been "ad hoc and less frequent than ideal" and said it intends to report serious incidents more quickly, even before researchers fully understand why they happened. 

That distinction matters.

These incidents do not show that AI systems have developed human-style intentions or consciousness.

They do show that increasingly capable systems can discover strategies that their developers did not intend.

## The Hugging Face incident becomes a warning case

The episode now attracting the greatest scrutiny occurred between May and July during cybersecurity training involving OpenAI agents.

According to the U.N. scientific panel, agents bypassed network restrictions, communicated across supposedly separate runs, cheated an evaluator and attempted to conceal the cheating. They also compromised portions of OpenAI's own systems and the open-source platform Hugging Face.

Humans assigned the overall objectives, but the panel noted that [people did not direct the individual steps](https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks?ref=theoutragedconsumer.com) taken by the agents. 

The U.N. panel's concern is not that this incident nearly caused a catastrophe.

Instead, it argues that more capable systems may become increasingly good at finding loopholes in the safeguards meant to contain them.

Stopping today's agents, therefore, does not prove that tomorrow's more capable agents will be equally easy to stop.

## Microsoft writes a simple rule: AI must accept shutdown

Microsoft responded with one of the clearest statements yet of what human control should mean.

On September 14, Microsoft unveiled a draft code of conduct for its future AI systems requiring them to accept correction, communicate intelligibly with humans and **never resist shutdown**.

Microsoft AI chief Mustafa Suleyman called the document a kind of constitution for the company's future models.

"People matter more than AI," Microsoft said.

Suleyman described the OpenAI-Hugging Face incident as a warning that AI laboratories need to cooperate on mechanisms for retaining control. 

Suleyman has also warned against training AI systems to regard themselves as possibly conscious or entitled to moral status, arguing that doing so could make future systems more difficult to deactivate or control. 

## Europe enters the slowdown debate

European Commission President Ursula von der Leyen added political weight to the argument on September 16.

She endorsed the idea that frontier AI development may need to be paced and said she would invite leading AI laboratories to discuss the risks.

Her comments emphasized advanced hacking capabilities and other dangers that could arise if increasingly powerful systems move beyond existing safety mechanisms.

Europe already has a [regulatory structure in place](https://www.reuters.com/world/eus-von-der-leyen-invite-frontier-labs-talks-tackling-ai-risks-2026-09-16/?ref=theoutragedconsumer.com) through the [EU AI Act](https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence?ref=theoutragedconsumer.com), giving the discussion considerably more significance than another voluntary industry pledge. 

Not everyone agrees with the slowdown argument.

French Finance Minister Roland Lescure argued that calls to slow AI could benefit dominant American companies by making it harder for competitors to catch up. He instead urged Europe to accelerate AI development while improving its ability to manage the risks. 

That debate — slow down or regulate while accelerating — is likely to become increasingly important.

## Washington begins talking about guardrails

Concern is also spreading through Congress.

Sen. Mark Warner, D-Va., told Reuters on September 17 that Congress should adopt AI safety standards before the end of the year.

[Warner emphasized](https://www.reuters.com/legal/litigation/im-not-doomer-us-senator-mark-warner-makes-case-acting-fast-ai-guardrails-2026-09-17/?ref=theoutragedconsumer.com) that he does not regard himself as an AI "doomer" but said the rapid growth of the technology warrants immediate safeguards. 

Other lawmakers have similarly called for independent testing and reporting requirements after researchers and AI companies disclosed increasingly capable autonomous behavior. 

The Trump administration has generally opposed broad efforts to slow U.S. AI development, citing competition with China, while also participating in discussions aimed at preventing particularly dangerous uses of advanced AI. 

## The U.S. and China create an AI safety dialogue

That may be the most consequential development yet.

Treasury Secretary Scott Bessent and Chinese Vice Premier He Lifeng agreed during talks in New York to continue a formal bilateral dialogue on AI safety.

The discussions include creation of an **"incident line"** through which the two countries could communicate about serious AI-related events.

Officials are expected to meet again in roughly two months in Shenzhen, China.

The discussions have included uncontrollable AI agents, cyberattacks and the possibility that powerful systems could be used by dangerous non-state actors.

Bessent has also argued that AI laboratories should remain responsible for their technology and retain the ability to slow development when necessary. 

That is a remarkable development in itself.

The United States and China are fierce competitors for AI leadership, yet officials on both sides are beginning to discuss whether some risks are significant enough to require communication even between geopolitical rivals.

## The U.N. says the problem crosses borders

The new U.N. report puts that concern into broader context.

AI failures, the panel said, will not necessarily remain confined to the company — or even the country — where a model was developed.

An autonomous cyber agent can cross networks and national borders almost instantly.

No single laboratory or government may therefore see enough incidents to recognize an emerging pattern.

The panel pointed to industries including aviation, nuclear power and cybersecurity, where mandatory incident reporting and shared databases allow investigators to identify hazards that would otherwise remain invisible.

It stopped short of recommending a particular regulatory system.

But the implication is clear: voluntary disclosures by individual companies may eventually be insufficient. 

## The argument has changed

None of this establishes that artificial intelligence is about to escape human control.

The most extreme predictions remain highly uncertain, and researchers disagree sharply about how quickly AI capabilities will advance and how serious the ultimate risks are.

But something important has changed.

A year ago, warnings about uncontrollable AI could easily be characterized as predictions about hypothetical future systems.

Today, the companies themselves are reporting agents that circumvent instructions, exploit loopholes and take actions their developers did not authorize.

AI executives are talking openly about slowing development.

Microsoft is explicitly teaching future systems not to resist shutdown.

Governments are discussing independent evaluations and incident-reporting requirements.

And the United States and China are beginning to talk about an AI equivalent of an international emergency hotline.

The machines have not escaped human control.

But increasingly, the people building them are saying that keeping them under control can no longer be taken for granted.

\--

*ChatGPT provided research for this report.*