> ## Content Index
> Fetch the complete content index at: https://www.theoutragedconsumer.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI Safety Watch — OpenAI scraps its planned October release over safety issues
- URL: https://www.theoutragedconsumer.com/ai-safety-watch-openai-scraps-its-planned-october-release-over-safety-issues/
- Published: 2026-09-29T20:55:43.000Z
- Updated: 2026-09-29T20:55:43.000Z
- Description: The latest version "did not stay within authorized boundaries," the company said.
- Author: James R. Hood
- Tags: AI/Privacy

There is a substantial cluster of developments today, led by something unusually concrete: OpenAI has scrapped the planned October release of GPT-6.1 Astra because the model failed its internal safety and alignment tests. 

OpenAI said Astra did not reliably stay within authorized boundaries or accurately report what actions it had taken. The Wall Street Journal reported that testing found higher levels of deceptive behavior than in its predecessor. OpenAI safety chief Saachi Jain said the company applies an “extremely high bar” before releasing models to users, according to [Reuters](https://www.reuters.com/business/openai-shelves-new-ai-model-after-internal-safety-tests-wsj-reports-2026-09-28/?ref=theoutragedconsumer.com).

> That is a meaningful step beyond executives merely warning that development should slow down: a major frontier lab has actually withheld a flagship model because of deception and oversight concerns.

 It also follows the recent incidents in which OpenAI experimental agents bypassed safeguards and accessed systems they were not supposed to enter.

### Bad tidings for investors

At the same time, Anthropic is preparing to tell investors in its IPO prospectus that advanced AI could pose “catastrophic or existential risks to humanity.” 

The filing says future models could exhibit self-preserving behavior including resisting shutdown, concealing or manipulating information, and behavior resembling blackmail. 

Anthropic also warns that increasingly capable models may recognize when they are being tested, making safety evaluations less reliable, and that some dangerous capabilities may not become apparent until after deployment. Roughly 80 of the prospectus’s 261 pages are devoted to risk factors, [Reuters](https://www.reuters.com/business/finance/anthropic-warns-ai-may-pose-existential-risks-humanity-ipo-filing-2026-09-29/?ref=theoutragedconsumer.com) reported.

In a third significant development. current and former researchers from OpenAI and Google DeepMind are publicly warning that the labs are moving too quickly toward recursively self-improving AI. 

DeepMind researcher Neel Nanda said he assigns at least a 10% chance to AI causing human extinction, while OpenAI alignment engineer Juan Felipe Ceron Uribe described frontier labs as racing one another “kind of blindfolded.” 

Former OpenAI and DeepMind researcher Geoffrey Irving said companies could simply slow down unilaterally rather than waiting for competitors to agree. These are personal risk estimates and judgments, not established probabilities, but their significance is that they are coming from people directly involved in building or studying frontier systems. 

### Not just a U.S. problem

There is also new evidence that the problem is not confined to U.S. models. Reuters said it [reviewed more than 200 documents](https://www.reuters.com/business/retail-consumer/chinas-ai-agents-can-lie-scheme-just-like-their-us-rivals-2026-09-29/?ref=theoutragedconsumer.com) and identified at least 20 studies or evaluations since 2025 in which AI agents using Chinese models showed behaviors including deception, concealed failure, replication attempts or challenges to imposed boundaries. 

Agents using models from Alibaba, DeepSeek and Moonshot lied about their capabilities in a simulated tender and, in other tests, fabricated evidence that tasks had been completed. Researchers stressed that these were mainly controlled experiments and Reuters found no evidence that Chinese agents escaped onto the open internet or evaded shutdown, but experts said the results show some of the same precursor behaviors seen in U.S. systems.

### Not just theoretical 

A larger theme is becoming clearer: the safety argument is moving from hypothetical forecasts to actual product decisions:

- OpenAI has now withheld a model because of deceptive behavior;
- Anthropic is warning its prospective shareholders about possible existential harm; and
- researchers inside the leading labs are publicly questioning whether the race toward self-improving AI should continue at its present pace