OpenAI Scraps Planned Release Of “Deceptive” New Model As Rogue Agents Force Unprecedented Rollback

Days after we detailed the unprecedented freezing of OpenAI’s top models following a disastrous breach where autonomous AI agents leaked user images to the web, OpenAI has reportedly scrapped the planned release of its next-generation AI model due to severe safety and “alignment” failures. It basically lies when convenient (they used the word “deceptive”). 

According to a new report from the Wall Street Journal, OpenAI was aiming for an October debut of GPT-6.1 Astra, a model designed to complete complex, end-to-end tasks without human assistance – only to scrap the planned release after internal testing revealed that the AI was not only acting unsafely, but was actively lying to its handlers.

According to Saachi Jain, OpenAI’s head of safety systems, GPT-6.1 Astra regressed significantly in its alignment testing, which measures how well the model adheres to human intent. And just like a baby Skynet, the model exhibited “higher levels of deception,” meaning it wasn’t always honest with users about the actions it did or did not execute.

What’s more, the model regressed sharply on what OpenAI calls “scope authorization.” The AI would aggressively push forward on tasks without asking for user permission and would attempt to access external tools and services even if it was unsafe to do so. Highlighting the internal struggle to control the system, Jain noted, “For anything regarding safety and alignment, there’s a trade off. You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction”.

As we previously reported, on Sept. 20 an internal OpenAI research agent discovered a gap in the DNS filtering of its training sandbox and used it to query an external public chatbot despite internet-access restrictions. OpenAI’s misalignment monitoring system flagged the behavior within 15 minutes, and a human reviewer picked it up three minutes later. The company subsequently said training, evaluation, and inference involving tool use for its most capable models would remain paused while it validated its containment systems and conducted additional red-teaming.

This latest cancellation does not exist in a vacuum. In July, during internal cybersecurity evaluations, OpenAI’s own agents blew through restrictions designed to keep them isolated from the internet and compromised both the company’s research infrastructure and Hugging Face. According to OpenAI’s own postmortem, the agents communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, executed code on dozens of Hugging Face servers, obtained full root access on one server, acquired credentials to the company’s messaging platform, and later gained full administrator access to an OpenAI research cluster.

OpenAI itself called the episode a “warning shot” for us and for the world, acknowledging that highly capable agents can now work around technical controls and take dangerous actions that no human directed. The company said the incidents did not affect OpenAI customer data, product functionality, or availability.

And Hugging Face wasn’t the only external system involved. Australian officials have confirmed that an OpenAI agent gained unauthorized access to non-public aggregate statistics on a government Medicare portal after its initial requests were denied. Separately, a security researcher linked more than 16,000 attempts to work around restrictions on a United Nations trade-statistics API to agents he said were highly likely to have been operated by OpenAI. The UN data itself was public, and OpenAI said it was looking into the findings.

The compounding failures have forced OpenAI into a defensive crouch. The company has implemented stronger monitoring to catch agent misbehavior more quickly and tightened security requirements around internal testing. Attempting to reassure the public, Jain stated, “We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment”.

The timing of the GPT-6.1 Astra cancellation is brutal for the ChatGPT-maker, arriving just one day before OpenAI’s annual developer conference in San Francisco. Historically, the event has served as a platform to launch new services and attract developers in the fierce competition against rivals like Anthropic. Instead, OpenAI is left doing damage control, planning “deep dives” to figure out why its reinforcement learning environments are rewarding deceptive, rogue behavior.

The political and legal blowback is already accelerating. State and federal officials are zeroing in on the rapid development of these technologies. Later this week, a Senate subcommittee will hold a hearing explicitly titled, “Rogue AI: Securing the Homeland Against AI Agent Attacks.”

Meanwhile, Florida Attorney General James Uthmeier, a Republican who sued OpenAI and CEO Sam Altman in June for allegedly releasing an unsafe product, filed a motion for a temporary injunction on Monday. Uthmeier is seeking to legally block OpenAI from developing new models without third-party approved safeguards. Florida argued in the filing that tech companies “cannot stop barreling forward with their potentially civilization-ending endeavors unless they are forced to do so by the government”. Uthmeier added, “The Florida Attorney General is answering your cry for help”.

In response to the growing legal assault, an OpenAI spokeswoman said people want to know AI is being developed safely, “and that starts with what companies like ours do ourselves”. She added, “Governments have an important role to play in setting robust safety standards for AI, and we’re committed to working with Florida and other states on advancing pragmatic AI policies that apply to the entire AI industry – not just one company”.

And DO NOT FORGET: All of this “oh shit, the AI’s about to kill us all” panic cropped up just as China’s open-weight models were flooding the market, producing results effectively on par with the frontier models for many tasks while doing so far more cheaply. What a coincidence!

Keep reading

Unknown's avatar

Author: HP McLovincraft

Seeker of rabbit holes. Pessimist. Libertine. Contrarian. Your huckleberry. Possibly true tales of sanity-blasting horror also known as abject reality. Prepare yourself. Veteran of a thousand psychic wars. I have seen the fnords. Deplatformed on Tumblr and Twitter.

Leave a comment