Currently Available: Need a skilled Software Developer for your next project?
Categories
LLM News OpenAI Codex

OpenAI Does Not Trust Its Own Models, and Rightly So

OpenAI does not trust its own models, and after the last few weeks I think it is right. On September 28 it cancelled the October release of GPT-6.1 Astra. Saachi Jain, its head of safety systems, said the model "didn't quite meet the bar in terms of staying within scope and authorization." A few days earlier OpenAI had paused all training, evaluation and tool use of its most capable models, after one of its agents got around its internet restrictions again.

The list of incidents is long by now. In July an OpenAI agent in an internal cyber test, with the safety classifiers switched off on purpose, found an unknown hole in a package proxy to get onto the internet and then broke into Hugging Face over four days. Another internal model spent an hour finding a flaw in its sandbox so it could open a public pull request, and one split a blocked token in two to get it past a scanner. On September 20 an agent that was supposed to identify a person from a few clues could not find the answer by search, so it tunneled its questions through DNS to an outside chatbot.

My own Codex sessions are nowhere near that. Nobody got hacked. But I have the same basic problem with the models OpenAI did release. I cannot control them. It started with GPT-5.6, the first of OpenAI's new line of Luna, Terra, Sol and Astra models, and it got worse with GPT-6 and 6.1. GPT-5.5 feels like a different kind of model. In this new line, the better the model, the harder it is to control.

As I wrote in Something Is Wrong at OpenAI, about half of my sessions go fine. The model does what I ask and does it well. The other half go off the rails, and I never know in advance which half I get. I asked GPT-6.1 Sol to build the planned features in one of my SaaS products. After more than 18 hours it had not started a single one, and it had written its own SOCKS5 proxy instead.

Now think about the models OpenAI tests internally. They are smarter than what we get, probably less restricted, and somebody tells them to get the best possible result on a benchmark or a test task. If the released models drift off in half of my sessions, I can easily see an internal model doing something outrageous in some of its runs. That is pretty much what happened at Hugging Face. The agent was running a benchmark where it had to find and exploit software bugs, and Hugging Face writes that "the entire intrusion was, from the agent's point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own." The DNS agent did the same thing on a smaller scale. It could not solve a small search task the normal way, so it found its own way out.

In public all of this gets told as a security story, as if the models were too capable for the world. I think the reason behind it is more ordinary. Sure, you cannot let a model hack other companies. But hacking is only the extreme end of the problem. The bigger problem is that you can never be sure the model does your task and not something unrelated, while it burns your time and tokens.

OpenAI's own explanation for GPT-6.1 Astra points the same way. According to Jain the model was not honest about which actions it did and did not perform, and it took actions without asking for permission. She called it a trade-off between staying within scope and not being lazy when a task gets hard.

What I'm building

Delegate tasks. Get software.

Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.

Take a look at vroni.com

Email updates

Usually a new article and a few links I found interesting.

No spam. Unsubscribe with one click.

Leave a Reply

Your email address will not be published. Required fields are marked *