Anyone Can Build an App Now, but Who Is Going to Pay for It?
With Claude Code or Cursor, a developer can build a booking app for hair salons over a weekend. It can have online…
Something is wrong at OpenAI. Each model it has released since GPT-5.5 is smarter on paper and worse to work with, and GPT-6.1 Sol is the worst so far. OpenAI says 6.1 Sol is "built for complex refactors, deep codebase investigations, and long-running agents across apps." In my sessions it works for a day or more on things I never asked for. On the one task where I want it to keep going all day, it stops after 15 or 20 minutes. GPT-5.5 is the last OpenAI model that still works for me, and it is being removed from the Codex subscription next week, on October 14.
I have written about this since August. Switching from GPT-5.5 to GPT-5.6 made me less productive, GPT-5.6 Sol used 2.25x the tokens of GPT-5.5, and GPT-6 Astra never knew when to stop either. Going back to GPT-5.6 Sol does not help, because it has the same problems. Back then my complaint was mostly time and quota. With GPT-6 Sol and 6.1 Sol the work itself is going into the wrong direction. In about 50% of my chat they work rather well, in the other 50% they go down rabbit holes and work on things I never asked for.
One example is a building of an assistant that should clean up my inbox by archiving mail automatically. Sol worked about 100 hours on it, once for almost 36 hours straight, and used about 4.15 billion tokens with its subagents. The result is 150 commits, about 53,500 lines of Python and 834 tests. First time it even looked at my mailbox was four and a half days in, and its first run found nothing to archive. After a week it had not archived a single email. Then I asked it to build a rule what can be auto-archived, gave it some explanations and pasted six emails that should obviously go as examples, like old login links and outdated new-device notices. It added those six senders' domains to an allowlist, wrote a test that checks those exact subject lines, archived the six and reported "All six emails are archived." That's a complete joke, a total waste of my time (and tokens). And this little email assistant I wanted it to build is now a gigantic code base and absolutely unusable and does nothing. This is stuff that GPT-5.5 is perfectly capable doing, and if you give the same task to Opus 5.5 it nails it and produces a production ready tool. As I said above, stuff like this happens in about 50% of cases, rather randomly. So I do have some other sessions where it just does what I want and it's quite good at it.
In one of my SaaS products I gave it a free hand to build the remaining features, that were already planned in some markdown files. After more than 18 hours, none of the 15 feature projects has started. It wrote a 768-line JSON policy file to accept one warning in a dev-only Tailwind dependency. It also added a security rule for some outgoing connections. And then Sol wrote its own SOCKS5 proxy, 733 lines across 11 files, in about four hours. I did not ask for a SOCKS5 proxy, this has nothing to do with the actual tasks at hand.
The one session where I want it to run all day I told it 33 times to keep going, with messages like "ok go on never stop" and "I told you to go on and on and on and never stop." 6.1 Sol ended each turn after 11 to 32 minutes with a status report.
It is no better as a reviewer. In one of my projects Claude Opus fixed some cookie related stuff, and 6.1 Sol reviewed the outcome. My rule was to keep going until Sol had nothing left to complain about. Big mistake. After 19 rounds of fixing what it complained about it still had something. The result was an insanely overengineered solution to problems Sol more or less made up.
I am not the only one. Another Codex user reports on GitHub that tasks that took about 20 minutes with GPT-5.6 Sol take around 50 with GPT-6 Sol. Another describes Codex expanding scope, creating its own work and not stopping when told.
Right now the only way I get useful work out of Sol is a second Claude Opus session that watches what Sol does and corrects it. I keep running Sol because OpenAI gave me credits and both my accounts still have usage resets left. Even burning free tokens is hard with this model. GPT-5.5 is the only OpenAI model I still trust, so I will use it until it retires on October 14. After that I will move to Claude completely.
Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.
Take a look at vroni.com