These moves all make sense when you take into account the enterprise market.
I remember when bandwidth was super expensive and now it’s dirt cheap.
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
It's not even anything controversial..
Astra is a pretty impressive model. Excited to try this.
Edit: for context, just Steam alone has ~200million monthly active users.
How many DCs are devoted solely to gaming?
Huge misstep releasing it.
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
Guess not?
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
The only way is for prices to go up. Way up.
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
The best thing is that we benefit from these constant back and forth gut punches :)
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.
not saying this is the case here but it does feel a bit like wine tasting sometimes, everyone claims to be an expert that can taste a few tokens and tell you exactly what region and vineyard its from.
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
Edit: removed a comment that was uncharitable and rude, for which I apologize.
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
We've gone from 80% in some places to 80% in some more places.
With Opus 5.5, it doesn't seem like model improvement is plateauing. And Fable 5.5 will likely be dropped this or next week.
You can only imagine what they've got going internally.
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
People were talking about plateau for years already.
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
it's also why there have been so many calls for regulation and slowdowns.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring for some.
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.
An entire planet. Just Steam alone has one or two hundres million monthly active users.
Yeah? Show me the big movements against computer gaming.
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
However it's less willing to obey your instruction so it's less usable for general runtine flows.
Looks like 5.5 is the new 4.6
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
> Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
Our VC-backed subscription days are numbered
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
I can justify $200/mo but more than double is not appealing to me.
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
Come on .. this is barely released and you can already make that assessment?
And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
$500 for the old $200 is definitely a fumble.
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
hlynurd•1h ago
tedsanders•1h ago
thejazzman•1h ago
https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%...
pkulak•1h ago
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
cmrdporcupine•1h ago
Will see if this remedies things.
pkulak•1h ago