frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Advancing the price-performance frontier with GPT‑5.6

https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
165•tedsanders•1h ago

Comments

bakugo•47m ago
> GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less

Looks like the Chinese models are really making a dent. Having 3 different price categories with the "most affordable" one still costing more than GLM 5.2 never made sense.

measurablefunc•41m ago
It all comes back to electricity cost. China has cheaper electricity so as long as China keeps pace there is no way for American companies to undercut them. Each boolean operation in China is cheaper than the one in America.
cbg0•18m ago
This doesn't seem correct.

Estimated final electricity price for large industrial customers in energy-intensive industries:

USA 50 USD/MWh

China 68 USD/MWh

https://www.iea.org/reports/electricity-2026/prices

preommr•38m ago
I thought the chinese models were cheaper per token, but about the same or more expensive on tasks because they used more tokens for reasoning. Cutting even further, seems like a really big leap.
simonw•46m ago
> The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%.

If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month?

We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we don't know how much of Anthropic's inference capacity that is (presumably a small fraction, since they were operating on top of AWS and other providers before the SpaceX deal.)

I've not seen any numbers that hint at OpenAI's per-month inference bill, but surely that has to be in the multiple billions of dollars as well.

So 20% is a really, really big deal.

dominotw•43m ago
imagine writing that on your resume

> reduced inference cost by 20 percent saving company x billion dollars per month

tekacs•40m ago
In this case, and I don't mean this critically, I guess it would technically be, "Instructed model to find efficiencies... reducing inference cost by 20% saving company x billion dollars per month."

I have no doubt that further work was required to enable this, but it's still very cool to be possible to say that.

da_grift_shift•31m ago
Does the model get the credit for its promo packet then? :^)
andai•26m ago
I think they meant that GPT-5.6-Sol can write that on its resume.
pavpanchekha•46m ago
Making Luna, which was already very cheap and extremely capable, 5x cheaper is crazy. I use Sol at work but Luna at home, and while there's definitely a difference, it doesn't feel like night-and-day. After a year of ever-increasing prices it suddenly feels (between this, Kimi K3, GLM 5.2) that prices are falling again.
jedberg•42m ago
> Sol vs Luna

> it doesn't feel like night-and-day.

I see what you did there. :)

deklesen•19m ago
Good observation! Kudos
handfuloflight•9m ago
You're absolutely right.
dominotw•41m ago
depends on what you are doing. if you are doing verifiable tasks like fixing bugs then any model would do as long as you write the right verification.
maxdo•26m ago
is kimi that cheap? it's a very expensive model
pixelesque•
measurablefunc•44m ago
Model segmentation & distillation like this that asks the consumers to pick exactly which version of the algorithm will solve their problem is evidence for lack of intelligence instead of its presence.
dominotw•39m ago
it is really hard to know upfront if you have fuzzy task. sometimes i would choose a cheaper model and it will spin and spin with bad outputs ending up costing more had i chosen a more capable model.
cute_boi•31m ago
there is mixture of experts which is also another routing. So, simple change in prompt can be a big difference.
beering•32m ago
You really really don’t need to pick. Just use Sol on high. That’s my daily driver and I don’t touch the model picker at all.

Now, if cost is your concern, then that’s a problem in all of computing. Hence why I’m sending you short plain text messages using an iPhone with a many-core CPU and gigabytes of RAM.

sidcool•42m ago
They don't mention Grok at all.
andybak•34m ago
Don Draper in the elevator meme?
hirako2000•33m ago
Of course. All comparison is with what makes them look good.
wilg•29m ago
What would they say about Grok?
paxys•27m ago
They also don’t mention a hundred other models.
goldsmith112•41m ago
Not sure who would use Terra anymore. Pair Luna High/Xhigh with Sol Medium and that's your power stack
fritzo•34m ago
Sounds reasonable. Is there a good benchmark on which make this decision?
espadrine•19m ago
I maintain this meta-benchmark leaderboard: https://metabench.organisons.com/

With this new price change, Terra does look pretty Pareto’ed by Luna.

On agentic coding, pairing Sol Medium for architecting with Luna High for coding does kinda make sense. But beware that architecting can be very read-heavy, and Sol is a bit read-pricey compared to Terra.

andai•21m ago
Sol as main agent, Luna for coding?
swingboy•40m ago
This is awesome. Luna is a pretty great model on xhigh.
preommr•40m ago
> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,

I don't have the words.

I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.

buckle8017•34m ago
They over purchased hardware.

This is very likely priced below recovering the cost of the hardware but still above operating expenses.

infecto•31m ago
What evidence is there?

I have no idea either way but one thing that detracts from these threads is folks claiming things as a fact without evidence.

qntmfred•21m ago
sama literally just said they wish they had bought more. the price drops are almost certainly due to good old fashioned hardware innovation (wafer scale with cerebras) and optimizing hardware development based on model architecture and inference costs. other inference providers will try to do the same if they can.

https://www.youtube.com/watch?v=XDB5beon4DY&t=4m20s

captainbland•33m ago
To be fair we don't really know in terms of prices what's real and what's just investor subsidised attempts at market capture at this point. It could well be OpenAI's attempt to drown Anthropic while they've got the halo product if they feel they've got deeper pockets.
gentlewater•38m ago
This is awesome. I’ve recently set up my opencode to use 5.6 terra for my main agent, who delegates work to a 5.6 Luna coder agent. So far it seems to work well, and reduce costs a lot. With this price reduction, it will work a whole lot better. Perhaps I can get my github copilot quota to last the whole month now.
tosh•31m ago
80% price cut for luna is a very aggressive pricing move

makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)

heaney-555•9m ago
Luna is meant to compete with Haiku. What tasks are you seeing it equal Opus on?
firasd•29m ago
This is one of the things OpenAI has been focused on for an year or so that led to the doomed autoswitcher in ChatGPT .com (switching models based on estimated task complexity) that was quickly reverted

Whereas Google with Gemini 3.x, Anthropic with Fable etc are happy to just go for 'big model with dense params'

It's hard to guess from the outside of course but just this kind of talking points focus on GPU efficacy is what we see from OpenAI and Chinese open source labs more often than from Anthropic or Google Deepmind and this benchmark chart seems to concur

bob1029•29m ago
This feels like the dialup->broadband transition to me.

I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the bank.

andai•25m ago
Do you have a sense of which tasks benefit from more agents and which don't?
bob1029•13m ago
Anything related to reading and interpreting the environment seems to always benefit from the addition of more agents to the search party, assuming you have some rational way to synthesize their results.

Taking actions that mutate the environment is a different story. I think this is where you run into diminishing returns very quickly. You generally want one strong agent to act given the results of all the searching that was done. If the plan is clear, you don't need a genius model to execute it.

handfuloflight•6m ago
I definitely think you want the genius model to synthesize everything that rolls up to them.
recitedropper•27m ago
Cool.. but impossible to gauge if this has to do with algorithmic optimizations; undercutting competitors; geopolitics; declining demand; the recent stock sellof; any other of the myriad of possibilities. OpenAI has not shown themselves to be an actor that you take at their word.

One hopes that, maybe one day, we'll have completely transparent inference profit, calculable carbon footprint per token, and no more predatory market-capture tactics. Maybe one day.

It is going to be remarkable how many question will be answered when we finally get to see the S-1.

alvis•27m ago
Basically lunar at extra level can cover all use cases scenarios other than those requiring opus up. Goodbye sonnet and haiku
sosodev•27m ago
Looks like I might have a reason to use something other than Deepseek V4 Flash.
andai•23m ago
I was curious so I photoshopped DSV4 Flash into the graph:

https://files.catbox.moe/csxl32.png

(2 cents to run AA index, score 40)

Looks like OpenAI broke the pareto frontier on the trust-me-bro benchmarks!

(One has to wonder if they used any of the neat tricks from the DSV4 paper :)

GodelNumbering•27m ago
“Half the money I spend on advertising is wasted; the trouble is I don't know which half.” -John Wanamaker

This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

fractorial•24m ago
Cosmically apt username given the substance of this comment.
quirino•26m ago
I generally just check the Price/Performance graph on Openrouter: https://openrouter.ai/rankings#performance#benchmarks. Activate the "Show Pareto" toggle on the right.

I was still using GLM-5.2 in my personal projects, but this just made Luna a very easy choice.

wronex•26m ago
What are your use case for these? I’m manly interested in coding where more capability is better - give me a 10x model at 10x the price and I’ll take it. A worse model at very low cost has no appeal to me. At-least not for coding. Translation maybe? OCR?
stri8ted•16m ago
Translation, moderation, classification, guardrails, etc..
kingstnap•25m ago
Those prices on luna are killer.

Haiku was already in a ditch.

But this is coming straight for the jugular of a ton of models on openrouter.

fractorial•24m ago
It would appear that rolling my own Anthropic-free harness / serving stack with a closed-weight carve out for Codex models is an absolute win.
peheje•22m ago
Might just resub. Will experiment with Luna next sessions. 5 h window is not working very well for me. But if I can drop down to Luna at 20-30 % left and comfortably ride out the wave then.. that might just work.
arcanemachiner•6m ago
They got rid of the 5hr quota, it's just weekly quotas now.
__jl__•20m ago
Didn't expect that. Luna pricing is crazy now. I don't think there is anything on the market that competes at this price-performance point.

For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.

dannyw•9m ago
OpenAI's APIs are extremely reliable for sure. I don't even remember when the last incident or downtime was.
jnakano89•19m ago
Seems like they cut the tiers(GPT-5.6 Luna) where GLM and Kimi compete and still held margin for their frontier models
ninjahawk1•18m ago
80% less for Luna is absolutely crazy, in my opinion we may reach a point in the next year where powerful models on the API could potentially be cheaper than subscriptions. Compute just keeps decreasing in price.
andai•18m ago
It says Luna is fastest, but doesn't it take way more steps to get the same job done?

https://deepswe.datacurve.ai/ - (See the Agent Steps view)

Or is the output speed so much higher that it cancels out?

I don't see a lot of benchmarks that record actual time. But on AA, Sol on Low beats Luna on High for Time Per Task.

incognito124•10m ago
While I can't deny this is a huge technological result, and it's laudable they reduced the price because of it, 80% is really a lot. I can't help but wonder, is this because of the model's capabilities, or was the initial system just really sloppy? The public will probably never know the details
baalimago•8m ago
We swapped an internal system from gpt-5-mini to gpt-5.6-luna and saw no benefit but 4x cost. Sufficed to say: we swapped back to gpt-5-mini.
purpleidea•7m ago
I would pay significantly more to use these models if there was a legal contract that guaranteed they weren't ever terfing them and some way to prove that.
hirako2000•34m ago
Contributed to. Can't be some IC who made a few nice PRs
paxys•30m ago
Where are you going to apply to with that resume that’s a step up from your current job though?
bpavuk•22m ago
lots of places, actually. not everyone wants to be attached to the Silicon Valley culture, and that line alone will guarantee practically any workplace. that person is going to find out what work-life balance is :)
NitpickLawyer•38m ago
~2 years ago gemini2.5 helped write better kernes for itself and (only) reached 1% efficiency gains. Today we're at 20%.
9m ago
It's cheaper currently on many of the inference providers.

Personally, I'm having surprisingly good results with DeepSeek 4 Pro at home, which is very good value for money: it's not as good as Claude / GPT 5.6 (I have Co-pilot license at work), but it's still really useful for code reviews, validating thoughts, and especially designing / writing unit tests for new (and old before refactoring) functionality.

And it's very cheap per task. (Flash is even cheaper, but I've had issues with that on more complex tasks where it starts forgetting things and arguing with itself "but wait, let me read the function again").

pioneer37•20m ago
Its just a matter of time at this point.These companies are working day and night to capture the market.
platinumrad•31m ago
We can guess based on the decisions of other inference providers who serve these models.
w29UiIm2Xz•27m ago
Enterprises implemented spending caps and inference providers are lowering prices. Seems they are jockeying for market share.
gentlewater•32m ago
This is gonna put Sonnet 5 in a really awkward spot.
bakugo•25m ago
Sonnet and Haiku were already in an awkward spot, likely by design.

Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it.

Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.

supern0va•13m ago
>figure out what the percentage split between Opus/Fable/Sonnet is.

This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.

baq•20m ago
I use sonnet as a smart grep and haiku never and that’s only when I have to use Anthropic at all
heaney-555•10m ago
Luna is comparable to Haiku, not Sonnet.
827a•9m ago
Totally untrue. Luna and Sonnet 5 are very comparable: https://artificialanalysis.ai/#intelligence

Luna is an extremely strong model.

foobar_______•27m ago
Hard to believe numbers. I don't mean that as a critique, but literally I am so impressed. Even if the model is a few percent lower for performance but is 80+% cheaper than competitors and is a US company hosted on US based hyperscaler clouds this is kind of a no brainer. Hard for most businesses to justify otherwise.
ismailmaj•24m ago
it's 80% less cost, not 80% in efficiency gains, could be that Luna was overpriced to begin with, we don't have much info on the models themselves.

Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?

dannyw•16m ago
Been using OpenAI models since ada/babbage/curie/davinci and at least from my own experience, their APIs feel the same.

If you use Codex it's different, the harness has a lot to do with it and there's definitely been changes including recently.

camel-cdr•12m ago
this type of thing usually means you are the product
827a•10m ago
Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
jpadkins•7m ago
When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X improvement in price/performance frontier.

The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.

Advancing the price-performance frontier with GPT‑5.6

https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
173•tedsanders•1h ago•90 comments

Read This Before You Buy That TV Streaming Stick

https://krebsonsecurity.com/2026/07/read-this-before-you-buy-that-tv-streaming-stick/
109•speckx•1h ago•39 comments

Gemini Robotics 2 brings whole body intelligence to robots

https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
258•ai2027•3h ago•259 comments

Physicists Solve a Muon Mystery. Now, Old Results Don't Add Up

https://www.quantamagazine.org/physicists-solve-a-muon-mystery-now-old-results-dont-add-up-20260729/
89•ibobev•2h ago•39 comments

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

https://www.bottlenecklabs.com/blog/autonomously-run-businesses
53•Areibman•46m ago•26 comments

Stacked PRs are now live on GitHub

https://github.blog/changelog/2026-07-30-stacked-pull-requests-are-now-in-public-preview/
62•tomzorz•1h ago•16 comments

The Economic Benefit of Refactoring

https://martinfowler.com/articles/exploring-gen-ai/refactoring-economic-benefit.html
113•javaeeeee•3h ago•48 comments

Why is everyone trying to build a solid-state battery?

https://www.construction-physics.com/p/why-is-everyone-trying-to-build-a
119•crescit_eundo•5h ago•148 comments

Rise Reforming (YC S26) Is Hiring

https://www.ycombinator.com/companies/rise-reforming/jobs/wJ9Q9nv-senior-chemical-process-engineer
1•george_rose25•1h ago

Hacker Public Radio

https://hackerpublicradio.org/
82•bmacho•3h ago•12 comments

Toot.community is shutting down

https://social.jorijn.com/@jorijn/statuses/01KYN00AP3NCZXCFB96KQB8GN2
30•speckx•1h ago•33 comments

Upper stage impacting the moon on 2026 August 5

https://www.projectpluto.com/25010d.htm
100•ryannevius•4h ago•31 comments

How Olinia Turns Mexico's EV Ambition into Reality

https://spectrum.ieee.org/mexico-olinia-car-electric-vehicle
17•rbanffy•1h ago•8 comments

RFC 8890 – The Internet is for End Users (2020)

https://mnot.net/blog/2020/for_the_users
89•notarobot123•5h ago•25 comments

Launch HN: Prized (YC S26) – Let non-engineer staff build secure internal tools

https://prized.dev
51•marinoseliades•4h ago•30 comments

How to Mount a Balcony Awning (2025)

https://solar.lowtechmagazine.com/2025/07/how-to-mount-a-balcony-awning/
24•karakoram•5d ago•0 comments

Ron Gilbert started production on Thimbleweed Park 2

https://www.grumpygamer.com/twp2_announce/
200•alberto-m•10h ago•95 comments

SDL_GPU minimal, single-header, high-performance 2D graphics painting library

https://github.com/n67094/sdl_gp
42•n67094•3h ago•13 comments

Are We Stuck with Lean?

https://mathoverflow.net/questions/513742/are-we-stuck-with-lean
93•jjgreen•6h ago•41 comments

Paging Through a Parquet File in DuckDB: File_row_number or Offset?

https://rusty.today/blog/paging-parquet-duckdb-file-row-number-vs-offset/
30•rustyconover•3h ago•3 comments

RCade: The Arcade Cabinet with CI/CD Deployment, Custom Graphics Card for CRT [video]

https://www.youtube.com/watch?v=W-OpIbLUOU0
20•evakhoury•21h ago•3 comments

Show HN: Claude-account – switch Claude Code accounts without logging in again

https://github.com/hamzarehmandeveloper/claude-account
22•hamza_rehman•3h ago•18 comments

How old is Ann?

https://quuxplusone.github.io/blog/2026/07/29/how-old-is-ann/
51•ibobev•6h ago•53 comments

Trusted URLs via Cryptographic Signatures

https://blog.certisfy.com/2026/04/trusted-urls-via-cryptographic.html
18•Edmond•3h ago•12 comments

3D Pinball for Windows (1995)

https://98.js.org/programs/pinball/space-cadet.html
77•mushstory•6h ago•39 comments

Gpiozero Flow

https://bennuttall.com/blog/2026/07/gpiozero-flow/
111•benn_88•7h ago•34 comments

Show HN: I made a game where you build a CPU from logic gates

https://select.supply/game/chipbuilder
58•laurentiurad•6h ago•53 comments

Building a native C# implementation of CEL engine

https://bsid.io/writing/building-a-cel-engine-for-net
38•jackedEngineer•2d ago•8 comments

The Alice and Bob After Dinner Speech (1984)

https://hex.ooo/library/alicebob.html
41•kamma4434•3d ago•6 comments

Azulejo

https://en.wikipedia.org/wiki/Azulejo
157•Amorymeltzer•1d ago•53 comments