frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Accelerating GPT-5.6 Sol Ultrafast

https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai
123•pr337h4m•1h ago

Comments

HawtAds•53m ago
Their dinner plate chips are impressive.
crazysim•49m ago
GPT 5.6 Luna Ultrafast when?
GodelNumbering•49m ago
The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/

There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding

rirze•42m ago
They're expanding access to companies that apply for the program and explain their use cases. So it's very real but limited imo.
WarmWash•24m ago
The stake in the side of cerebras has always been that the economics are pretty poor.

Who knows if they will subsidizes it to mitigate sticker shock, but it's a safe assumption that it will be scarily expensive. However if you are in a "cost is no obstacle, speed is god" position, it will likely be pure magic.

fcarraldo•18m ago
Can anyone explain why Cerberus needs to be _fast_ instead of _cheap_?

I don't think I understand why they aren't leveraging the increased speed to do batching to serve more customers at a "normal" tok/s.

Is the limitation, even on cerberus, still that the cache can only serve so many concurrent sessions over time? Is there no scaling advantage? I genuinely do not understand how any of this works.

jaggederest•3m ago
They're cache limited, almost certainly, so more slower sessions doesn't solve the problem - you still have to load and unload the whole cache hierarchy at some level and that's a network bandwidth and memory bandwidth problem between the external systems and the waferscale chip.

Also worth looking into how they do cooling for it, because that's kind of absurd and awesome as well.

wxw•49m ago
> Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.

Awesome work. I'm personally very excited for faster models/inference.

I think speed is underrated to some degree in the current conversation. For a while, I was using Cursor's Composer quite a lot, even over frontier models, just because of how darn fast it was.

kilroy123•43m ago
I've been using DeepSeek flash a lot this week to try it out. Now, I deeply want the smart frontier models to be just as fast.
arw0n•6m ago
What do you need speed for? That's a genuine question, I feel like the limiting factor already is my creativity, attention span and budget. And I'm not even yet optimizing cost by batching things like review to slow local models over night, or schedule tasks to take full advantage of my subscriptions.
poly2it•47m ago
I guess Gemini 3.7 Flash is no longer at the pareto frontier of speed to intelligence.
odo1242•38m ago
Well, there’s still price
thraway3837•42m ago
This is really cool. Someone here commented about similarity between this and hardware advancements for AV encode/decode.

I think it's only a matter of time before miniaturization can have a thumbnail sized user-replaceable accessory that contains the LLM built onto the hardware. I admit I don't know how any of that works, but would be amazing to experience. Fully local, fully offline, ultra fast local inference better than any personal computing product.

christkv•34m ago
https://chatjimmy.ai/ Is that. Company behind it just got acquired by AMD
iamcoder18•31m ago
I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration.

> In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.

This is actually insane.

Hopefully the release ultrafast of Terra and Luna too.

piyh•26m ago
Feels like the 90's again where single threaded speed is improving fast. ASICs and wafer scale rather than node shrinks, but end result to me the consumer feels the same.
storus•29m ago
Wow, that's even faster than diffusion LLMs but with the Fable-level quality! Congrats!
scotty79•29m ago
I swear that now frontier AI stuff comes out few times a week.
applfanboysbgon•19m ago
This kills the crab.

Compilation time will be a genuine bottleneck for slop coding if this becomes the standard generation rate over the next few years. Go, Zig or even C99 with TCC for dev builds, any language that can get you systems-level performance (or close to it) in a dev environment where you can iterate in ms rather than minutes is going to be immensely more appealing than generating a potential prototype in 10 seconds and waiting 15 minutes for it to compile.

yetihehe•13m ago
Maybe then LLM's will switch to outputting raw machine code?
behnamoh•18m ago
Fast mode is already 1.5 times faster and 2x more expensive in the Codex subscription plan. If this thing is 14 times faster, then I can imagine running out of my quota in one session.
fg137•16m ago
> allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical work

Curious, what are some of the use cases?

ricardobeat•15m ago
The omission of Mimo v2.5-Pro Ultraspeed, released in June, which can achieve 1000tok/s is an interesting flaw in the comparison graphs.

It is a bit outdated (scores ± 40% lower), but smart enough for a lot of coding tasks, and can cost under 1/10th of Sol.

https://mimo.mi.com/models/en-US/mimo-v2.5-pro-ultraspeed

anthonypasq•11m ago
I'd just like to point out that the largest model Cerebras has ever served is Kimi K2.6 which is 1T parameters, so that either means that theyve had a breakthrough on the hardware engineering side of things, or GPT-5.6 Sol is likely a lot smaller than people think.

If it truly is only ~1-2T parameters, then this kinda kills 2 narratives for me.

1. all the handwringing about open source catching up via Kimi K3 (3T params) is complete nonsense. All that matters imo for determining which labs are leading is intelligence per parameter. Anyone with a enough compute can train a giant model, but being able to squeeze capabilities into smaller models gives you a massive inference and training edge.

2. Inference margins are clearly insane, and this explains why OpenAI was able to lower the price of Luna by 80%. Id guess that thing is probably 120b params based on the TPS they are serving it at.

manmal•4m ago
Isn’t the fact Fable is more expensive than Sol-Max by multiples already an indication that Sol is way smaller?
owentbrown•9m ago
Whoa. This looks both powerful and expensive.

My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference.

Within labs, I've heard spend is already far beyond this per developer.

jaggederest•5m ago
I mean I don't think $50k is the ceiling, unless you're talking about actual cash out. Claude code subscriptions right now can easily clear you $25-35k a year in nominal value for $2400 out of pocket cost.

Given sufficient budget and scope, I could certainly productively burn a half million dollars in tokens a year or more. I think that's where we're headed anyway, buying a 2nd or 5th claude max subscription feels slightly excessive for personal usage, but at a corporate level...

Topfi•6m ago
Unless I have read over it, besides the animation in the intelligence vs speed graph which only mentions internal data and not whether they truly reran the AA suite, there is no actually solid statement on the important aspect of performance.

Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfully) scream that off the rooftops. A line such as "this is the same performance, just faster, with no downsides" would go a long way in clarity and communication. Along with no pricing information, I'll hold out on further information.

[0] https://openai.com/index/previewing-ultrafast/

Gemini 3.7 Flash

https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-fl...
307•thisisauserid•2h ago•201 comments

Accelerating GPT-5.6 Sol Ultrafast

https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai
128•pr337h4m•1h ago•30 comments

Choose Boring Technology (2015)

https://mcfunley.com/choose-boring-technology
95•tosh•1h ago•52 comments

Donkey.bas is 45 Years Old – 131 line of Glory

https://donkeybas.com/
68•jkrauska•1h ago•24 comments

Mistral OCR 4.1

https://docs.mistral.ai/models/ocr-4-1
121•spelk•2h ago•37 comments

Spaghettifying DRAM

https://github.com/xoreaxeaxeax/skitter-creek-bath-salts
336•matt_d•5h ago•104 comments

Where did the old web go? We followed 657,607 links to find out

https://0.mk/blog/link-rot
40•tdx•1h ago•17 comments

Tocharian Online

https://lrc.la.utexas.edu/eieol/tokol/0
32•Bluestein•2h ago•1 comments

DeepSeek Harness developer preview

https://deepseek.com/harness/en/
460•bjin•6h ago•209 comments

Come for ENIAC, Stay for UNIVAC and Skeduflo

https://uniqueatpenn.wordpress.com/2026/08/05/come-for-eniac-stay-for-univac-and-skeduflo/
46•cainxinth•2d ago•12 comments

Kubernetes on Oxide: How customer needs shaped our integrations

https://oxide.computer/blog/kubernetes-on-oxide
102•stevehipwell•5h ago•48 comments

Gloomberb

https://gloom.sh/
307•rbanffy•5h ago•163 comments

AI At Home Part 1: A Box Of Scraps

https://jdagostino.github.io/ai-pt1-box-o-scraps/index.html
37•timmmmmmay•3h ago•18 comments

How art invented humanity

https://aeon.co/essays/humans-did-not-invent-art-it-was-the-other-way-around
52•prismatic•20h ago•12 comments

Ordinary abundance

https://ordinaryabundance.com/
123•yen223•5h ago•52 comments

Choosing an AI model: one prompt, 11 models, different results

https://www.netlify.com/blog/one-prompt-11-models-very-different-results/
134•toddmorey•6h ago•58 comments

Show HN: Pixy, visual editor for coding agents, like Figma on your live site

https://pixydesignapp.com/
3•astroscout•11m ago•1 comments

Codex in ChatGPT desktop app for Linux is now in preview

https://community.openai.com/t/codex-in-chatgpt-desktop-app-for-linux-is-now-in-preview/1390027
411•allanrbo•14h ago•286 comments

JDK 27 G1/Parallel/Serial GC Changes

https://tschatzl.github.io/2026/08/10/jdk27-g1-serial-parallel-gc-changes.html
18•0x54MUR41•2h ago•5 comments

I built a 500k-domain search engine for makers in a weekend for $10

https://alexmorleyfinch.github.io/marlin/history/v1/article/the_birth.html
97•dreamforever•5h ago•57 comments

ATG (YC F25) Is Hiring Member of Technical Staff (Data Platform)

https://atg.science/careers
1•dkobran•7h ago

GoAccess – Open-source real-time log analyzer and interactive viewer

https://goaccess.io/
21•gregsadetsky•3h ago•3 comments

Show HN: Inertia – Not your Grandfathers animation editor

https://inertiagraphics.com/
7•hpen•4d ago•1 comments

Graduate student proves a quantum uncertainty principle for fractals

https://www.quantamagazine.org/graduate-student-proves-the-fractal-uncertainty-principle-20260812/
50•bookofjoe•5h ago•6 comments

Show HN: MCP Memory – Fast Agent Memory Using Google's OKF and SQLite FTS5

https://github.com/fellowgeek/mcp-memory
47•pcbmaker20•5h ago•30 comments

Launch HN: Bullet (YC S26) – A Faster Coding Agent

https://www.codewithbullet.com
28•adi1•11h ago•31 comments

Better Gaussian Splatting in Julia

https://pxl-th.github.io/blog/better-gs-julia/
100•pxl-th•4d ago•15 comments

I requested a copy of my data from McDonald’s loyalty program

https://www.wired.com/story/mcdonalds-built-a-515-page-dossier-on-me-it-says-ill-never-leave/
171•thehoff•4h ago•211 comments

Show HN: OJCP – an open protocol for agent-consumable job data

https://ojcp.dev/
17•fraywing•1d ago•3 comments

Deutsche Bank becomes first foreign yuan clearing bank in Europe

https://tradersunion.com/news/central-banks/show/2973571-deutsche-bank-becomes/
351•Markoff•7h ago•385 comments