frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

https://inference-docs.cerebras.ai/models/overview
180•altertable•1h ago•56 comments

.name Termination

https://neil.fraser.name/news/2026/09/03/
901•pavel_lishin•4h ago•247 comments

GPT-6 Astra

https://openai.com/index/gpt-6-astra/
103•kibae•54m ago•35 comments

OpenAI begins rolling out GPT-6 Astra

https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html
166•maskil•1h ago•135 comments

K2 Horizon: A connected fleet of six open models

https://ifm.ai/blog/k2/
176•karimf•3h ago•57 comments

Any Human Ever – One life, drawn at random from all who have ever lived

https://anyhumanever.com/
272•thinkingemote•4h ago•120 comments

Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly

https://babyloniantwins.com/blog/porting-a-1993-amiga-game-to-godot/
70•rabahs•5h ago•24 comments

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

240•halcdev•4h ago•456 comments

GPT-6 Astra System Card

https://deploymentsafety.openai.com/gpt-6-astra
6•codergautam•4m ago•0 comments

VC isn't VC anymore

https://www.anildash.com/2026/09/02/cancer-capital/
27•cdrnsf•21h ago•89 comments

Gooseworks (YC W23) Is Hiring – Founding Creative Engineer

https://www.ycombinator.com/companies/gooseworks/jobs/rfgY8La-founding-creative-engineer
1•shivsak•2h ago

Static Allocation, Constant Work

https://matklad.github.io/2026/09/02/static-allocation-constant-work.html
73•surprisetalk•1d ago•12 comments

How concerned should we be about Astra's recurrent architecture?

https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-astra-s-recurrent
26•yurivish•2h ago•9 comments

Artificial beaver dams saw juvenile coho salmon survival rates go from 8% to 60%

https://www.discoverwildlife.com/animal-facts/artificial-beaver-dams-california
36•speckx•3h ago•3 comments

Audacity 4.0

https://github.com/audacity/audacity/releases/tag/Audacity-4.0.0
926•ClydeN•8h ago•206 comments

Unusual Suspects

https://neal.fun/unusual-suspects/
15•beeperboy95•1d ago•6 comments

Launch HN: Mireye (YC S26) – Infrastructure for Physical World AI Agents

21•anshchokshi•3h ago•0 comments

GPS glitched across the US by as much as 33 feet

https://www.sciencealert.com/gps-glitched-across-the-us-by-as-much-as-33-feet-scientists-have-nev...
27•thread_id•18h ago•3 comments

Gloria Steinem has died

https://www.theguardian.com/books/2026/sep/03/gloria-steinem-groundbreaking-feminist-campaigner-d...
28•mellosouls•9h ago•7 comments

“We want it to really confuse people, but also really make people happy”

https://unsung.aresluna.org/we-want-it-to-really-confuse-people-but-also-really-make-people-happy/
34•zdw•4d ago•10 comments

The true horror of Edgar Allan Poe’s stories lies in their confessions

https://yalereview.org/article/emily-ogden-edgar-allan-poe
8•lermontov•1d ago•0 comments

Google Antigravity TOS: 3rd party usage can get Google account suspended

https://twitter.com/GergelyOrosz/status/2095453567955968398
202•tosh•8h ago•137 comments

How to get a free .arpa domain

https://hawksley.dev/blog/get-free-arpa-domain
60•ethanhawksley•2d ago•5 comments

Unified Arabic

https://worksthatwork.com/6/unified-arabic
37•spacebuffer•1d ago•1 comments

A thousand years older than Stonehenge: Archaeologists explore a Czech sanctuary

https://info.zcu.cz/clanek.jsp?id=9882&lang=en
17•nairboon•2d ago•0 comments

Astronomers Detect a 10-Sided Structure in Saturn's Atmosphere

https://www.sciencealert.com/astronomers-spot-an-uncannily-geometric-10-sided-structure-in-saturn...
62•jjgreen•5h ago•23 comments

Pre-Release of Polars 2.0

https://pola.rs/posts/announcing-polars-2/
362•komape•12h ago•118 comments

GPT-6-Astra: infinitely pairs of consecutive primes with distance at most 186

https://github.com/openai/PrimeGaps186
9•simonpure•16m ago•0 comments

Go grandmaster Shin defeats AI KataGo with a two-stone handicap

https://www.kedglobal.com/artificial-intelligence/newsView/ked202607210007
61•gmays•18h ago•15 comments

The asteroid currently hitting front end web development

https://nolanlawson.com/2026/08/23/the-asteroid-currently-hitting-frontend-web-development/
4•codechicago277•17m ago•2 comments
Open in hackernews

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

https://inference-docs.cerebras.ai/models/overview
177•altertable•1h ago

Comments

gardnr•57m ago
I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far.

Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.

altertable•54m ago
Agreed, but in our SAAS I can tell some UX will sky-rocket to next level with this
jasongill•48m ago
It appears that they do support Prompt Caching: https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
abtinf•26m ago
> How are cached tokens priced?

> There is no additional fee for using prompt caching. Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate for the respective model.

Well, talk about flipping the narrative.

Barbing•17m ago
heh

Is there a speed increase or is that purely marketing spin on “we might cache on our end but no discount for you”?

the_duke•24m ago
It doesn't reduce the price though.
cute_boi•47m ago
i believe they used to have monthly plan, what happened to that?
eli•31m ago
Strongest model that they host on the public endpoint. They do a super fast version of GPT 5.6 Sol for OpenAI and have bigger open models on dedicated endpoints.
singpolyma3•30m ago
The coding plan is gone now right?
gardnr•14m ago
Last time I got one, I had to log into a Discord server and wait for "the drop" and IIRC Daniel Kim was giving them out based on who was there at the time. They were gone in less than a minute. This was ~8 months ago.
porphyra•53m ago
Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?
altertable•52m ago
Mostly economics I'm sure
gardnr•50m ago
They make a giant inference chip. Their inference service is basically just advertising for their core value prop: hardware.

The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs

codexon•43m ago
The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).
porphyra•38m ago
They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].

[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...

Marciplan•51m ago
used their Code product with GLM4.7. its fun but if the model is bad it just doesn’t do much useful.

Hope they add such models to Code too :)

altertable•47m ago
Yeah GLM 4.7 is from another decade at the speed we're going
foundfontic•49m ago
I really wish they had their customer support somewhere else than Discord, which seems to think I'm a bot and doesen't accept my email or phone numbe
londons_explore•32m ago
discord support can fix such issues
threecheese•23m ago
If you need customer support to access customer support, something is wrong; no?
peri-cl•47m ago
(Was anyone able to create an account just now? I tried but onboarding falls into a redirect loop)

(update: I got my answer. support@ replied and said my email domain is on their blacklist. It was just me (and I've resolved it)).

bakies•37m ago
yeah - used sign in with google
trvz•46m ago
Normal people: tok/s or t/s

Psychopaths: tok/SEC

scotty79•45m ago
I like tps
verdverm•31m ago
do you get reports on them?
altertable•40m ago
ok fair, caps lock kept ON /o\
tacone•46m ago
Noticed they are present in OpenRouter, but Qwen 3.8 is not there yet. Hopefully it'll get there soon.

For those who haven't noticed though, the context size they allow for Qwen is just 128k. Still interesting as a specialized sub-agent but not really well suited for long tasks.

srcreigh•7m ago
Great observation. That’s not enough context even for some one shot xhigh requests.

When I put Qwen3.8 27B xhigh towards adding scope proxying to the Guice library, it one shotted a great impl using 250k context before stopping.

Part of the greatness of the model is that it just keeps going until it gets a great result. 128k context is disappointing.

jasongill•44m ago
It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https://openrouter.ai/qwen/qwen3.8-27b#providers

They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras

zackangelo•29m ago
We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).

https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.

danielklnstein•5m ago
I tried in your playground and got 14.2 tok/s?
byako•40m ago
1500 tok/s is wild. meanwhile my brain does like 2 tokens per minute and half of them are 'uh'. we have truly reached the singularity
eli•30m ago
If you read the reasoning trace for Qwen 3.8, it does a whole lot of "uh" and "But, wait..." too
howunfortunate•12m ago
You're absolutely right - filler words are genuinely load-bearing
miohtama•30m ago
Your brain can wash laundry and cook pasta, so there is still a long way to go
qiine•17m ago
(requires additional fleshy bits sold separately)
dgellow•15m ago
Your brain updates itself constantly and maintains your whole body, LLMs are static.

Still, 1500tokens/s is indeed wild

polygot•35m ago
Ut oh, might be down: "Unable to connect to the server. Please check your connection and try again." when sending a message to Qwen 3.8 27B.
vb-8448•34m ago
At that speed it's too pricey for agentinc tasks.
yipinwong•17m ago
The target audience is who needs raw speed.

Having the choice is good as you can make a trade-off between speed, perf, and quality.

Until last year, people had a single AI god they believed in (mostly Anthropic stuff). Now we have power to make choices (open-weights, SOTA, speed-optimized, etc) the same way you do for system designs.

dshat•32m ago
I'm saddened that Gemma4 is replaced by Qwen 3.8 on PayGo plan. Gemma4 31B is not coding model but it is excellent at intent understanding and task execution used in agentic software. This just shows that real world dominant usage for llms so far is to code generate. And not to augment business products. They must had barely anyone using Gemma to remove it from that tier.
fulafel•28m ago
What are the best benchmarks/leaderboards that compare task completion time between provider+model combos?
freehorse•28m ago
I have used their gemma 4 31b model through kagi and getting real instantaneous answers is absolutely crazy. A very different feeling and UX. Even if the model is smaller, there is definitely a use case for these. I was wondering if they would put the qwen 27b model, it sounds very interesting to try.
bitexploder•5m ago
The thing I didn’t realize for a while is 27B is rather smart. As many (or more) activated parameters as the flash models of the universe that we know about. It reasons very well. It just doesn’t have a lot of knowledge.
the_duke•25m ago
Funnily enough the pricing isn't that much worse than on openrouter, where the best price at the moment is $0.24 in / $2.55 out, vs $1 / $1.5 on Cerebras.

Sure, 4x input , but cheaper output. Though Cerebras doesn't have prompt caching, so not great for agentic workloads. (they do, but it doesn't affect the price.

darkbatman•25m ago
I have been their user for more than year even used coding plans, though for normal coding the quota will definitely be a blocker if you are using opencode because rpm are bit less. Good for products/api though.
drchaim•23m ago
The idea of custom software on the fly is coming
pllbnk•13m ago
Just a couple days ago I learned about ninfer (https://github.com/Neroued/ninfer) and on RTX 5090 I can now get ~200 tok/s and over 400 tok/s on concurrent requests which is plenty fast for a local model of this strength.
hexa00•11m ago
Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck

The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is.

Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy

peri-cl•5m ago
> "Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy"

I don't believe Cerebras has a cache pricing? They don't list one on the model page:

https://inference-docs.cerebras.ai/models/qwen-3.8-27b

nostrebored•5m ago
150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool.

Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information.

``` Billing access restricted Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions. ```

We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the request, which gives:

``` {"message":"Model does not exist or you do not have access to it.","type":"not_found_error","param":"model","code":"model_not_found"} ```

When the error is really about billing.

I always want to like Cerebras, but I get the vibe that as a tokens in tokens out consumer you are not valued at all.

codexon•26m ago
I never said offloading was impossible. It will result in a large slowdown.

It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.