frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

New Teardown: Combination Lock

https://mechanical-pencil.com/products/lock
1•crescit_eundo•59s ago•0 comments

Interactive demo of MINIX1-like O/S on emulated CPU

https://swtos.softwarewrighter.com/
1•softwarewright•1m ago•1 comments

Nybble – Describe a binary format in a small DSL and watch it parse live

https://trynybble.com/
1•bitcask•1m ago•0 comments

RSA-260 Factored

https://www.johndcook.com/blog/2026/09/03/new-rsa-number-factored/
1•floxy•2m ago•0 comments

Arrested for a Late Manuscript: Seicho Matsumoto's 'Tokyo Express'

https://www.millersbookreview.com/p/arrested-for-a-late-manuscript-seicho-matsumoto-tokyo-express
2•benbreen•2m ago•0 comments

'Fakers' by Rory Cormac Review

https://www.historytoday.com/archive/review/fakers-rory-cormac-review
2•pepys•3m ago•0 comments

Every Novel Is Boring–Until It Isn't

https://www.publicbooks.org/every-novel-is-boring-until-it-isnt/
2•samclemens•3m ago•0 comments

Understanding Rust's Pin Type

https://gmcgoldr.github.io/2026/08/27/pin-in-rust.html
1•garrinm•4m ago•0 comments

Data Center Infrastructure: A Decision Reference [pdf]

https://www.cs.mu.edu/~papers/StevenGoodman/DataCenters/DataCenterInfrastructure-ADecisionReferen...
1•sgoodmanabc•5m ago•0 comments

Why Does Free Time Become Another Thing You Have to Use Correctly?

1•nullifypreset•5m ago•0 comments

Florida bans highway license-plate readers as backlash over surveillance spreads

https://www.reuters.com/legal/government/florida-bans-highway-license-plate-readers-backlash-over...
2•1659447091•5m ago•0 comments

Former Ukraine Defense Minister to open defense tech company backed by Palantir

https://kyivindependent.com/ukraines-former-defense-minister-fedorov-to-open-defense-tech-firm-wi...
1•hentrep•5m ago•0 comments

MapQuest app reaches No 1 on US Apple list

https://www.theguardian.com/us-news/2026/sep/01/mapquest-lake-ontario-trump
1•tony101•6m ago•0 comments

OpenAI Cut Off a Billion-Dollar Customer to Avoid Elon Musk

https://www.wired.com/story/openai-elon-musk-cursor-billion-revenue/
2•sbulaev•6m ago•0 comments

Astra is at 98.6% score on ARC-AGI-3

https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot-2026-09-03-at-10.51.35-am-1024x880.png
2•TheJCDenton•6m ago•0 comments

China, Egypt to do away with the use of the US dollar

https://africa.businessinsider.com/local/markets/the-worlds-second-largest-economy-just-decided-w...
2•mikhael•7m ago•1 comments

GPT 6 Astra

https://openai.com/gpt-6-astra/
7•skogstokig•10m ago•4 comments

Robot startups are trying everything they can think of to get more data

https://www.understandingai.org/p/robot-startups-are-trying-everything
1•alehlopeh•11m ago•0 comments

GWM Worlds 2

https://runway.com/research/introducing-gwm-worlds-2
1•tobr•11m ago•0 comments

Linode's $12 VPS Didn't Outperform Its $5 Plan in Six Fresh Deployments

https://webbynode.com/articles/linode-5-12-48-los-angeles-six-fresh-deployments
1•gsgreen•12m ago•0 comments

DiskSpace: Building an app 3 (ok 4) times in a day with Claude

https://shaneosullivan.wordpress.com/2026/09/03/clean-your-mac-windows-linux-pc-with-disk-space/
1•shaneos•13m ago•1 comments

Comfortable Fluency in Consuming Information Is Not a Proxy for Actual Learning

https://www.justinmath.com/comfortable-fluency-in-consuming-information-is-not-a-proxy-for-learning/
2•yarapavan•15m ago•0 comments

Cincinnati police drones respond in mins, here's what they learned after 1 year

https://local12.com/news/local/cincinnati-police-drones-first-responders-program-results-what-cpd...
2•thinkcontext•17m ago•1 comments

After Reflection: The Runtime Story – Saksham Sharma – C++Now 2026 [video]

https://www.youtube.com/watch?v=bUmt9K1o1d0
2•matt_d•19m ago•0 comments

Final Ruling After 30 Years, Kraftwerk Sample Stays Legal

https://www.gearnews.com/sampling-copyright-eu-kraftwerk/
1•thm•20m ago•0 comments

Matching Puzzle Pieces and Disappointing Benchmarks

https://llogiq.github.io/2026/03/20/case.html
1•speckx•20m ago•0 comments

Truchet Tiles

https://twitter.com/pickover/status/2095581910365802964
2•Ariarule•22m ago•0 comments

Redwood: A Frontier AI Accelerator Designed from Scratch in 2 Weeks by AI

https://arxiv.org/abs/2608.26418
1•imakwana•22m ago•0 comments

Restorekit – Reformat Any T2 or Apple Silicon Mac from macOS, Linux or Windows

https://www.restorekit.org/
1•bnb•23m ago•0 comments

Repeat in Excel 97

https://unsung.aresluna.org/unsung-heroes-repeat-in-excel-97/
2•Bluestein•23m ago•0 comments
Open in hackernews

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

https://inference-docs.cerebras.ai/models/overview
117•altertable•41m ago

Comments

gardnr•35m ago
I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far.

Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.

altertable•33m ago
Agreed, but in our SAAS I can tell some UX will sky-rocket to next level with this
jasongill•27m ago
It appears that they do support Prompt Caching: https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
cute_boi•26m ago
i believe they used to have monthly plan, what happened to that?
eli•9m ago
Strongest model that they host on the public endpoint. They do a super fast version of GPT 5.6 Sol for OpenAI and have bigger open models on dedicated endpoints.
singpolyma3•9m ago
The coding plan is gone now right?
porphyra•32m ago
Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?
altertable•30m ago
Mostly economics I'm sure
gardnr•28m ago
They make a giant inference chip. Their inference service is basically just advertising for their core value prop: hardware.

The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs

codexon•22m ago
The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).
porphyra•16m ago
They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].

[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...

Marciplan•29m ago
used their Code product with GLM4.7. its fun but if the model is bad it just doesn’t do much useful.

Hope they add such models to Code too :)

altertable•25m ago
Yeah GLM 4.7 is from another decade at the speed we're going
foundfontic•27m ago
I really wish they had their customer support somewhere else than Discord, which seems to think I'm a bot and doesen't accept my email or phone numbe
londons_explore•11m ago
discord support can fix such issues
peri-cl•26m ago
(Was anyone able to create an account just now? I tried but onboarding falls into a redirect loop)

update: I got my answer. support@ replied and said my email domain is on their blacklist.

bakies•15m ago
yeah - used sign in with google
trvz•25m ago
Normal people: tok/s or t/s

Psychopaths: tok/SEC

scotty79•24m ago
I like tps
verdverm•10m ago
do you get reports on them?
altertable•18m ago
ok fair, caps lock kept ON /o\
tacone•24m ago
Noticed they are present in OpenRouter, but Qwen 3.8 is not there yet. Hopefully it'll get there soon.

For those who haven't noticed though, the context size they allow for Qwen is just 128k. Still interesting as a specialized sub-agent but not really well suited for long tasks.

jasongill•23m ago
It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https://openrouter.ai/qwen/qwen3.8-27b#providers

They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras

zackangelo•8m ago
We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).

https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.

byako•19m ago
1500 tok/s is wild. meanwhile my brain does like 2 tokens per minute and half of them are 'uh'. we have truly reached the singularity
eli•9m ago
If you read the reasoning trace for Qwen 3.8, it does a whole lot of "uh" and "But, wait..." too
miohtama•8m ago
Your brain can wash laundry and cook pasta, so there is still a long way to go
polygot•14m ago
Ut oh, might be down: "Unable to connect to the server. Please check your connection and try again." when sending a message to Qwen 3.8 27B.
vb-8448•12m ago
At that speed it's too pricey for agentinc tasks.
dshat•11m ago
I'm saddened that Gemma4 is replaced by Qwen 3.8 on PayGo plan. Gemma4 31B is not coding model but it is excellent at intent understanding and task execution used in agentic software. This just shows that real world dominant usage for llms so far is to code generate. And not to augment business products. They must had barely anyone using Gemma to remove it from that tier.
fulafel•7m ago
What are the best benchmarks/leaderboards that compare task completion time between provider+model combos?
freehorse•6m ago
I have used their gemma 4 31b model through kagi and getting real instantaneous answers is absolutely crazy. A very different feeling and UX. Even if the model is smaller, there is definitely a use case for these. I was wondering if they would put the qwen 27b model, it sounds very interesting to try.
codexon•5m ago
I never said offloading was impossible. It will result in a large slowdown.

It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.