frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

https://github.com/Niko1221/Strata
83•snehesht•1h ago

Comments

snehesht•1h ago
I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

https://huggingface.co/Qwen/Qwen3.8-Flash-Next

proc0•1h ago
Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
incognito124•43m ago
Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore
snehesht•40m ago
Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
nicce•23m ago
I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
mickeyp•37m ago
I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

snehesht•24m ago
This is interesting, thanks. - https://github.com/Neroued/ninfer
thatsabadlook•19m ago
Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
geye1234•13m ago
I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
thatsabadlook•21m ago
Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
esafak•51m ago
Has anyone calculated the effective intelligence of these quantized models?
mkl•43m ago
There's some info in the README, including:

> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

https://github.com/Niko1221/Strata#which-model-should-i-pick

javier2•38m ago
ok that is getting interesting!
nicce•26m ago
I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?
nsagent•12m ago
See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

  We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

[1]: https://arxiv.org/abs/2608.08188

quietFalcon•50m ago
Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
snehesht•38m ago
They have some community benchmarks published https://github.com/Niko1221/Strata/tree/main/bench/results
merbanan•11m ago
Q2_0 does 33 tok/s decode and ~600t/s prompt processing at 128k context on RTX2060 8GB VRAM.

ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing

gdevenyi•41m ago
I had this working with the FreeToken inference engine a month ago when they launched.

https://github.com/FlashML-org/FreeToken

deadbunny•33m ago
> Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

And I thought piping to bash was bad

snehesht•32m ago
Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.
gchamonlive•20m ago
Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
prettyblocks•27m ago
I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
hypfer•27m ago
Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

The Readme doesn't say, but it's all AI generated, so..

panny•17m ago
I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
MrDrMcCoy•13m ago
Ternary Bonsai 2 might be for you.
0xbadcafebee•15m ago
Lol, sure, if you quant it to hell (Q2) it'll go real fast...
Tepix•11m ago
Q2 quantization. Not interested.

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

https://github.com/Niko1221/Strata
90•snehesht•1h ago•32 comments

Glashütte Trash Clock – A 30-minute pendulum clock made from trash

https://niklasroy.com/gtc/
28•r0r0•2d ago•4 comments

VGHF Digital Archive passes 5000 magazines. Here's what's next

https://gamehistory.org/5k-magazines/
53•rdmuser•4h ago•6 comments

Tell HN: Bob Cringely has died

589•paveworld•13h ago•116 comments

Why don't more developers “use the platform”?

https://nolanlawson.com/2026/10/03/why-dont-more-developers-use-the-platform/
192•vinhnx•9h ago•174 comments

Show HN: AI search for every photo and every frame of video on macOS

https://github.com/allenv0/SCM
28•allenleee•4h ago•15 comments

The work by Valve's Timur Kristóf on improving old AMD GPUs on Linux

https://www.phoronix.com/news/XDC-2026-Valve-Timur-AMDGPU
351•speckx•18h ago•61 comments

gpuvis: GPU Trace Visualizer

https://github.com/mikesart/gpuvis
42•luu•1d ago•5 comments

Rejection Sensitivity in Gifted and Twice-Exceptional Children

https://teachyourkids.substack.com/p/rejection-sensitivity-in-gifted-and
41•actfrench•2h ago•14 comments

Agents don't need memory, they need documentation

https://liao.gg/blog/agents-dont-need-memory
244•kmeh•21h ago•132 comments

Treachery in the Rodin Museum 3D scan verdict

https://cosmowenman.substack.com/p/rodin-museum-3d-scan-verdict
243•CosmoWenman•20h ago•117 comments

Emitting metadata early makes building/checking Rust up to twice as fast

https://github.com/PowderworksCode/headstart
39•knuckleheads•7h ago•3 comments

Hole Punch: Sling your spaceship around gravitational fields

https://notoriousbfg.com/hole-punch/
317•trwhite•20h ago•75 comments

We're working on a new RuneScape MMO

https://play.runescape.com/4
110•droidjj•12h ago•55 comments

So you think you could be an electrician?

https://asteriskmag.com/issues/15/so-you-think-you-could-be-an-electrician
340•zdw•3d ago•267 comments

LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents

https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-s...
216•Anon84•20h ago•318 comments

Celebrating the 100th birthday of the kidney donated to him as a teenager

https://www.whec.com/top-news/webster-man-celebrating-the-100th-birthday-of-the-kidney-his-mom-do...
216•gscott•2d ago•57 comments

The Heilbronn Problem

https://math.tejstead.com/heilbronn/
3•tejstead•1d ago•1 comments

We're going to need default hard budget caps on pretty much everything

https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/
502•elffjs•13h ago•257 comments

Reasons I didn't become an EMT, ranked

https://ben.stolovitz.com/posts/reasons-not-emt-ranked/
182•citelao•17h ago•81 comments

We want you to build the next Git platform on Cloudflare

https://blog.cloudflare.com/next-git-platform-on-cloudflare/
162•geoffbp•18h ago•138 comments

Religious scholars met with Anthropic

https://www.nytimes.com/2026/09/29/us/anthropic-claude-morals-ai.html
119•bookofjoe•11h ago•255 comments

I quit OpenAI because its culture is broken

https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/?gift=v5U_Uz...
316•Brajeshwar•1d ago•552 comments

Math's pedagogical curse – Grant Sanderson [video]

https://www.youtube.com/watch?v=UOuxo6SA8Uc
59•bobajeff•1d ago•28 comments

Your body of work thinks back at you

https://photoni.st/index.php/2026/09/25/your-body-of-work-thinks-back-at-you/
77•surprisetalk•4d ago•13 comments

What Meta got right with Muse

https://metedata.substack.com/p/what-meta-got-right-with-muse
75•young_mete•19h ago•100 comments

How to hack time, with C2PA

https://www.da.vidbuchanan.co.uk/blog/hacking-time.html
47•Retr0id•19h ago•6 comments

Dirty Optimization Secrets (C for Playdate)

https://devforum.play.date/t/dirty-optimization-secrets-c-for-playdate/23011
29•ibobev•1d ago•4 comments

Surely you have ultra-wideband radios on your bins too?

https://sjg.io/writing/binrange-have-you-actually-put-the-bins-out/
99•simonjgreen•17h ago•52 comments

FTL: A new operating system for clouds

https://ftl-os.org/
186•romac•23h ago•74 comments