frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

https://github.com/Niko1221/Strata
47•snehesht•47m ago

Comments

snehesht•47m ago
I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

https://huggingface.co/Qwen/Qwen3.8-Flash-Next

proc0•36m ago
Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
incognito124•15m ago
Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore
snehesht•13m ago
Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
mickeyp•9m ago
I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

esafak•23m ago
Has anyone calculated the effective intelligence of these quantized models?
mkl•15m ago
There's some info in the README, including:

> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

https://github.com/Niko1221/Strata#which-model-should-i-pick

javier2•11m ago
ok that is getting interesting!
quietFalcon•23m ago
Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
snehesht•10m ago
They have some community benchmarks published https://github.com/Niko1221/Strata/tree/main/bench/results
gdevenyi•14m ago
I had this working with the FreeToken inference engine a month ago when they launched.

https://github.com/FlashML-org/FreeToken

deadbunny•6m ago
> Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

And I thought piping to bash was bad

snehesht•4m ago
Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.

Show HN: B1 Sprechen – Practise German B1 Speaking Exams with AI

https://b1sprechen.com/
1•HumphreyZ•3m ago•0 comments

Show HN: Untyped – check recorded agent runs against a TLA+ spec

https://github.com/untyped-ai/untyped
1•damianabramov•6m ago•0 comments

How Effective Altruism Conquered the World

https://www.economist.com/international/2026/10/01/how-effective-altruism-conquered-the-world
1•bazzmt•8m ago•0 comments

Can an Android app without the INTERNET permission phone home?

https://arjun.maniyani.com/gander/phone-home.html
1•mokshablr•12m ago•0 comments

Ask HN: If you're struggling with p(doom), how are you handling it?

1•pyronite•14m ago•2 comments

The Thirty Million Line Software Problem [video]

https://www.youtube.com/watch?v=kZRE7HIO3vk
1•childintime•15m ago•0 comments

Tell HN: Python.com Is for Sale

2•mococa•16m ago•0 comments

What a Jurassic rainforest may have sounded like

https://www.science.org/content/article/what-jurassic-forest-may-have-sounded
1•bookofjoe•17m ago•1 comments

Teaching LLVM a trick it knew

https://rhyadav.dev/blog/teaching-llvm-a-trick-it-already-knew
1•slowrah•18m ago•1 comments

Writing an Operating System in 1k Lines

https://github.com/nuta/operating-system-in-1000-lines
1•teleforce•21m ago•0 comments

One Thousand Origami Cranes

https://en.wikipedia.org/wiki/One_thousand_origami_cranes
1•vismit2000•22m ago•0 comments

The Evolution of Smalltalk – From Smalltalk-72 Through Squeak [Ingalls 2020]

https://dl.acm.org/doi/pdf/10.1145/3386335
1•AlexeyBrin•22m ago•0 comments

All You Need Is Cable TV? (2019)

https://www.tandfonline.com/doi/epdf/10.1080/00220388.2018.1506581?needAccess=true
1•vismit2000•22m ago•0 comments

What AI debates have to do with alchemy

https://www.programmablemutter.com/p/what-ai-debates-have-to-do-with-alchemy
1•rwmj•24m ago•0 comments

Latency Implications of Virtual Memory

https://rigtorp.se/virtual-memory/
2•porridgeraisin•26m ago•0 comments

GTA V – Graphics Study

https://www.adriancourreges.com/blog/2015/11/02/gta-v-graphics-study/
1•porridgeraisin•27m ago•0 comments

The Laffer Curve

https://en.wikipedia.org/wiki/Laffer_curve
2•baxtr•29m ago•0 comments

Why I tried to kill token billing (and why we kept it)

https://stripe.com/blog/where-pricing-is-headed
1•torutofu•30m ago•0 comments

macOS ProactiveHarvesting and AppleIntelligenceReportingSELFIngestor

https://knowledgeisuserdata.medium.com/apple-store-dfu-installs-appleintelligencereportingselfing...
1•cududa•31m ago•1 comments

US Hackers Face $1k-a-Month Pay Cuts in Controversial Plan

https://www.bloomberg.com/news/articles/2026-10-02/us-hackers-face-1-000-a-month-pay-cuts-in-cont...
1•sbulaev•32m ago•0 comments

Over 20 NHS trusts have ditched Palantir waiting-list tools, analysis shows

https://www.ft.com/content/cc1f2e28-117f-41bb-96b9-030746cf9530
1•sbulaev•32m ago•0 comments

AI's 'Thought' Process Can No Longer Be Trusted

https://www.wsj.com/tech/ai/ai-monitoring-chain-of-thought-research-b46a05fd
3•1vuio0pswjnm7•42m ago•1 comments

Mac OS 9 Platinum desktop recreated in the browser

https://wieslawsoltes.github.io/MacOS9/
3•wiso•45m ago•0 comments

Six operating systems that inspired Windows and Linux

https://www.howtogeek.com/these-operating-systems-inspired-windows-and-linux-and-one-is-probably-...
2•dxs•46m ago•0 comments

Herbarium – keep and review the HTML pages AI tools generate

https://github.com/abdoufermat5/herbarium
1•abdouyaya1998•47m ago•0 comments

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

https://github.com/Niko1221/Strata
52•snehesht•47m ago•13 comments

Harden Your AI Agent

https://sometechblog.com/harden-your-ai-agent
1•l5870uoo9y•53m ago•0 comments

Ya-GPT: free in-browser ML training

https://gpt.codexes.io/
1•zteppenwolf•53m ago•0 comments

Improving and Stabilizing the Racoon2 IKE Daemon in NetBSD

https://blog.netbsd.org/tnf/entry/gsoc2026_racoon2
1•jaypatelani•54m ago•0 comments

Tokyo's record rainfall streak ends after 37 days

https://english.kyodonews.net/articles/-/87138
1•giuliomagnifico•54m ago•0 comments