frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

AWS Acquires DuckDB

https://ducklabs.com/news/2026/08/26/ducklabs-to-join-aws
309•onderkalaci•1h ago•67 comments

Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

https://qwen.ai/blog?id=qwen3.8-flash-next
158•tosh•1h ago•51 comments

RAG Is Simpler Than You Think

https://www.lighthousenewsletter.com/p/rag-is-simpler-than-you-think
231•j0selit0•5h ago•107 comments

Meta reaches $16.68B settlement over social media harms to children

https://www.reuters.com/world/us/meta-settles-with-us-states-over-social-media-harms-2026-08-26/
78•bhouston•41m ago•40 comments

A curmudgeon tries a language server

https://entropicthoughts.com/curmudgeon-tries-language-server
29•crescit_eundo•1h ago•7 comments

Oldinsurancemaps.net is now a Charter Project

https://openstreetmap.us/news/2026/08/oim-charter-project/
118•altilunium•5h ago•19 comments

XCancel and Nitter are receiving C&D letters from XCorp

235•mobilio•4h ago•89 comments

Fake US thinktank set up and funded by Israel sought to game AI for propaganda

https://www.theguardian.com/world/2026/aug/26/fake-thinktank-israel-ai-propaganda
166•n1b0m•1h ago•25 comments

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-...
274•garo-pro•4h ago•107 comments

Proliferate (YC S25) Is Hiring

https://www.ycombinator.com/companies/proliferate/jobs/OgpCKYJ-founding-product-engineer
1•pablo24602•2h ago

Apple introduces M6 and M5 Ultra

https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-perform...
1237•interpol_p•1d ago•1185 comments

Stalking the Wily Hacker: 40 years later – Cliff Stoll [video]

https://www.youtube.com/watch?v=656058JxTM0
174•zoenolan•4d ago•51 comments

Beyond Recall and the Illusion of Competence

https://var0.xyz/posts/beyond-recall-and-the-illusion-of-competence.html
49•tuxie_•4h ago•16 comments

FDA authorizes first wearable device that monitors ketone and blood sugar levels

https://www.fda.gov/news-events/press-announcements/fda-authorizes-first-wearable-device-continuo...
461•sunnynagra•19h ago•210 comments

Value Classes Still Need Compiler Sympathy

https://johan-sjolen.github.io/post/compiler-sympathy/compiler-sympathy/
58•lichtenberger•5h ago•21 comments

OpenAI Jalapeño: Better than Nvidia Blackwell

https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
547•bmulholland•1d ago•343 comments

Queryable Executables

https://fzakaria.com/2026/08/24/actually-queryable-executables
262•rguiscard•13h ago•73 comments

New Mac Studio with M5 Max and M5 Ultra

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
797•interpol_p•1d ago•527 comments

Harvest (IBM 7950): Supercomputer for cryptanalysis at the NSA in the Cold War

https://spectrum.ieee.org/cold-war-codebreaker-nsa-ibm
65•jnord•9h ago•17 comments

Black hole singularity is a surface not a point

https://arxiv.org/abs/2608.21590
283•raattgift•21h ago•196 comments

New Mac mini, featuring M6 and M5 Pro

https://www.apple.com/newsroom/2026/08/apple-unveils-a-more-powerful-mac-mini-featuring-the-all-n...
522•runako•1d ago•324 comments

Omarchy development practices lead to predictable security issues

https://blog.happyfellow.dev/merchants-of-insecurity/
229•arn3n•1h ago•312 comments

Show HN: Buslens – where can I get to by bus? (UK)

https://rupertlinacre.com/buslens/
69•RobinL•6h ago•46 comments

Maiao: Gerrit-style code review workflow for GitHub, GitLab, Gitea, others

https://github.com/runetes/maiao
101•zdw•15h ago•63 comments

When str.lower() is a security vulnerability in Python

https://sethmlarson.dev/when-str-lower-is-a-security-vulnerability
158•rbanffy•17h ago•68 comments

Building a backyard office, the build and cost breakdown

https://www.imkylelambert.com/articles/building-a-backyard-office-the-build-and-cost-breakdown
388•surprisetalk•23h ago•237 comments

Nitter and XCancel receive cease and desist notices

https://github.com/zedeus/nitter/issues/1442
1063•Banditoz•21h ago•900 comments

Tooltips need a delay, and then they need to skip it

https://blog.master.dev/tooltips-need-a-delay-and-then-they-need-to-skip-it/
208•ibobev•21h ago•73 comments

Run OpenBSD on DigitalOcean for $4/month

https://nil.wallyjones.com/run-openbsd-on-digitalocean-for-4month/
189•speckx•20h ago•86 comments

Bomb fishing is wreaking havoc on Indonesia's coral reefs

https://e360.yale.edu/digest/bomb-fishing-coral-reefs
341•speckx•23h ago•184 comments
Open in hackernews

Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

https://qwen.ai/blog?id=qwen3.8-flash-next
152•tosh•1h ago

Comments

whwhyb•1h ago
looks like it's better than deepseek v4 flash
freakynit•1h ago
Those benchmarks look seriously impressive.. considering how small of a MoE model this is.
skarz•1h ago
do we really need breaking news about qwen posted every single day?
iAMkenough•1h ago
yes there’s no shortage of online real estate
tosh•56m ago
this is a new architecture (foreshadowing qwen 4)

> trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board

https://x.com/Alibaba_Qwen/status/2092591393424515114

pseudony•50m ago
I and presumably quite a few others with AMD AI or Apple Mac platforms are very impacted by this.

:)

It is very relevant and for a certain group of us, far more impactful to our work the next month(s) than any blog post could be.

KronisLV•47m ago
If there’s news, then yes. This is a pretty great new release for those still stuck on Qwen3.6 35B A3B if they have enough memory but don’t have super powerful compute.

I wonder if I could get this running through vLLM on 6x Nvidia L4 - the 3.6 worked great on 4 cards but sadly TP6 just isn’t a thing and I don’t have 8 cards available, maybe it’s gonna be okay with like TP2 and MTP. I have no idea at this time, probably need to test out what even might be possible.

NitpickLawyer•43m ago
This particular release is interesting because it's a preview of qwen4 architecture. And, while benchmarks are iffy, this is a direct comparison, by the same team, with qwen3.8-27b that was pretty well received for a local model.

This "next" release adds a new concept, first public release with n-grams, I think. And it's in a MoE size that is likely to be very fast and cheap to serve (faster than 27b for sure). It's also well suited for inference on alternative compute (i.e. sparks, macs, etc) so it's relevant to local users.

dofm•40m ago
This actually is meaningful news, I think. Pretty wide audience appeal in the local LLM space too.
c16•40m ago
There are many topics, personalities and politicians we hear about daily who have no merit.

Qwen's advances do (currently) have merit.

christkv•57m ago
Looks like a good model for strix halo
Iolaum•45m ago
indeed, can't wait for it to be supported by llama.cpp (or other engines)?
tosh•56m ago
this is a new architecture (foreshadowing qwen 4)

> trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board

https://x.com/Alibaba_Qwen/status/2092591393424515114

rohansood15•55m ago
Didn't expect it to beat 3.8 27B so cleanly.

Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

Squarex•48m ago
I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.
LaurensBER•40m ago
If/when we can get larger context this will mostly be mitigated by these smaller models being able to search the internet.

Self-learning/improving would be even better but that's still a long way to go.

rohansood15•35m ago
For world knowledge, you'd want it to find and reference the source material to be sure. At that point, it doesn't matter if the knowledge is embedded.
quev•12m ago
Keep in mind a web search might not include scanned books baked in the weights ;)
hedora•11m ago
I think the big models have adequate recall, so tool use is probably unnecessary, but the user said the correctness of my response is important. Let me look up the data instead of relying on my memory.
lnenad•40m ago
Adding to my homelab stack, hopefully it doesn't overthink like the little model. Actually, hoping it thinks a bit less. Wait actually I'm really praying it reasons a bit more directly. But wait, I'm really sure that it must be a bit better.
grim_io•37m ago
That's low reasoning for a model, but max for a HN comment.
cyanydeez•26m ago
My stack is basically deer-flow with Qwen3.5-122B-A10B; this hopefully will be a speed and intelligence improvement. Running deer-flow overnight on any research topic or verify clear scoped programming issue is really neat.

Also, heating my home during the winter is nice.

Oh, also, I use llamacpp with --reasoning-budget; very simple way to move on.

redrix•4m ago
You’re absolutely right to be hopeful. Three honest possibilities, and I’ll be straight with you about each:

1. It overthinks — Just like the previous iteration. High confidence. 2. It doesn’t overthink — Improvement from the last model for your use case. Regression for others. 3. It sometimes overthinks — Best case all around. A feature, not an impairment.

One final thing worth mentioning: (I made myself irrationally angry writing this)

martinald•39m ago
FYI: nothing seems to be able to run this (easily) yet. llama.cpp, vllm etc I couldn't get working because of no support in the mainline version.
kzrdude•24m ago
They are giving pointers to how to run it now using for example https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next (and an especially provided vllm release).
amclennon•39m ago
It looks like this also undercuts the already absurdly inexpensive Deepseek Flash in pricing. Wild.
kzrdude•31m ago
Great. It's also smaller than DSV4 Flash, so it makes sense that way.
geooff_•29m ago
Where are you seeing that? At the bottom of this post from Qwen I see:

Qwen 3.8 flash: $0.16 / $0.47

Compared to

Deepseek 0723: $0.03 / $0.075

(units in USD/m tok)

twohaibei•15m ago
0.03 / 0.075 ? Where can i get that prices? Especially during peak hours DS4flash became much more money hungry than last month.

https://api-docs.deepseek.com/quick_start/pricing

ls_stats•11m ago
Deepseek 0732 is $0.22/$0.66 off peak
pram•39m ago
It's in Unsloth Desktop already. Looks like it's 73GB, so 128GB Mac or Strix Halo etc will work. Exciting!
cwizou•34m ago
Download is available, but likely need to wait for an update, I get this which is understandable with the architectural change :

Original error: llama.cpp does not support this GGUF's model architecture ('qwen4exp')

Edit : Saw the pull request, should arrive soon enough https://github.com/ggml-org/llama.cpp/pull/27742

andy99•7m ago
I only see a 1-bit quant posted on unsloth HF and it’s 72.5 GB. Is that what you mean? That’s much bigger than I expected. If you can’t run a 4 bit quant in on Strix Halo it becomes a lot less interesting. https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
dist-epoch•7m ago
73GB for the 1 bit model...
KolmogorovComp•34m ago
Will this be cheaper than DS4flash ?
armcat•29m ago
How is input token efficiency/verbosity on this model? Has anyone tried? GLM 5.2 was doing lot of turns and thinking piling up input tokens in the context (compared to Claude and GPT models). Then Qwen3.8-27B was 2x of that. Both delivered good output results but those cumulative input token costs were not cheap. Note this is on our specific business workloads. Genuinely interested in other people's experience (if you are able to try it out).
petu•20m ago
Haven't tried, would be surprised if it's any different.

It's new arch demo for future Qwen 4 family, but (as I understand) training recipe/data is same as any other 3.8 model.

a_humean•24m ago
Waiting for llama.cpp support to land, but this might be a big deal for Strix Halo users.

6B active params helps around the memory bandwidth constraints, but a 128GB box can probably run the Q3/Q4 quants fairly easily with a decent context size. This might actually be better for strix users than 27B, which was already very good.

Imustaskforhelp•15m ago
Pelican: https://gist.github.com/SerJaimeLannister/8fdef9c00175da0ca6...

Aside from the pelican, I am sort of impressed by the fact that things are going the way in terms of really impressive small models.

Also I love how this uses N-gram embedding. I think that Longcat was the first one who used it (I submitted that submission on hackernews because I really just loved the idea of it that I understood), I am certainly more interested in local LLM models and its interesting how they are utilizing new architectures to do some really impressive optimizations!

(Do note that I created it using a free rate limited end-point that I found on the huggingface space section: https://victor-chat-with-qwen3-8-flash-next.hf.space)

andai•13m ago
Father, I cannot scroll the website.
loclol101•7m ago
Definitely need to try this out locally.
dist-epoch•8m ago
World knowledge also means knowing the various algorithms and ways particular programming problems are solved.

You can't search what you don't even know exists.

gruez•46m ago
>Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

How much memory does this translate to and what quantization (if any) were applied?

rohansood15•34m ago
128GB, 4-bit quantized.
user43928•45m ago
For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen.

Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.

rohansood15•32m ago
This is a good counter argument. But you have to note that this is after OpenAI cut Luna costs by 80%. If you compare launch pricing, Qwen probably comes out ahead on a cost-performance basis.
jrflo•16m ago
The luna cost cuts were real though, not a one time promotion or something, due to some optimization (probably distillation?) that openai did.
QwenGlazer9000•12m ago
Was it?

Given the timing, I think they A. shat their pants since Deepseek flash just came out with insane pricing before the price hikes, and B. Anthropic is really struggling in model tiers below opus.

It was smart for them to cut prices regardless of whether they had 80% efficiency gains or not

dist-epoch•8m ago
It's a much bigger model, with a next-gen architecture. It's expected to be much better.