frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Why Claude's Watermarking Won't Fix Anything

https://www.plagiarismtoday.com/2026/08/13/why-claudes-watermarking-wont-fix-anything/
1•speckx•36s ago•0 comments

IC 1101 is the largest galaxy identified, so far, in the universe

https://disassociated.com/ic-1101-largest-galaxy-identified/
1•Brajeshwar•56s ago•0 comments

The Jobless Boom Has Arrived

https://www.wsj.com/economy/jobs/the-jobless-boom-has-arrived-41361a06
2•donsupreme•59s ago•0 comments

Dreams, Reflections, and Inceptions

https://papercompute.com/blog/dreams-on-paper/
2•eigenBasis•1m ago•0 comments

Reviving the Famicom Network System – Throaty Mumbo

https://www.youtube.com/watch?v=xiyCKUl93Uo
2•mrinfinite•2m ago•0 comments

Federal government finalizes ownership reporting exemption for US firms

https://www.reuters.com/legal/government/trump-administration-finalizes-ownership-reporting-exemp...
4•anigbrowl•3m ago•0 comments

Hypothetical Document Embeddings

https://arxiv.org/abs/2212.10496
2•kaycebasques•3m ago•0 comments

Ask HN: What are your unfinished projects?

2•adam_gyroscope•4m ago•0 comments

I bought the most expensive cable I could find. It still died

https://medium.com/@pokhts/i-bought-the-most-expensive-cable-i-could-and-it-still-died-welcome-to...
2•PhilYeh75•5m ago•0 comments

Show HN: Online SNMP MIB database - upload/view your own MIBs

https://mib-viewer.com/
2•beefstew123•6m ago•0 comments

"Solving a largely imaginary user goal"

https://unsung.aresluna.org/solving-a-largely-imaginary-user-goal/
3•euthymiclabs•7m ago•0 comments

Is AI Coding Effecting Our Mental Health? [video]

https://www.youtube.com/watch?v=iPUn1Fnfn0k
2•torartc•7m ago•0 comments

The Java Story – The official documentary [video]

https://www.youtube.com/watch?v=ZqGSg4b_cZA
2•kimi•8m ago•0 comments

Show HN: Eve Software Factory Template

https://github.com/vercel-labs/eve-software-factory-template
2•flashbrew•8m ago•0 comments

A Plane Ran Out of Fuel over the Atlantic. The Pilots Saved 306 Lives

https://www.popularmechanics.com/flight/a73413706/air-transat-flight-fuel-glider/
2•ilamont•10m ago•0 comments

Agentic engineering optimizes for rejecting output, not generating it

https://dparkmit.substack.com/p/how-i-used-agentic-engineering-to
2•dparkmit•11m ago•0 comments

Google Is Making Private AI Practical with Homomorphic Encryption

https://blog.google/security/how-google-is-making-private-ai-practical-with-homomorphic-encryption/
2•u1hcw9nx•13m ago•2 comments

Taylor Farms' connections to Trump admin spurs probe into Cyclospora response

https://arstechnica.com/health/2026/08/taylor-farms-connections-to-trump-admin-spurs-probe-into-c...
3•cratermoon•13m ago•0 comments

Qwen 3.8 27B Pelican

https://www.reddit.com/r/LocalLLaMA/comments/1voa3ch/comment/p3o0om9/
2•T0mSIlver•14m ago•0 comments

Activation Energy is a good model for a lot of things

https://homosabiens.substack.com/p/activation-energy-is-a-good-model
2•surprisetalk•16m ago•0 comments

The Other Sean Byrne Doesn't Exist

https://conic.al/writing/the-other-sean-byrne-doesnt-exist/
2•seanieb•17m ago•0 comments

Ed Zitron Was Wrong About AI

https://www.drjoshcsimmons.com/writing/ed-zitron-ai-predictions
3•joshcsimmons•18m ago•1 comments

Too Little, Too Late: Flock Admits Their Technology Needs Reforms

https://www.eff.org/deeplinks/2026/08/too-little-too-late-flock-admits-their-technology-needs-ref...
4•Brajeshwar•18m ago•1 comments

Agent Host Protocol

https://microsoft.github.io/agent-host-protocol/
2•eatmyshardz•18m ago•0 comments

Show HN: SVC16 = a specified virtual computer to build compilers for

https://janneuendorf.github.io/SVC16/
2•JanNeuendorf•20m ago•0 comments

Show HN: Remarc – contextual feedback for coding agents via MCP

https://github.com/metedata/Remarc
2•young_mete•21m ago•0 comments

Animica – a post-quantum Layer 1 blockchain built primarily in Python

https://animica.org/
3•animica•21m ago•0 comments

The Treasury market's toxic codependency

https://www.ft.com/content/e3cd352e-5202-4748-952f-ed623ccdc774
2•petethomas•21m ago•0 comments

Explain to Me in Simple Technical English

https://allaboutcoding.ghinda.com/explain-to-me-in-simple-technical-english/
2•speckx•22m ago•0 comments

Chery didn't exist in the UK last year but just outsold Honda and Mazda combined

https://www.carscoops.com/2026/08/chery-uk-market-share/
2•daau•22m ago•0 comments
Open in hackernews

Qwen3.8-27B

https://twitter.com/alibaba_qwen/status/2088280182356611304
155•mfiguiere•52m ago

Comments

chvid•48m ago
These are massive improvements - and something you can actually run on a laptop.
brcmthrowaway•45m ago
My Strix Halo is about to go overdrive!
tosh•44m ago
27b dense model at Opus 4.6 level

Opus at home

I hope there also will be a new ~10b variant

yassa9•23m ago
can you tell me ideas of usecases of 9 or 10B language models ? I cant find any usecases other than training a lora on them to give good bash commands for example
tosh•6m ago
they are all overlapping but:

categorization, information retrieval, semantic search, image description

also with the model as part of an agentic system with tool calling

scrlk•44m ago
Beats Opus 4.7 Max (w/ Claude Code) on DeepSWE (42.2 vs 40). Looks like Qwen's 27B models continue to pack some punch.

Unsloth's GGUF quants are up: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

edg5000•42m ago
That's crazy, considering the massive size difference. But the small Qwen models are known for punching above their weight.
nblgbg•41m ago
Is there any advantage to using the model from Unsloth compared with https://huggingface.co/Qwen/Qwen3.8-27B-FP8 ?
benxh•40m ago
Depends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus
petu•28m ago
Unsloth one is gguf for llama.cpp (and some other on-device engines).

So advantage is not having to produce your own quantisation / gguf from .safetensors you've linked.

4chandaily•27m ago
kristopolous•43m ago
q4km is about 48 tps on a 4090. my llama.cpp params are --flash-attn on --parallel 1 --load-mode mmap
m_ke•41m ago
With spec decode should easily get to >100tps

on my dual 3090s qwen 3.5 27b was running at around 110tps using the config from https://github.com/noonghunna/club-3090

make that 200tps on a single 5090, 4x faster than opus https://x.com/radixark/status/2088285681131110446

devs about to get handed a two 5090 box each and told to max that out

KronisLV•43m ago
I hope really badly that we'll get a new 35B A3B or similar MoE model!

I also miss the Qwen 3 Coder Next, which was 80B A3B, there are quite a few use cases where a non-dense model <100B would be the sweet spot (when you have the VRAM but not the TDP or compute power). Heck, I'd gladly take A5B or A8B or even A10B as a sort of middle ground.

Also alternate link for viewing the images without signing in: https://xcancel.com/Alibaba_Qwen/status/2088280182356611304

Casteil•37m ago
I'm hoping too that they'll put out some MoE variants.

Qwen3.5:122b:a10b can run about twice as fast as this 27b dense model.

peri-cl•35m ago
Same here! Qwen3.6-35B-A3B is the only local model I've found that runs reasonably on my iGPU. Looks like me and and my noisily-wheezing laptop will be sitting out this upgrade.
Alifatisk•16m ago
> I'd gladly take A5B or A8B or even A10B as a sort of middle ground.

Whats up with focusing on the active param count? Do yall fiddle with the weights or something?

martinald•12m ago
You can run these on CPUs at a somewhat reasonable speed.
altruios•42m ago
remember to let llama.cpp catch up to anything new in this model. Save your judgment until about 2 weeks of use.
chrismartin•22m ago
'Good' news, there seems to be nothing new architecture-wise. Same as Qwen 3.5 and 3.6, so llama.cpp doesn't know the difference.
anana_•40m ago
Monstrous benchmarks! Hoping it is not benchmaxxed.
kunver•40m ago
Welcome deepseek flash flash!
NorwegianDude•40m ago
If the benchmarks are a real indication, we now have a local model that is runnable on a high-end personal PC that trades blows with the leading model Claude Opus 4.6 Max from half a year ago.

Insane if that is the case. Downloading now!

kunver•39m ago
Looks like a pretty significant improvement on the DeepSWE benchmark compared to the previous 27B model.
alpha_trion•39m ago
NICE, i've been waiting for this drop, thanks for posting this
TomGarden•39m ago
Any tips on the best approach at running this at an M4 Max 128GB? Token throughput was a bit slow with the last 27B one (MLX), ended up using the A3B variant but if I could get this one to reasonable speed I'd much prefer it.
LoganDark•32m ago
Unfortunately, that chip just doesn't really have the memory bandwidth to run this (or nearly any) model at acceptable speeds. I have the exact same chip (M4 Max 128GB) and I've been trying to optimize a completely purpose-built implementation with Fable and this is just not possible. Even if you could reach the full 576GB/s, it's just physically impossible to exceed these numbers with the model's architecture:

2 bpw - ~85.7t/s

3 bpw - ~58.0t/s

4 bpw - ~43.9t/s

6 bpw - ~29.5t/s

8 bpw - ~22.2t/s

16 bpw - ~11.2t/s

without cheating. You'd have to exclude layers, skip operations, etc. basically do stuff the model wasn't trained for. And speed collapses so fast with context that even 2 bpw would be looking at ~37.6t/s after just 128K tokens.

MTP only improves the situation by up to 2x in the ideal case, while drastically reducing the performance floor. While optimizing a 9B model on this hardware, I've found that the GPU just doesn't have enough FLOPS to handle speculating more than one or two tokens ahead on a single stream, regardless of quant level, simply because of the arithmetic cost of the forward pass. The 27B model would be even more expensive than that, potentially such that it's already bottlenecked by the GPU itself rather than memory.

I wouldn't get my hopes up for the 35B-A3B either. Not only is it reportedly much less intelligent, but I hit the exact same 85t/s wall in practice (again with highly specialized inference).

Without speculation I can reach around 120t/s on Qwen3.5-9B and with n-gram speculation (not even MTP; this derivative didn't come with one) around about 150t/s on average. This is on the very very edge of what I'd consider acceptable for me to even consider using such a compact model. YMMV due to the silicon lottery but the situation isn't good.

brcmthrowaway•32m ago
Check out MTPLX and limit your context size.
jedbrooke•35m ago
I hope the bonsai team makes another 1bit quant of this model (or releases code/instructions on how to do it), using the Qwen3.6 27B on my 16GB mac mini has been wild . The 1bit quant feels like opus level… for the first couple turns. Then it has trouble eg switching from plan mode to act mode. This is mostly mitigated by starting a new session. (tbf this limitation is called out on the hf page)

I saw unsloth has 1bit quants too so I might check that out, anybody have experience with those?

tosh•35m ago
also cool: Qwen 3.8 27b is multi modal!
gurkwart•13m ago
strong visual reasoning apparently, which is nice. still lacking native audio however. hoping for more companies to embrace the spirit of something like `gemma-4-12b-qat` for actual multi-modality (text, image, video, audio).
WithinReason•34m ago
another thread: https://news.ycombinator.com/item?id=49294502
LeBit•32m ago
https://xcancel.com/Alibaba_Qwen/status/2088280182356611304
looksjjhg•21m ago
I could kiss you right now
brcmthrowaway•31m ago
This with ddg mcp to fill in world knowledge. Are local models the future when computer architectures catch up?
pu_pe•31m ago
Seems to be SOTA for its size. Hopefully independent benchmarks will come soon.
ramon156•27m ago
need another fable uncensored merge with 3.8, really curious what it can deliver
ThouYS•26m ago
3.6-27B on little-coder was already mind blowing. looking forward to this guy!
mickeyp•24m ago
Model benchmarks are useful, to a point, but it is the long tail of things you do with the model that determines if it's good at a wide range of activities. Ant/OAI, to their credit, build their models -- even the small ones -- so they follow instructions and do tool calling well, without the system prompts confusing them. This is especially important for long-horizon tool calling.

So one open weight model might "meet" Opus or whatever on benchmarks, but then fail to follow a simple answer format and also tool call correctly. The models are whipped to within an inch of their lives to strictly adhere to their post training quality gates.

hathym•24m ago
better than opus 4.6 max ╰(°□°)╯

  __        __   ___   __        __
  \ \      / /  / _ \  \ \      / /
   \ \ /\ / /  | | | |  \ \ /\ / / 
    \ V  V /   | |_| |   \ V  V /  
     \_/\_/     \___/     \_/\_/
yassa9•20m ago
Can anyone who has that specific personal test he tries on different models , and tries this model , to tell us here if possible , how good or bad is this new model ? compared to others ?

I only trust those users genuine personal tests

alyandon•17m ago
There is a down to earth guy on YT that performs a series of tests against LLMs running on non-god-tier commodity hardware. He will likely be testing this soon enough.

https://www.youtube.com/@lukesdevlab

I don't know if that is what you are looking for or not and as always your experiences may be different.

theanonymousone•15m ago
I'm wondering whether any provider can offer this for cheaper $/token than the new DSv4 Flash, which is both cheaper and smarter :/

Completely local use is a different story, of course.

ramon156•13m ago
People will claim it's not comparable to Opus despite it beating the score. I'm not sure I disagree, but I'm also unsure whether I care. Most new models nowadays are "good enough". I cannot complain because I'd rather spend that time improving my prompts and docs. Opus might be a _slight bit better_ at picking up vague hints, but it's also extremely expensive, and I hit the 5 hour limit way too quick.

I care a lot about speed and efficiency right now. For my setup I would like to have 2-3 different model families. I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting, Deepseek V4 Pro 0813 for developing, and Gemini flash lite (any recent cheap model) for repo scouting. I'll add another one in the mix for reviewing (in this case Gemini 3.7) and that's all I need.

I've tried most models except Grok.

Qwen is too expensive IMO (Alibaba Cloud subscriptions are hard to come by and I'm not spending 50 euros a month for a tool, so 18 euros it is). If it ever becomes efficient enough to run locally I will definitely look back.

Claude is slow and expensive (the cache hit prices are absurd).

OAI is pretty good, I might add it to my arsenal seeing how cheap it is.

These opinions change every day. Last week I would've never picked Deepseek until I read about the pricing. even post aug 16 it's worth it (although it's getting close to gemini pricing).

Right now my costs are 12 euros a month (z.ai) + whatever deepseek consumes. This typically isn't more than 8 euros a week. 44 euros a month and I have a setup that is doing pretty well.

hypfer•11m ago
> I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting

Dude, GLM-5.3 released _today_.

The phrasing "I've settled on" is incorrect for this context.

ramon156•10m ago
hence the "former deepseek v4 pro". I tried it out this morning and have had no complaints. I already liked glm 5.2
hypfer
filup•9m ago
https://news.ycombinator.com/item?id=48403639

my prediction was way too far out. 4.6 at home! Woo.

T0mSIlver•8m ago
Unsloth Q4_K_M on a single 3090, llama.cpp "Generate an SVG of a pelican riding a bicycle" first try https://www.reddit.com/r/LocalLLaMA/comments/1voa3ch/comment...
jlkivey•6m ago
Note: on the model card the comparison to Opus is Opus 4.6 Max, not 4.7
Run the unsloth if you are using llama.cpp (GGUF)

Run the one you linked if you are running vllm (safetensors)

WithinReason•32m ago
I wish each quant was benchmarked on the same tests as the original network so we could compare their performance
scrlk•7m ago
Unsloth publishes KL divergence numbers which measures how much the quantised probability distribution changes compared to the unquantised weights: https://unsloth.ai/docs/models/qwen3.8#quantization-analysis

It's a bit bare at the moment, I assume they are going to add further detail later, similar to their other releases.

NitpickLawyer•31m ago
> Beats Opus 4.7 Max

I'm a huge open model fan, and have used them since forever, even have daily drivers for on-prem dev, but no. They do not beat opus on real-world usage.

Qwen models are impressively good for what they are, are "good enough" for plenty tasks, can be ran locally on decently priced hardware, and so on. They certainly have their uses, and the field in general has advanced faster than my early expectations. But to compare a 27B model to SotA behemoths from a few months ago is doing everyone a disservice, especially people who pick it up, try to use them just like API models, and leave disappointed and confused. Number goes up on a benchmark isn't it.

KronisLV•24m ago
> ...but no. They do not beat opus on real-world usage.

I agree, but then we just need meaningful benchmarks that clearly show that! Otherwise it's hand waving about something that should be put on paper in quantifiable terms.

niek_pas•19m ago
A wise man once said, "not everything that counts can be counted, and not everything that can be counted counts".
bewareofscams•18m ago
Only useful benchmarks are those you (in particular) don't have access to.
xienze•14m ago
> but then we just need meaningful benchmarks that clearly show that!

That's the rub. AI benchmarks are IMO, by and large totally unreliable. We think of them as similar to traditional benchmarks of deterministic processes where the number of variables is low. But they're anything but that. Non-deterministic processes with an astounding number of variables and fuzzy acceptance criteria.

It leads to results like these, where if you take it at face value, the only conclusion you can draw is "wow Anthropic must be stupid if Opus takes 1T parameters to do what Qwen can do in 27B."

spmurrayzzz•22m ago
> They do not beat opus on real-world usage

We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios (mostly Rust, some C for microcontroller stuff). Qwen3.6-27B scores only 4% lower for pass@1, n=250 compared to Opus-4.8.

For the labeled dataset, the average PR size they're being measured against is around 1.5k SLOC.

This is very much "real-world usage" for us. The sort of change sets that come in daily/weekly and are solving non-trivial issues in the respective codebases.

As is usually the case, the most broad claims from both the labs and from the consequent pushback are talking past each other.

Foobar8568•11m ago
Considering the clusterfuck that is opus 5 or even fable, if Qwen 27B is trully better than Opus 4.7 Max, I will rejoice.
mft_•27m ago
Go for a slightly more quantised version, and experiment with different MTP settings. I find that MLX versions are marginally faster on my 64GB M1 Max, but I usually use Unsloth's GGUFs via llama.cpp as there's a much greater range of quants available and I prefer llama.cpp. MTP sometimes also helps a little, but I suspect it's less helpful on my system than others.

Unsloth: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

This might work for you, but I didn't get on very well with MTPLX when I tried it a while back; YMMV: https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized...

•
7m ago
The sentence still doesn't make sense, because "settled on" implies a long testing phase with a verdict eventually emerging out of that.

What you're currently doing is "testing out"