frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Meta Muse Glimmer – open weights 30B local coding model

https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
158•riordan•1h ago•57 comments

Docker Sandboxes – Disposable, isolated sandboxes for AI agents

https://www.docker.com/products/docker-sandboxes/
279•etoxin•5h ago•160 comments

What Happened to HackerOne?

https://blog.teknogeek.io/posts/what-happened-to-hackerone/
265•hipparchus•9h ago•129 comments

Run Android ARM64 VR APKs on Apple Vision Pro

https://github.com/shinyquagsire23/Klepton
105•LorenDB•8h ago•11 comments

An Interesting Fourier Transform – 1/F Noise

https://www.dsprelated.com/showarticle/40.php
57•q7m•3d ago•13 comments

Show HN: Voice driven murder mystery, Interview AI suspects with your voice

https://www.whodunnitai.com/
112•MrRowTheBoat•8h ago•44 comments

Tail-Call Interpreters in Rust – Jimmy Ostler

https://lordgoati.us/blog/tail-call/
33•amatheus•2d ago•9 comments

How I use LLMs to learn complex topics

https://laurentiugabriel.github.io/blog/articles/how-i-use-llms-to-learn/
677•laurentiurad•16h ago•423 comments

Taxi drivers rarely die of Alzheimer's

https://theconversation.com/taxi-drivers-rarely-die-of-alzheimers-how-complex-mental-maps-and-spa...
307•jader201•20h ago•217 comments

COLDCARD's Random Numbers Weren't

https://coldcard.rip/
4•bialyalibaba•4d ago•0 comments

How We Pushed CDC into Postgres

https://www.snowflake.com/en/blog/engineering/postgres-to-snowflake-replication-mirroring/
102•craigkerstiens•10h ago•14 comments

How Blackwing Pencils are Made [video]

https://www.youtube.com/watch?v=fow-LsdaH2E
19•NaOH•4d ago•3 comments

Ask HN: What are you working on? (August 2026)

263•david927•18h ago•906 comments

Turn satellite imagery into a paper globe you fold yourself

https://foldingglobes.com/
68•dango2506•8h ago•22 comments

Cool URIs Don't Change (1998)

https://www.w3.org/Provider/Style/URI
249•Klaster_1•21h ago•61 comments

An alias-based formulation of the borrow checker (2018)

https://smallcultfollowing.com/babysteps/blog/2018/04/27/an-alias-based-formulation-of-the-borrow...
12•parksb•2d ago•1 comments

ATProto for Distributed Systems Engineers

https://atproto.com/articles/atproto-for-distsys-engineers
93•LelouBil•3d ago•17 comments

Picophysics: Single file physics for games on platforms like N64, PSX, DC

https://gitlab.com/Kazade/picophysics
63•klaussilveira•4d ago•20 comments

Tuxedo No. 2 – Cocktail recipes

https://tuxedono2.com
103•smartmic•14h ago•34 comments

Nearest Pint

https://knowwhereconsulting.co.uk/maps/pubs/
37•bookofjoe•5d ago•22 comments

Show HN: I made alchemical-cosmological PCB badges

https://github.com/KaiPereira/Alchemical-Cosmological-Badges
15•kaipereira•4d ago•1 comments

Everything you do is being recorded

https://www.theatlantic.com/technology/2026/05/ai-wearable-surveillance-countermeasures/687203/
328•ike_usawa•1d ago•278 comments

Auto mode is now the default in Claude Code

https://claude.com/blog/auto-mode-default-in-claude-code
232•sbehere•7h ago•236 comments

I made tinnitus my friend, then it disappeared [video]

https://mynoise.net/vlog.php?ep=20260803
160•gregsadetsky•17h ago•135 comments

New Zealand lost its music media, and what we're building to replace it

https://propelmusic.co.nz/articles/the-sound-went-quiet-nz-music-media
119•berghoffer•15h ago•77 comments

The Ambition Project

https://www.betonit.ai/p/the-ambition-project
50•herbertl•12h ago•5 comments

Slap ROM Patcher

https://nyuu.page/projects/slap/
20•apsec112•3d ago•12 comments

OpenChamber: An Agentic Development Environment

https://openchamber.dev/
158•hexomancer•18h ago•76 comments

"The Persian MâR-Nâmeh Or, the Book for Taking Omens from Snakes" (1892)

https://publicdomainreview.org/collection/marnameh/
64•Thevet•3d ago•12 comments

Windows 11's built-in Weather app wastes more than 1 GB of RAM

https://www.notebookcheck.net/Windows-11-s-built-in-Weather-app-wastes-more-than-1-GB-of-RAM.1364...
571•akyuu•20h ago•509 comments
Open in hackernews

Meta Muse Glimmer – open weights 30B local coding model

https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
154•riordan•1h ago

Comments

tosh•1h ago
good to see new open weights releases from meta
jauntywundrkind•59m ago
good looking showing too, which is excellent.
InfiniteLoup•50m ago
The least they could do, after ruthlessly bombarding my employer's servers with requests, ignoring the robots.txt, scraping everything, and incurring significant Google Maps costs for us in the process.
gunalx•58m ago
Meta did not abandon opensource. I would love to see a smaller distill, or a moe of this size but the benchmarks seems competetive as long as it isnt benchmaxed witch i would not be suprosed if it is.
ignoramous•11m ago
> Meta did not abandon opensource

Open weights*

I don't think outside of the Big 3 (Ant, OAI, GDM), given the strong competition from China, any other Lab has a chance at capturing the coding market if they aren't open weights (save for xAI whose latest Grok looks every bit good & will probably rely on Cursor for distribution instead of going open weights). There's literally no other selling point, as the capabilities have mostly converged by now among the chasing pack.

scrlk•55m ago
Will be interesting to see how Qwen3.8 27B compares against this once it releases this week. Seems like dense 30B is back in fashion?

EDIT: An open weight version of Muse Spark 1.2 is going to be released as well:

https://x.com/alexandr_wang/status/2086756152034066792

https://xcancel.com/alexandr_wang/status/2086756152034066792

wronglebowski•53m ago
It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.
lostmsu•23m ago
It seems worse than 3.6, but a bit smaller.
IsTom•2m ago
How is 30B smaller than 27B?
Gecko4072•21m ago
Makes me feel hopeful. Things felt more positive around the llama 3 era. Now it’s like a dark, dreadful race.
karimf•13m ago
Yes, and also waiting for the next iteration of Gemma. Muse or Qwen are optimized for coding, while IMO Gemma is still better for non-coding tasks.

https://x.com/osanseviero/status/2086107547535122767

richardfey•54m ago
Looking forward to giving this a try with llama.cpp. I’m watching the open-weights competition with high expectations.
Gecko4072•53m ago
What I think would be perfect is a model that could run on a single DGX spark and be competitive with DSV4 Flash 731. Flash is already a game changer. Hopefully meta plans on this, like the old 70b. V4 flash is smart enough for any use but slightly too big. 27b-30b isn’t intelligent enough.
127•44m ago
DSV4 Flash 0731 already runs on RTX 4090 24GB + 128GB system RAM at a usable tok/s and quantization.
Gecko4072•39m ago
You personally? Just curious. Context window is also a factor and ram isn’t really cheap. Sparks are assembled units which I like.
cmrdporcupine•30m ago
This model I think will be too slow for that on Spark, even at 4 bit quant.

It's a dense model, not MoE like e.g. Qwen 35b or Gemma 4 26B A4B. On a Spark it will be memory bandwidth limited

I haven't tried yet (working on it) but back of the napkin estimate puts it at around 15tok/s even after converting to NVFP4. Prefill would be much higher though. That 15tok/sec is pretty typical for dense models of this size:

NVFP4 Q/K/V/O and MLP projections: ~13 GB/token

BF16 attention gates: ~3 GB/token

BF16 LM head: ~2.5 GB/token

Total: ~18.9 GB/token

At 273 GB/s, that gives a bandwidth-only ceiling of about 14.5 tok/s; actual performance would be lower.

sajithdilshan•49m ago
Still needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.
Gecko4072•42m ago
There have been discussions on language specific not really being a relevant change to reduce size.
Manfrednotfunny•29m ago
I would love to see any good research projects about it but i have the feeling that Frontier with MoE is making too fast of a progress so that a customized model would always be worse and that the MoE part is actually going somehow in this direction.

On the other hand, at the GTC was a talk about coding in different lanugage (like spanish) and explaining that the quality between spanish and english is relevant different.

But i have not found a good article about the impact of learning data with practical experiments or even if the order of the learning data matters.

At least I think i remember that Meta mentioned having better and less data can be better than more data with lower quality.

As long as these models can explain to you facts about any other topics, its still overfitted for the task though.

solarkraft•42m ago
I feel like we’ve had this discussion before. From what I remember, specialized models rarely do that much better than general ones, hence no mode Codex models.
sparkling•41m ago
solarkraft•45m ago
Wow, Meta is back (at least for now)!

I like this class of model. Multi-token prediction makes it viable to run dense models at not-too-far-off speeds as MoE models with much better intelligence.

The submission’s title (open weights 30B local coding model) is luckily wrong: This is meant to be a general agentic model.

It even comes pre-quantized and with a MTP/drafter model. Looking good!

Let’s hope they aren’t dishonest with the benchmarks this time …

Havoc•44m ago
The favourable comparisons to Gemma 4 and qwen3.6 look promising!
cmrdporcupine•17m ago
Those two offer MoE variants, this doesn't seem to.

Dense model makes it dog slow on anything without HBM. Max 15tok/sec on decode on DDR5 systems like a Spark or a Strix Halo -- and that's at 4 bit quant.

petu•7m ago
3090/4090 probably would do 40 t/s, for 5090 75 t/s is shown in the blog.
nutjob2•44m ago
The more open weight models get released the greater the market for personal and small business oriented hardware to run these models. This will drive lower cost hardware, which has stagnated in recent years due to most software not needing the performance and capacity.
cmrdporcupine•5m ago
The opposite happening because foundries are full to capacity making higher margin stuff.
jkwang•42m ago
The memory math is the part I keep re-reading. 4-bit gets the LM under 20GB, but they're explicitly budgeting the KV cache, the perception encoder, and the DFlash drafter into the same 24GB envelope. On a 4090 that leaves maybe 3-4GB of KV once the drafter and encoder are resident, so long agent traces are going to spill or truncate — which is exactly the workload this model is trained for. I'd like to see the K-Quant-17GB numbers reported at 64k+ context, not just conversational length.

"Minimal to no degradation on agentic tasks" from quantization is also a strong claim. In my experience 4-bit shows up first in tool-call schema adherence — malformed JSON args, wrong enum values — before it moves benchmark averages. Does the report break down tau-Bench / MCP-Atlas per quant level?

zmmmmm•39m ago
Meta knows how to win back developer's hearts .... let's see if they have the goods
xandrius•35m ago
If there is anything meta can do to regain hearts other than owning up their evil deeds, radically change their business model and paying up for taxes and damages, then the world is truly fucked and corporations will continue to win.
maxignol•33m ago
Optimizing speed is really the way to go. Yet 24GB is not what everyone can afford. Maybe we could take some of those 56tk/s and transfer into some free RAM space using MoE loading ? I'd be glad with a less than 10GB and more than 6tk/s model.
Manfrednotfunny•25m ago
I don't thinnk just MoE will solve it. If you hit constantly different expert layers, you can't outsource layers efficently and have to swap it in.

MoE will be faster because it will read less memory for sure, you still have to have it though.

lisplist•24m ago
Unfortunately this is just the entry price for LLMs. With the exception of the Qwen 27B models, I personally haven’t found a ton of use cases for models less than 200B. With the right setup, fine tuning, etc, you can make small models do cool things, but hard to please everyone given the insane hardware costs at the moment and the comparably cheap API costs.
bronxbomber92•14m ago
I wish they would release the quantized versions in a safetensor format. Many frameworks can't load PTE and GGUF.
_ache_•12m ago
It is interesting but it does look like a careful distillation of (Spark and) biggers open-weight models.

The progress compared to Qwen3.6 27B is good, not that impressive, it's a 4 months old model. (kuto to them to compare to 27B dense and not 35B MoE, it's more fair to do so). It is very probable that Qwen3.8 27B will crush Glimmer-30B on most benchmarks.

petcat•10m ago
As an industry, I wish we would stop calling these things "open weight" because it is too easy to confuse with actual "open source", which they are not.

Photoshop source code+ OSI license = open source

Photoshop binary you can run on your own computer = open weight

Photoshop SaaS web app = closed, proprietary (Opus, GPT, etc.)

"Open weight" models are still just binary blobs that are completely inscrutable. It's like bringing home a dog from the rescue and just hoping that it doesn't have a tendency to bite kids in the face. You just can't know. The only thing that you can do is try to add more training (fine tuning) telling it not to bite kids.

I don't think the FOSS community has ever accepted this, but somehow we're feeling like it is okay now.

piker•8m ago
It is useful to indicate you can run the weights on your own hardware. That’s categorically different from most other commercial offerings. It’s as if your adobe example ignores the reality that would exist had photoshop been invented in 2019: cloud only.
microtonal•8m ago
Photoshop source code+ OSI license = open source

Photoshop binary you can run on your own computer = open weight

I don't think this is a correct analogy. You are not allowed to distribute modified versions of the Photoshop binary. Most open weight model licenses allow you to make and distribute your own finetunes, etc.

monster_truck•8m ago
This analogy is terrible and seems to be extremely misinformed about how rescues evaluate dogs before they are put up for adoption
petcat
pu_pe•12m ago
Based on the benchmarks, it seems that Muse Glimmer barely edges out against Qwen3.6 27B, except for tool-calling skills (MCP, etc.). I wouldn't be surprised if they released it now because they are afraid they wouldn't beat Qwen3.8 27B.
Even if you had a 64GB machine: Are you willing to reserve 90% of your memory to run a LLM? With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no.
zoobab•32m ago
"With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no."

When it's free, you are the product.

IMTDb•23m ago
Deepseek flash is open weight, this means we can download and run that model without any connection to deepseek, no data/tokens/usage data ever reaches them. They cannot make us their product.
prplxd_nihilist•22m ago
I see many people saying deepseek and other chinese providers have always been profitable. Also they show their training costs publicly. Can't say for sure since I have not used it personally, but I think they'll for sure outlive the western SOTAs.
Manfrednotfunny•26m ago
I'm waiting for the speed/quality per dollar metric to go down a little bit further and then I will def run it at home.

Its not just that you send a sentence to an API endpoint, you always send EVERYTHING to that agent as a context.

You want to analyse your spending history? You now send everything to someone.

Either no one cares but understands this implication on how easy it is to really capture you or no one really things about it.

But i'm a lot more diligent on what I send. I disabled the gemini activity feature for example because google started telling me that my stuff could be reviwed by humans.

halJordan•21m ago
It's the size of a big vm. There's nothing wrong with reserving that much working space for one item.
mihaelm•27m ago
I'm sooo happy I pulled the trigger on upgrading and getting a new laptop (with 64 GB RAM) last summer. Feels like it was just in time before the exponential price jumps.
ishtanbul•24m ago
Pulled the trigger?
idiotsecant•19m ago
Common phrase.
karolist•14m ago
Parent used "pulled the plug", are you saying it's applicable here and not "pulled the trigger" like suggested?
Hinrik•13m ago
That commenter you're replying to knows that. The original commenter before them wrote "pulled the plug" which is different and doesn't quite apply here (actually implies the opposite of what they meant to say).
mihaelm•6m ago
lol, you're right, the brainfart completely changes the meaning.

I corrected it.

mettamage•21m ago
Bought an M1 64 GB for 2000 euro’s second hand a year ago. That was sweet
karolist•13m ago
paid 2.7k € for this same build new in Dec 2023, that was also sweet (still is)
karimf•16m ago
Practically ~20GB with KV cache

> We quantize weights to ~4-bit, bringing the LM under 20 GB. We validated minimal to no degradation on agentic tasks under compression.

https://www.reddit.com/r/LocalLLaMA/comments/1vkgsum/introdu...

dbbk•2m ago
Well if you're spending thousands on API tokens already, you could just drop the same amount on a 128GB MacBook Pro and that's a one time cost.
•
7m ago
I am extremely well aware of how rescues evaluate dogs. And I'm also fully aware that they do not know the full history of the dog. They go through a limited set of testing and interrogation to evaluate the safety of the dog. That's it.
QuadmasterXLII•5m ago
Given an open weights model trained to sometimes bite kids, we can’t train it to not bite kids, even though billions of dollars of research have been thrown at this open problem.

Given an open weights model trained to never bite kids, you can get it to bite kids with 10 prompts and a linear projection, the known simple algorithm doesn’t even need a backwards pass.

yay asymmetry!