frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Ember-1

https://fireworks.ai/blog/ember-1
132•gmays•2h ago•66 comments

Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi

https://loficities.com/
36•safaelmali•51m ago•6 comments

In an $80 motel room, a discovery to shed light on the origins of life

https://www.nytimes.com/2026/09/26/science/motel-science-discovery.html
143•danso•5h ago•56 comments

Writing Efficient C++ Code (2013)

https://asawicki.info/articles/writing_efficient_cpp_code.php
112•ibobev•1d ago•48 comments

Replacing the old battery on rechargeable bike lights

https://jvns.ca/blog/2026/09/27/replacing-the-old-battery-on-rechargeable-bike-lights/
103•surprisetalk•6h ago•51 comments

Show HN: TinyAIArena watch AI agents battle it out

https://tinyaiarena.com/
60•hp6•3h ago•31 comments

Oral history of John Chowning, inventor of FM synthesis [video]

https://www.youtube.com/watch?v=e1Xn3030IvM
6•Rochus•1h ago•0 comments

The Normalization of Inexplicable Failures

https://www.ihatethefuture.com/2026/09/the-normalization-of-inexplicable.html
192•pxx•4h ago•69 comments

On caring for user data: NeoVim caused Vim undo files to be deleted

https://unsung.aresluna.org/they-had-no-concept-of-a-duty-of-care-to-their-users/
297•jandeboevrie•4h ago•256 comments

Fragment of oldest known peace treaty found in Turkey

https://www.livescience.com/archaeology/ancient-egyptians/we-have-found-traces-of-peace-thousands...
24•gmays•5h ago•5 comments

The Cartesian Hand: In-Hand Manipulation with All-Linear Fingers

https://generalroboticslab.com/cartesian_handv1
18•AareyBaba•1d ago•4 comments

Flip Fluid on Flip Dots

https://mitxela.com/projects/flipflip
315•blutack•1d ago•21 comments

Fakecloud: Local AWS cloud emulator for integration tests

https://fakecloud.dev/
75•theanonymousone•1d ago•41 comments

John Coltrane Centenary's – Impulse Records Release the Legendary Tiberi Tapes

https://www.jazzwise.com/content/news/john-coltrane-centenary-celebrations-see-impulse-records-re...
20•gregsadetsky•2d ago•5 comments

Show HN: Building a Markdown editor for Mac, iOS and web

https://www.markdown.beauty/
47•thiagoperes•5h ago•28 comments

There are no "rogue" AI agents

https://eoinhiggins.substack.com/p/there-are-no-rogue-ai-agents
240•zzzeek•3h ago•176 comments

Show HN: Trail – new kind of logic game

https://trail.franzai.com/
17•franze•12h ago•7 comments

Kicki: A DECsystem1060 – Interim Computer Museum

https://icm.museum/blog/?p=207
7•rbanffy•1d ago•1 comments

Video CDs Break Windows Explorer

https://clydesnotes.blogspot.com/2026/08/video-cds-break-windows-explorer.html
43•ClydeN•1d ago•14 comments

Faster prompt lookup drafting in llama.cpp

https://jadidbourbaki.github.io/blog/prompt-lookup-llama-cpp/
35•pptadversary•23h ago•3 comments

Improving site performance by shipping more CSS

https://github.blog/engineering/architecture-optimization/improving-site-performance-by-shipping-...
69•torutofu•1d ago•56 comments

Walgit: A Git server that is one binary in front of an object store

https://github.com/rgodha24/walgithub
58•handfuloflight•1d ago•7 comments

PostmarketOS is rebranding as Nura

https://nura.eco/blog/2026/09/27/nura-rename/
117•HotGarbage•4h ago•20 comments

Go Concurrency Distilled

https://antonz.org/go-concurrency-distilled/
348•chmaynard•1d ago•147 comments

C's Flexible Integer Sizes Were Not a Design Mistake

https://pikuma.com/blog/c-integer-sizes-not-a-mistake
50•ibobev•2d ago•76 comments

Ten lines of code that changed my world

https://pixelambacht.nl/2026/ten-lines-of-code/
100•dimonomid•6h ago•29 comments

PipePipe: NewPipe hard fork implementing SponsorBlock

https://github.com/InfinityLoop1308/PipePipe
477•Qision•2d ago•262 comments

Show HN: A CC0 museum of retro 3D tricks you can paste into a page

https://3d-retro.com/
59•SouthWestAtlas•3d ago•10 comments

Rusty thoughts on "Parse, don't validate"

https://eli.thegreenplace.net/2026/rusty-thoughts-on-parse-dont-validate/
60•ingve•10h ago•29 comments

Finally, A True Blue Rose Exists

https://www.sciencenews.org/article/true-blue-rose-pigment-copigment
86•bookofjoe•1d ago•33 comments
Open in hackernews

Ember-1

https://fireworks.ai/blog/ember-1
131•gmays•2h ago

Comments

andsoitis•1h ago
> The problem: thinking models think too much

Analysis paralysis stifles not just human intelligence, but other intelligences too.

AraneaDev•15m ago
Yes and thar makes you wonder if the Paradox of Choice would apply as well ;)

The more options you have, the harder it becomes to be satisfied with the one you picked.

minimaxir•8m ago
The thinking traces on some Chinese models just output the full response in the thinking trace, then output it again to the user, which is redundant.
monkey_monkey•1h ago
I don't think the article mentions Pareto frontier enough.

Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?

AnodicElegy•1h ago
I guess they figure "best bang for your buck" comes off a little too colloquial.
DonsDiscountGas•47m ago
They want it to be the best at something. And it's obviously not the absolute smartest. So here we are.
user43928•36m ago
Pareto frontier on some benchmark that I am hearing of for the first time.

Kimi K3 with less reasoning tokens isn't exactly exciting either, and particularly so if the license is less open than original Kimi K3.

erichocean•1h ago
Need this done for DeepSeek, ideally one of the Flash models.
atemerev•1h ago
If you have the compute, I have the expertise.
drob518•47m ago
And GLM. Both Deepseek 4.1 Flash and GLM 5.3 Flash are quote verbose when thinking.
tomrod•1h ago
Well done, and great iteration.

The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).

drob518•48m ago
Unfortunately, it’s hard to make a chart of that.
esafak•1h ago
It looks like it would be similar to GLM 5.3 Flash, had they tested it...
jamienk•1h ago
Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
intothemild•1h ago
Yes, absolutely, but only if people keep contributing in the open.
jack_pp•1h ago
not necessarily, just knowing something is possible will motivate others to achieve it somehow. Which is why there are so many LLMs and OAI doesn't have a monopoly
segmondy•1h ago
No, because close labs/models borrow but don't contribute back.
andsoitis•1h ago
> Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.

So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.

ls612•1h ago
On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
spijdar•1h ago
I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality.

That's just vibes, though.

KaoruAoiShiho•21m ago
https://www.reddit.com/r/LocalLLaMA/comments/1wj3s31/thank_y...
tdhz77•1h ago
Does anybody know if this would be a good model for creative writing?
intothemild•1h ago
So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?
netvarun•1h ago
Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.
Evidlo•49m ago
weight available
DonsDiscountGas•50m ago
It happens. Most open licenses aren't GPL style copyleft.
swagatkonchada•40m ago
It happens with open source software all the time, why would we expect any different with open source weights.
reactordev•37m ago
Because we do. The GPL isn't a suggestion. If you can take open source code and make private software out of it then what are we all doing? No, license requirements and agreement are law for a reason.
bloggie•
netvarun•1h ago
Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15
drob518•49m ago
Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost.
nostrebored•41m ago
Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design
pornel•4m ago
Competition is good. Without K3/GLM/DS4 etc. there probably would be no reason for OpenAI to drop Sol's price.
logicallee•55m ago
This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.
nostrebored•40m ago
I can’t help but think it’s more expensive tinker.
dbuxton•53m ago
Do they mean Opus 5.5 or Opus 5?
nico•51m ago
> The problem: thinking models think too much

This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking

It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications

demibabs•42m ago
What are the useful applications of Jev so far? Not to sound dismissive, I just haven’t seen what people are using it for yet.
neosat•36m ago
Lots of use cases! I've personally used it for the following:

1. Evals (once you have your rubric defined and tuned using a reasoning model, jev can be great for running periodic evals especially those that run daily.

2. e-commerce catalog classification 3. quick search using anything as context and query mapping to a pre-defined set.

elcomet•24m ago
Why not using a cheap LLM with thinking completely disabled ? I don't think it will be much more expensive than jev.
nico•11m ago
I’ve tested this with some local LLMs and their accuracy is in general better than Jev/Laya, but they are super slow in comparison as well

For example, a typical/stock LLM can’t really play Doom in real time, but a Jev-like model can. Just because of latency

Of course, if you want the best Doom player, there are way better and faster adhoc models

themgt•51m ago
The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.

"Pareto": 8 hits

"Opus 5.5": zero hits

wmf•44m ago
Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.
GodelNumbering•35m ago
This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
shriphani•28m ago
what hardware are you using to train?
GodelNumbering•25m ago
I didn't have a local GPU, so I asked it to go out and find hardware. It found a google TPU v6e which seemed reasonably priced. I gave it my google api key. I told it to use TPU only when training and bring it down afterwards. That's about it.
shriphani•20m ago
Neat!
otterley•17m ago
What kind of observability did you have over this process? I’m interested in how my peers are operating these efforts.
GodelNumbering•6m ago
On the cloud side, nothing valuable existed, so the training couldn't ruin anything it didn't create. On the laptop side, I usually ask the agents to create named scripts for everything it needs to access, then those local script directory is green-lit with approve all. For cost, I kept giving it new budget in the 20-30 dollar increments.

I had to intervene a few times. For instance, as smart as the models are said to be (Astra), it would copy the full training run, train on the server, pull every checkpoint to the local machine, then run tests, update. So, the bandwidth bill was as high as training bill for the first 6 hours. It could have simply tested each checkpoint on the server, saved time and money, didn't occur to it until I said.

Arcuru•9m ago
Over on /r/LocalLLaMA there's a group that's been getting popular doing the same thing for the Qwen 27B (and other) models. - https://huggingface.co/ukisai
jamienk•56m ago
Diff people have diff motives to experiment, then new work is done on top of stuff that "hits" in a way no one anticipated. Then work gets piled on top in a way that might make it hard to port
andsoitis•54m ago
> then new work is done on top of stuff that "hits" in a way no one anticipated.

Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.

Greatness cannot be planned.

mirekrusin•6m ago
I think people make mistake here, google’s approach is not to spend $2.3 on every $1.0 earned, they’re riding on serving to masses “luna”, they absolutely have way more powerful models internally but they don’t clutter their infrastructure with fragile and costly intelligence/size frontier. I think “underdog” perception is illusory/temporary, not stupidity, but calculated, conscious, longer term bet.
zeroq•39m ago
The difference between contributing to OS and AI, is that the first is a hobby alternative to woodworking or hiking, while the other can easily bootstrap you a company you can get millions in investment, at least for time being.
swagatkonchada•38m ago
Won't the "frontier" labs figure out whatever techniques were used and apply them to their closed models?
k__•30m ago
If they can keep up.

The lock-in is less pronounced as it is with AWS or MS.

cyanydeez•28m ago
Like how the last 2 decades of tech companies are thinly veiled open source pilfering into business units.
21m ago
Kimi K3 has its own license which is permissive, it isn't at all like GPL https://github.com/MoonshotAI/Kimi-K3/blob/main/LICENSE
otterley•15m ago
Because the licenses that apply to software make no sense in the context of LLMs. With the latter, there is no source code to license.

The words of a license are what the license is.

makeramen•38m ago
Aren't Cursor Composer models like this too? At some point all the extra RL you do can be considered as proprietary information added.

Not suggesting this is right or wrong, but is sort of the nature of the technology.

kingstnap•34m ago
There is little to no point reading the article as well. It's stripped of all alpha.

> task and environment feedback

> on-policy planning and learning

> feedback connects decisions to their consequences

These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.

Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.

PEe9bB7D•27m ago
i also need more info!
GodelNumbering•24m ago
I am thinking about opensourcing everything, although this is not my main domain or my main startup, so the overhead of huggingface etc seems a bit unnecessary
jack_pp•7m ago
just ask the agent to write it up if you don't have time to do a write-up yourself
luisfmh•3m ago
Curious about how you generated the training data? Was it just asking an existing model to generate a bunch of examples?

I ask cause would this be a kind of model distillation?

I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.