frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Using an open model feels surprisingly good

https://matthewsaltz.com/blog/using-an-open-model-feels-surprisingly-good/
63•msaltz•1h ago

Comments

irishcoffee•59m ago
It surprises me that this concept took as long as it did to gain traction in… hacker news. 15 years ago folks here were compiling kernels and gentoo distros. Lately it’s been “you should just pay the man, it’s cheaper than running these things yourself”
brcmthrowaway•53m ago
Rsync the dropbox, dude
loire280•44m ago
Rolling your own Linux is nearly free and could be done on any computer you had lying around. Dropping >$5k on a computer to run a local model (badly) doesn't really scratch my "hacker" itch. The author of this article works for an AI infrastructure company and ran Kimi K3 on their infrastructure - this post reads like marketing.
eru•39m ago
It's fairly easy (and cheap!) to run your own email server, but most people don't bother and just use gmail. And rightly so!
khimaros•49m ago
this seems to be an advertisement for Modal. is it really your own infrastructure if the hardware is leased?
jdw64•47m ago
Actually, if you break tasks down into small enough units, there's not much difference between open models and closed models. The only reason to use agents is the hope that they'll work with natural language input—but that's where the difficulty lies.

For example, if you modify things at the level of small functions, open models seem to perform just as wel

paradox460•39m ago
Yup. Been using big expensive models like opus, gpt, or glm for planning, then switch to something like gpt-oss-120B on cerebras and watch it fly. Total cost will be less than $5 for even very large tasks
jdw64•20m ago
People tend to think alike. I work in much the same way.First, I write the draft myself, then let AI handle the revisions, and finally run it through an open model in small units before I check the results.
verdverm•14m ago
I think this is where they say "loop engineering" / orchestration comes in, though I would also include context/harness engineering in the toolbox
m_ke•46m ago
GLM 5.2 feels better than Opus and K3 is as good as Fable.

Now I can't wait for someone to distill K3 into a Qwen 3.6 27b or Poolside S 2.1 sized models for a proper fast local Composer 2.5 replacement.

arjie•42m ago
To be honest, I am surprised by how good DeepSeek V4 Flash is. I use Claude Code and Codex on Claude 5 Opus and GPT-5.6 Sol most of the time, but when I use DS V4 Flash I don't feel like it's really that bad. And with oh-my-pi and just plain pi it's pretty good. To be honest the frontier models are much better at tool calling so in an assistant flow they're better but I did the dumb thing and optimized the harness for the model instead, rewriting the tools so they match what it guesses at, and DS V4 Flash does just fine.

The TTFT and tok/s are much higher on the small model so that makes it competitive for a bunch of things. It feels like what old Sonnet used to by the end of last year which is honestly damned good.

pimeys•29m ago
I wrote my own OpenClaw one weekend and I am running it as my assistant through Matrix with DeepSeek v4 Flash (and Qwen). It probably costs me about 2 dollars a month and is even more useful than ChatGPT would be due to me having full control on what tools it has access to.

I can do things like take a photo of a doctor's note among add the appointment to my calendar, send a PDF to my archive tagged, OCR'd etc, search info from internet, look data from Google Maps.

I can even integrate this to home assistant and talk to my agent with my open source Alexa-like system.

mayank•21m ago
Oddly enough, I did the exact same thing this past weekend, for a couple hundred usd in Fable overage.

It was truly remarkably easy to build and package it exactly to my whims, in my case as a single docker container with a process reaper that runs llama, my Go code, tts, chat harness, browser in xvfb, and even a mailer daemon. With Gemma, it even runs on a RPi 5.

What’s absolutely wild to me is that over a couple hours, I could probably have it import parts of Home Assistant for my devices directly, and do other wacky stuff in what is essentially software for one.

arjie
loufe•42m ago
This is only a thinly veiled ad. It's fine, I was curious about this exact setup, anyways.

What would be useful is a cost metric. I'm curious how much I'd be willing to spend as a premium to not have those companies piping my conversations directly to the NSA. Maybe only some conversations? Claude and OpenAI are heavily subsidized, by all accounts, so Kimi K3 on a private endpoint might end up costing more or less - that's what I want to know.

mayank•30m ago
I was curious about the cost angle too, i.e. how much "free" coding agent I can get for what cost. Here's the research by Fable if you're interested: https://claude.ai/public/artifacts/2c9a5001-0b7e-4944-beb1-9...
nujabe•41m ago
This is such a poor quality post, reads more like a diary entry than a substantive technical post.
PontifexCipher•35m ago
https://en.wikipedia.org/wiki/Blog
sroerick•20m ago
lol
nujabe•19m ago
??
Gigachad•22m ago
I will take 5000 "diary entry" blog posts over LLM generated walls of text that say nothing.

People need to learn to just post the prompt rather than the LLM output which just fluffs the prompt.

sudo_cowsay•38m ago
I love OpenCode and the blankness of it too. Clean, light, and manual. It's a good change of pace from what we are used to on the internet (very cool but slow). Also, can you hmu with that modal plan XD
roywiggins•31m ago
Pi Agent is even leaner feeling, it's worth a try.
wps•27m ago
The reason Claude code is so popular is because it’s really good at taking super vague human prose “Claude build me a million dollar SaaS”-type prompts and spitting out thousands of lines of code which cover tons of surface-level edge cases, build in tons of functionality, etc

The smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with it. If you’re using it as an aid to traditional software dev, iterating on small, targeted functions, GLM works as good if not better than Claude. Anthropic expects low quality prompts. If you rubber duck GLM, you get absolutely pristine output in most cases.

Gigachad•24m ago
Personally I've found Deepseek v4 Flash to be as useful to me as Opus. But I don't do these silly one shot tech demos. I have the technical understanding to ask for exactly the change I want with the right terminology. I loaded up some credit on openrouter and it took me ages to hit $1 in spend.
wps•18m ago
I’m the exact same way. $15 on openrouter lasted me so long it would’ve got me fired at FAANG. Despite this, my number of commits is dramatically higher. Small, beautifully scoped changes is just good software development, and good for the wallet as well.

I think the issue is that no one is content with incremental progress. We all know one shots are mostly possible, so the age of the personal project is kind of over. There’s no motive to invest dozens of hours getting a working prototype when Claude can give you something right now. So you can’t make small changes until you have that codebase in place already. You’re forced to make sweeping changes if you use AI from the beginning. And it’s not like it matters, there’s no personal attachment to any one part of the code, it’s not even seen!

Gigachad•11m ago
spicyusername•19m ago
Something I've been thinking about a lot is that this new technology's primary UI is natural language. Human's are REALLY primed for natural language.

If it sounds good it must be good. The code this new model writes is incredible! It told me so!

•
7m ago
Haha! Love to hear it. That's exactly what I have too: a claw-like[0] system which I originally used Sonnet for and now use my DeepSeek v4 Flash with.

In my case, I was foolish enough to run it all on my hardware which is pretty damned fast but has a duty cycle of 5% and runs idle most of the time. It's definitely better done via API.

Tell me more about the Alexa-like system! I have mine at home on a custom OpenWakeWord model trained on the word 'Aurora'. And I, too, have it hooked up to Home Assistant. I have a bunch of Eufy E21 baby cameras mounted on the walls[1] so it has vision through the house. The vision model uses GPT-5.5 on the subscription because I haven't yet set up a Qwen or Gemma multimodal that can see.

With Frigate on my home server I can even watch for events like my daughter waking up! And at night the agent sends out the vacuum if we've cleared the floor of baby toys. I feel this close to the dream of sci-fi AI. Because I auto-forward a bunch of my email to it etc. it knows about what's going on with me.

My wife will sometimes ask the agent information about me etc. and it's way easier to get a fast response rather than waiting for me to see the message etc.[2]

0: https://news.ycombinator.com/item?id=47538158

1: https://wiki.roshangeorge.dev/w/Blog/2025-12-01/Grounding_Yo...

2: https://imgur.com/a/tP12lfj

I can see why Anthropic is freaking out right now. China has undermined their whole business. AI models will be basic commodities where hosting providers earn a tiny margin over the raw costs rather than the predicted fortunes from being the gatekeepers to the technology.
basch•4m ago
the writing should have been on the wall when meta did it with llama. it was very apparent that one or two entities could release a "good enough" product into the wild commoditizing the frontier of yesterday.
drnick1•8m ago
Frontier models like Opus 5 are also extremely good a "research" in an academic sense, including combining elements from different fields or literatures and creating something quite novel sometimes, publishable even, from vague/speculative prompts. This is on top of implementing known methods in about 1/100th of the time it would take by hand, allowing for fast exploration of ideas. They can also roll things out from scratch, removing essentially all third party dependencies from your code base, other than Numpy or Scipy. This really buys you a lot of clarity and/or performance sometimes.

Our position on open-weights models

https://www.anthropic.com/news/position-open-weights-models
582•surprisetalk•5h ago•820 comments

Using an open model feels surprisingly good

https://matthewsaltz.com/blog/using-an-open-model-feels-surprisingly-good/
65•msaltz•1h ago•32 comments

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

https://fermisense.com/when-machines-take-the-wheel/
41•ilreb•1h ago•7 comments

Benchmarking Opus 5 on SlopCodeBench

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarki...
178•dhorthy•5h ago•41 comments

Astronauts describe persistent 'observer' sensation after 6 month missions

https://spacedaily.com/sd-v-astronauts-returning-from-six-month-missions-describe-a-persistent-ob...
119•zdw•4h ago•64 comments

RTX 2080 Ti Memory Upgrade to 22 GB

https://gpusolutions.net/rbservices/graphics-card-upgrade/
51•wslh•3d ago•25 comments

Residential Proxies Are a National Security Threat

https://jacob.gold/posts/residential-proxies-are-a-national-security-threat/
20•joshbetz•2h ago•13 comments

EYG: A Programming Language for Humans

https://crowdhailer.me/2026-06-08/a-programming-language-for-humans/
25•crowdhailer•1h ago•9 comments

DConf 2026 in London

https://dconf.org/2026/index.html
61•teleforce•4h ago•26 comments

Watching Go's new garbage collector move through the heap

https://theconsensus.dev/p/2026/07/19/observing-gos-garbage-collector-old-and-new.html
182•matheusmoreira•2d ago•17 comments

C/C++ projects packaged for Zig

https://github.com/allyourcodebase
37•jcbhmr•4h ago•23 comments

Show HN: Yap – OSS on-device voice dictation for macOS with no model to download

https://github.com/FrigadeHQ/yap
32•pancomplex•9h ago•6 comments

An Uncomplicated Man

https://www.lrb.co.uk/the-paper/v48/n14/emily-wilson/an-uncomplicated-man
9•andsoitis•1h ago•2 comments

Three Theses on the Literacy Crisis

https://trevoraleo.substack.com/p/three-theses-on-the-literacy-crisis
30•samclemens•2d ago•19 comments

Vehicle Motion Cues

https://support.apple.com/guide/iphone/iphone-comfortably-riding-a-vehicle-iph55564cb22/ios
23•Austin_Conlon•2h ago•10 comments

The Burau representation of the braid group is faithful for n = 4

https://arxiv.org/abs/2607.05283
26•wglb•4h ago•9 comments

Self-contained highly-portable Python distributions

https://gregoryszorc.com/docs/python-build-standalone/main/
125•jcbhmr•9h ago•27 comments

Launch HN: Rise Reforming (YC S26) – Turning Waste Gases into Valuable Chemicals

https://www.rise-reforming.com
63•george_rose25•8h ago•30 comments

Glue bonds to nonstick surfaces and wipes clean with ethanol

https://cen.acs.org/materials/adhesives/glue-bonds-nonstick-surfaces-wipes-clean/104/web/2026/07
160•gmays•4d ago•89 comments

Securing Services with Rootless Containers

https://blog.coderspirit.xyz/blog/2026/07/06/securing-services-with-rootless-containers/
76•speckx•4d ago•24 comments

Ray tracing massive amounts of animated geometry using tetrahedral cages

https://gpuopen.com/learn/ray-tracing-massive-amounts-animated-geometry/
83•LorenDB•4d ago•10 comments

A Dying Art: The last of the morticians

https://harpers.org/archive/2026/08/a-dying-art-john-semley-mortuary-sciences-competition/
9•Petiver•3d ago•0 comments

Netflix employee fired for sharing personal details in retreat trust exercise

https://nypost.com/2026/07/26/us-news/netflix-exec-goes-ballistic-after-being-fired-for-stunning-...
206•softwaredoug•4h ago•157 comments

Some combinatorial applications of spacefilling curves

https://www2.isye.gatech.edu/~jjb/research/mow/mow.html
24•shraiwi•2d ago•1 comments

A missing underscore sent innocent man to prison for 18 months

https://arstechnica.com/tech-policy/2026/07/police-missed-one-underscore-and-sent-the-wrong-man-t...
191•quantified•5h ago•101 comments

How real are real numbers? (2004)

https://arxiv.org/abs/math/0411418
59•surprisetalk•12h ago•47 comments

Exploiting Volvo/Eicher's fleet platform to gain control over all users/vehicles

https://eaton-works.com/2026/07/27/my-eicher-hack/
143•EatonZ•12h ago•47 comments

UpCodes (YC S17) is hiring remote AE's to help make buildings cheaper

https://up.codes/careers?utm_source=HN
1•Old_Thrashbarg•11h ago

Paged Out #9 [pdf]

https://pagedout.institute/download/PagedOut_009.pdf
192•laurensr•13h ago•22 comments

The computer that helped win World War II

https://spectrum.ieee.org/colossus-computer-ieee-milestone
182•baruchel•5d ago•73 comments