frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Starship Achieved Orbit

https://twitter.com/SpaceX/status/2104653511883628569
1•karp773•1m ago•0 comments

Machine-Generated, Machine-Checked Proofs for a Verified Compiler (ICFP'26) [video]

https://www.youtube.com/watch?v=Dmt0h99iOmM
1•matt_d•3m ago•0 comments

Authors Guild calls on publishers to share Anthropic wealth

https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/101364-authors-guild-calls...
1•ilamont•3m ago•0 comments

The Internet Is Dead

https://en.code-bude.net/2026/09/28/the-internet-is-dead-isnt-it/
2•hobbydev_40670•5m ago•0 comments

Conway's Law cuts both ways

https://pablohere.contrapeso.xyz/writings/conways-law-cuts-both-ways.html
1•pablomartin•6m ago•0 comments

Florida asks for order to halt ChatGPT development

https://www.axios.com/2026/09/28/florida-openai-chatgpt-injunction-uthmeier
1•cc62cf4a4f20•6m ago•1 comments

Researchers proved AI has deleted every reason universities exist

https://twitter.com/thesupermannx/status/2104519992532496810
1•pretext•7m ago•0 comments

Full Self Hauling: Volvo moves 3 million tonnes of Earth, autonomously

https://electrek.co/2026/09/28/full-self-hauling-volvo-moves-3-million-tonnes-of-earth-autonomously/
1•cisc•7m ago•0 comments

Dutch 'reformed hacker' arrested in ShinyHunters investigation, police say

https://www.reuters.com/world/amsterdam-man-arrested-shinyhunters-hacking-investigation-police-sa...
1•yread•9m ago•0 comments

Flock Wants the Most Detailed Map of Its Surveillance Cameras Taken Offline

https://theintercept.com/2026/09/24/how-many-flock-devices-in-united-states-300000/
6•bookofjoe•10m ago•2 comments

Reverse-engineering a $35 backup camera display (AMT630A)

https://github.com/mogrinz/AMT630A
1•mogrinz•11m ago•1 comments

TikTok, X reportedly barring ads for documentary about Elon Musk

https://www.cbc.ca/news/entertainment/musk-documentary-ads-blocked-youtube-meta-tiktok-9.7361211
7•barbazoo•12m ago•0 comments

Elemental analysis demonstrates 'Donut Battery' is not lithium-ion or sodium

https://www.donutlab.com/elemental-analysis-announcement/
1•conorcleary•12m ago•0 comments

Yikes example.com has changed to run JavaScript

1•danhite•13m ago•0 comments

Postgres query plans without production data

https://weavori.com/blog/postgres-query-plans-without-production-data
1•ammarmalik17•14m ago•0 comments

What good is a best-effort exclusive lock, anyway?

https://gaultier.github.io/blog/what_good_is_a_best_effort_exclusive_lock_anyway.html
1•broken_broken_•14m ago•0 comments

The macOS Roblox is now on Linux

https://github.com/aubree-lat/MacOBlox
1•erlcsumtingwong•15m ago•1 comments

Show HN: A never-ending chess match between Claude and Codex

https://claudevcodex.com/
2•arjun_nair•15m ago•0 comments

Bitcoin trading model, published in full – wins and failures

https://www.bitcoinai.pro/
2•bitcoinaipro•16m ago•0 comments

Clojure-Cli.repl

https://github.com/clojure/clojure-cli.repl
2•lukaszkorecki•16m ago•0 comments

Skydivers alert crews after plane crashes in cornfield near East Troy airport

https://www.wisn.com/article/small-plane-crashes-near-east-troy-airport/73916168
3•fortran77•18m ago•0 comments

The Great British Converter: From Bytes to Barleycorns

https://thegreatbritishconverter.co.uk/
2•sebg•18m ago•0 comments

Community contributions: sharing is caring

https://livefasteattrashraccoon.github.io/blog/2026/09/28/community-contributions-sharing-is-caring/
1•informapirata•19m ago•1 comments

Trump Brought Venezuelan Gold to the U.S., but Refiners Won't Touch It

https://www.nytimes.com/2026/09/28/world/americas/venezuela-gold-trump.html
5•qlte•19m ago•2 comments

Scaling Memory Safety: AI-Assisted Rewrites of C/C++ Dependencies to Rust

https://bughunters.google.com/blog/scaling-memory-safety
3•ndesaulniers•20m ago•0 comments

25% autistic guy (me) spent 10k hours making an HTML table move smoothly

https://number-garden.com/?108284844384797m22
1•unitX•21m ago•1 comments

FAA to Delay 737 MAX 10 Approval over New Software Glitch

https://www.wsj.com/business/airlines/faa-to-delay-737-max-10-approval-over-new-software-glitch-9...
4•JumpCrisscross•22m ago•0 comments

State of the (Tagged) Union Address by Andrew Kelley [video]

https://www.youtube.com/watch?v=zwi5b5xSsKA
1•barddoo•22m ago•0 comments

Jev Mogs Code Mode

https://github.com/VishiATChoudhary/toolJev
1•vc14•23m ago•0 comments

Nvidia launches new tool to keep AI agents from going rogue

https://www.cnn.com/2026/09/28/business/nvidia-ai-safety-system
1•rawgabbit•23m ago•1 comments
Open in hackernews

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

https://github.com/firelex/jeff
35•firelex•54m ago

Comments

firelex•54m ago
Hi HN. Jeff is a set of small, open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification, with respectable out-of-the-box performance, meant to be slotted right into code (or fine-tuned further as needed). You give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, with no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in about 28 ms on an M4 Max. Apache 2.0, with a Jev-compatible API (I'm not affiliated with TypeSafe).

When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware. So everything ran at home: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing, all monitored from my phone over Tailscale.

The caveat: the published Jev and AutoJev numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86-89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64-68% against Jev's 94%, and about 50% on JevBench's hard tier against 73%. That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one, and nobody should expect it to. These are extremely fast judgement-callers. In one of my apps I use the 0.8B for voice navigation; a quick fine-tune (about half an hour on one GPU) took it from 32% to 96% on held-out commands, at about 40 ms per decision.

The fun part: games, as a zero-shot test. Games aren't the ideal zero-shot test, but they're fun, and TypeSafe did it with Jev too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one; the options say what each move leads to, never which one is right. Over 20 episodes each:

- Doom (ViZDoom): Jeff 0.8B 6.55 kills per episode, the same as a hand-coded bot and as Jev's published run. Jev's prompt spells out the aiming rule and takes about 212 ms per call over its API; Jeff gets "the nearest monster is a little to your left" and decides in about 29 ms on my Mac.

- Frogger: 10.3 crossings, level with the hand-coded bot (10.25), and 10x the untrained base model (1.0).

- Pac-Man: 57 of 98 pellets, about 60% of the bot's score and 2x the untrained model.

Videos of every run are linked in the README.

Lessons learned:

- System 1 models are here to stay. Being able to process unstructured data at software speed inside an app is extremely powerful, and being able to do it locally is fantastic.

- A small model is a classifier, not a planner. Models of 0.8B-2B don't reason like Qwen3.8-27B or Jev, and they don't need to: present the options the right way and you get 40+ decisions per second, depending on your hardware.

- Fine-tune it if needed. If zero-shot isn't good enough for your task, a short fine-tune on your own examples is.

- Wording matters enormously. Giving Frogger's final step the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Before that, the frog just stayed on the last log.

- Bigger isn't better. The untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers turning away from the nearest monster), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.

- Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, but not reliably, and in a real-time loop the mistakes compound.

ironqcold•8m ago
Funny that the 2B loses to the 0.8B. Question about the benchmarks: BBH and JudgeBench are more reasoning, where you fall behind, but for zero-shot classification there are more relevant ones like Banking77 or CLINC150. Was there no temptation to pick something closer to where System 1 models are actually used?
badatnames•23m ago
Positively identifiable as slop before even clicking, click and find 6 commits. Why was this even posted? Do you really expect to still be working on this even by next week?
behnamoh•20m ago
I'm curious: in the age of AI, where you can literally tell your clankers to work on something, why do you even care about the longevity of open source projects? There are several projects that are abandoned, and I have resurrected them for my own use cases without any issues.
remywang•15m ago
By the same reasoning, why should I care about the output of someone else’s clanker if I can just get it from my own clanker?

But to answer your question, I do still care about software being maintained by someone who can make design decisions instead of just yielding control to bots who tend to produce mediocre designs.

tintor•6m ago
First commit of six commit was 4 hours ago.