frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

https://github.com/firelex/jeff
79•firelex•1h ago

Comments

firelex•1h ago
Hi HN. Jeff is a set of small, open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification, with respectable out-of-the-box performance, meant to be slotted right into code (or fine-tuned further as needed). You give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, with no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in about 28 ms on an M4 Max. Apache 2.0, with a Jev-compatible API (I'm not affiliated with TypeSafe).

When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware. So everything ran at home: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing, all monitored from my phone over Tailscale.

The caveat: the published Jev and AutoJev numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86-89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64-68% against Jev's 94%, and about 50% on JevBench's hard tier against 73%. That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one, and nobody should expect it to. These are extremely fast judgement-callers. In one of my apps I use the 0.8B for voice navigation; a quick fine-tune (about half an hour on one GPU) took it from 32% to 96% on held-out commands, at about 40 ms per decision.

The fun part: games, as a zero-shot test. Games aren't the ideal zero-shot test, but they're fun, and TypeSafe did it with Jev too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one; the options say what each move leads to, never which one is right. Over 20 episodes each:

- Doom (ViZDoom): Jeff 0.8B 6.55 kills per episode, the same as a hand-coded bot and as Jev's published run. Jev's prompt spells out the aiming rule and takes about 212 ms per call over its API; Jeff gets "the nearest monster is a little to your left" and decides in about 29 ms on my Mac.

- Frogger: 10.3 crossings, level with the hand-coded bot (10.25), and 10x the untrained base model (1.0).

- Pac-Man: 57 of 98 pellets, about 60% of the bot's score and 2x the untrained model.

Videos of every run are linked in the README.

Lessons learned:

- System 1 models are here to stay. Being able to process unstructured data at software speed inside an app is extremely powerful, and being able to do it locally is fantastic.

- A small model is a classifier, not a planner. Models of 0.8B-2B don't reason like Qwen3.8-27B or Jev, and they don't need to: present the options the right way and you get 40+ decisions per second, depending on your hardware.

- Fine-tune it if needed. If zero-shot isn't good enough for your task, a short fine-tune on your own examples is.

- Wording matters enormously. Giving Frogger's final step the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Before that, the frog just stayed on the last log.

- Bigger isn't better. The untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers turning away from the nearest monster), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.

- Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, but not reliably, and in a real-time loop the mistakes compound.

ironqcold•37m ago
Funny that the 2B loses to the 0.8B. Question about the benchmarks: BBH and JudgeBench are more reasoning, where you fall behind, but for zero-shot classification there are more relevant ones like Banking77 or CLINC150. Was there no temptation to pick something closer to where System 1 models are actually used?
danbrooks•8m ago
This type of project looks extremely useful. There was a lot of buzz around Jev, but having models that run locally and can be fine-tuned is extremely helpful.
AgentMasterRace•2m ago
I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.

World Labs Is Joining AMD

https://www.worldlabs.ai/blog/amd-announcement
86•mfiguiere•1h ago•21 comments

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

https://github.com/firelex/jeff
80•firelex•1h ago•8 comments

Pirating the Pirates

https://mubi.com/en/notebook/posts/pirating-the-pirates
343•piotrgrabowski•5h ago•165 comments

MicroLLM Lab – Try 7 tiny LLM's in the browser

https://stateofutopia.com/experiments/microllmlab/
83•logicallee•2h ago•29 comments

12,000-year-old Göbeklitepe burials explain scattered bones

https://archaeologymag.com/2026/09/gobeklitepe-burials-hundreds-of-scattered-bones/
32•yusufaytas•2d ago•3 comments

Pacing the Frontier is not the actual goal for AI labs

https://www.lesswrong.com/posts/Nm4ewbYovtjq69dvH/pacing-the-frontier-is-not-the-actual-goal-for-...
20•brlewis•59m ago•13 comments

Scientists solve 1840s space weather mystery

https://arstechnica.com/science/2026/09/scientists-solve-1840s-space-weather-mystery/
23•gumby•1h ago•9 comments

Hijacking the PS5's RTMP stream

https://yashgarg.dev/posts/hijacking-ps5-rtmp-stream/
159•ibobev•6h ago•50 comments

Joseph Szabo’s pictures of American adolescents

https://www.newyorker.com/culture/photo-booth/the-teen-portraits-that-captivated-sofia-coppola
62•prismatic•4h ago•25 comments

Parley: Federated, decentralised chat that speaks plain IRC

https://git.mills.io/prologic/parley
277•davidcollantes•11h ago•142 comments

It's Time to Investigate the AI Labs

https://calnewport.com/its-time-to-investigate-the-ai-labs/
82•ibobev•1h ago•14 comments

Sonnet 5.5

https://www.anthropic.com/claude-sonnet-5-5
444•D2OQZG8l5BI1S06•3h ago•305 comments

First Steps of the PLC Organization – Independent Public Ledger of Credentials

https://blog.plcred.org/3mwlphq42d227
28•embedding-shape•2h ago•15 comments

Palantir founder purchases large swath of forest in Sweden

https://www.arctictoday.com/palantir-founder-purchases-large-swath-of-forest-in-sweden/
56•davidja•55m ago•48 comments

Launch HN: Vespper (YC F24) – SOTA Docx MCP

https://www.vespper.com/blog/launching-vespper-docx-mcp
27•topaztee•4h ago•8 comments

Cf: The Agentic CLI for the Cloudflare API

https://blog.cloudflare.com/cloudflare-cf-cli-launch/
82•macleos•6h ago•37 comments

I made a visual workspace for AI Automations

https://www.biom.dev/
23•jduhking•3h ago•18 comments

Who wrote Elizabeth I's most scathing letters?

https://www.smithsonianmag.com/history/who-wrote-elizabeth-is-most-scathing-letters-new-research-...
28•benbreen•4h ago•16 comments

Best of British Design

https://best-of-british-design.vercel.app/
25•ArisC•1h ago•13 comments

What heraldry and Japanese mon can teach about visual-identity generators

https://benovermyer.com/blog/2026/09/japanese-vs-western-heraldry/
57•bovermyer•6h ago•19 comments

What reversing, modernising old games tells us about the economic impact of AI

https://this.os.isfine.org/blog/posts/what-reverse-engineering-and-modernising-an-old-war-game-te...
5•keeda•1d ago•1 comments

GrapheneOS – When an app is slow

https://blog.wirelessmoves.com/2026/09/grapheneos-when-an-app-is-slow.html
68•speckx•3h ago•33 comments

MongoDB CEO resigns to join Meta

https://www.reuters.com/technology/mongodb-ceo-desai-steps-down-lead-metas-enterprise-platform-20...
284•diek•6h ago•236 comments

Nvidia wants to put a watchdog chip next to every AI agent

https://www.cnbc.com/2026/09/28/nvidia-releases.html
62•jonbaer•6h ago•106 comments

Flock Wants the Most Detailed Map of Its Surveillance Cameras Taken Offline

https://theintercept.com/2026/09/24/how-many-flock-devices-in-united-states-300000/
24•bookofjoe•39m ago•6 comments

Show HN: Destroy Any Website with Stickman

https://destroy.spritefusion.com/
70•HugoDz•5h ago•21 comments

Kids turned low-traffic NPR Spotify comments into a secret group chat

https://www.thisamericanlife.org/897/transcript
200•simonpure•6h ago•132 comments

I switched to Brave

https://kevquirk.com/i-switched-to-brave-browser
87•mindracer•7h ago•143 comments

Neal Stephenson responds with wit and humor (2004)

https://slashdot.org/story/04/10/20/1518217/neal-stephenson-responds-with-wit-and-humor
49•jdkee•2h ago•19 comments

Windows 11½

https://definitelynotwindows.com/
375•jjbinx007•3h ago•115 comments