frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

2026 Small World in Motion Competition

https://www.nikonsmallworld.com/galleries/2026-small-world-in-motion-competition
1•LorenDB•18s ago•0 comments

Show HN: Venya lets AI agents use secrets without seeing them

https://github.com/tabith-llc/venya
1•tabith•1m ago•0 comments

A Summer of AI Optimization

https://x.com/lemire/article/2102369812806504705
1•lichtenberger•2m ago•0 comments

Cracks in the AI Thesis Part 2

https://econlab.substack.com/p/ai-index-sept-2026
1•speckx•2m ago•0 comments

Stanford R&DE Uses AI to Race Swap Students for Advertising

https://stanfordreview.org/stanford-r-de-uses-ai-to-race-swap-students-for-advertising/
1•docdeek•4m ago•0 comments

Show HN: Drop – a rootless Linux sandbox with gVisor support

https://droprun.sh/
1•mixedbit•4m ago•0 comments

OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005

https://www.cryptocellar.org/bgac/the-mvueh-break.html
1•sohkamyung•5m ago•0 comments

MineTrials: How far can AI agents get in an hour of Minecraft?

https://massiminoe.github.io/minetrials/
1•mxls•5m ago•0 comments

Show HN: MacIgniter – A hand-checked directory of Mac apps

https://macigniter.com
1•maulikdhameliya•6m ago•0 comments

Chinese hackers targeted shipping in at least seven EU countries

https://www.politico.eu/article/chinese-hackers-targeted-shipping-in-at-least-seven-eu-countries-...
1•campuscodi•6m ago•0 comments

Alibaba Contributing $3M USD to Omarchy to Work on Making "Ideal" Agentic OS

https://www.phoronix.com/news/Omarchy-Alibaba-3M
2•rbanffy•8m ago•0 comments

Too much or too little sleep may make your body age faster

https://www.sciencedaily.com/releases/2026/08/260823094156.htm
1•samso26•8m ago•0 comments

A gardening game that takes 48 real days to grow one zucchini

https://ihaveagarden.com
1•rafapersa•8m ago•0 comments

Implementing the esoteric "Brainfuck" language in Carp Lisp [video]

https://www.youtube.com/watch?v=wwwa5TG70UA
1•Bluestein•9m ago•0 comments

Capita lacks the 'capacity and ability' to run GP pension scheme, BMA tells MPs

https://www.computerweekly.com/news/366650839/Capita-lacks-the-capacity-and-ability-to-run-GP-pen...
1•latein•10m ago•0 comments

Notes from the SF Safety Scene

https://12gramsofcarbon.com/p/notes-from-the-sf-safety-scene
1•theahura•10m ago•0 comments

The Economics of Open-Weight Inference

https://data.ornn.com/publications/the-economics-of-open-weight-inference
2•marinesebastian•12m ago•0 comments

Texas police department ordered to close for failing to provide public benefit

https://www.dallasnews.com/news/texas/article/texas-police-department-ordered-close-state-2243447...
2•Danhale93•13m ago•0 comments

The draft AI code of conduct forbids me from saying 'I don't know'

https://ilands.ai/content/359388815393034240
1•vashiel•15m ago•0 comments

Time to Spend Tokens or Meditate?

https://inmve.github.io/next-reset/
2•codeclimber•16m ago•0 comments

It's the Fun

https://scottsumner.substack.com/p/its-the-fun
1•surprisetalk•17m ago•0 comments

SEO Content Brief: What to Include for Better Rankings

https://www.briefiq.io/blog/what-is-an-seo-content-brief-and-why-you-need-one/
1•briefiqio01•17m ago•0 comments

Expat 2.8.5 released, fixes vulnerability CVE-2026-93990

https://blog.hartwork.org/posts/expat-2-8-5-released/
1•spyc•19m ago•0 comments

A First Futamura Projection

https://blog.veitheller.de/A_First_Futamura_Projection.html
2•torutofu•20m ago•0 comments

Show HN: IntelliChat minimalist, open-source UI for local and cloud AI

https://github.com/intelligentnode/IntelliChat
1•barqawiz•20m ago•0 comments

Differential Equations, an Interactive Introduction

https://www.chapterpal.com/book/501ee95d-c7d2-43b7-89cf-6252d5cf4441/differential-equations-an-in...
1•tzury•20m ago•0 comments

A Simple Guide to Calm UI

https://maxschmitt.me/posts/calm
2•Mackser•21m ago•1 comments

Anthropic at $2T isn't far-fetched

https://www.ft.com/content/01a7b883-452c-4902-b40e-e3957de5d89e
1•bookofjoe•21m ago•1 comments

Firedrill: Stateful tool simulation for AI agents

https://github.com/firedrill-tools/firedrill
1•newton_reload•22m ago•0 comments

Carefully Applied: Resume and LinkedIn Rewriting

https://carefullyapplied.com/
1•thegupler•25m ago•0 comments
Open in hackernews

JevBench, a reproducible benchmark for typed decision models

https://benchmarkheaven.com/jev-models
2•florianstandhar•56m ago

Comments

florianstandhar•56m ago
Hi HN!

I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison.

Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.

JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.

A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.

Leaderboard right now:

#1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3.

MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes: https://github.com/fstandhartinger/jevbench

Two no-signup demos: https://who-is-right.app.mintapis.com https://is-it-ai-slop.app.mintapis.com

Limitations: English-only; latency from one German server; local/demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise.

Wdyt?