frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Apple Pass Designer

https://developer.apple.com/pass-designer/
102•soheilpro•51m ago•45 comments

Around 2-6% of World Bank foreign aid got siphoned into crypto wallets

https://www.nber.org/papers/w35655
51•bko•1h ago•6 comments

Court agrees with EFF: Utah's VPN law demands a technical impossibility

https://www.eff.org/deeplinks/2026/10/court-agrees-eff-utahs-vpn-law-demands-technical-impossibility
301•hn_acker•21h ago•137 comments

With most information hidden, the game Stratego had stumped AI until now

https://arstechnica.com/science/2026/10/ai-finally-beat-the-best-stratego-player-in-history-and-d...
75•PaulHoule•5h ago•20 comments

A 12-year sequence of telescope images of a star and four planets orbiting

https://bsky.app/profile/theplanetaryguy.com/post/3mwucf5ert22f
48•mariuz•8h ago•8 comments

Greg Kroah-Hartman – Security in the LLM Age [video]

https://www.youtube.com/watch?v=NnV_cWeoo5Q
62•usernomdeguerre•17h ago•8 comments

Loss of cell identity drives human aging: Two new papers

https://erictopol.substack.com/p/loss-of-cell-identity-drives-human
72•bookofjoe•23h ago•11 comments

On social reality in China

https://www.lesswrong.com/posts/b5cSYh4emQb2qrGmK/on-social-reality-in-china
43•thicTurtlLverXX•8h ago•24 comments

Mike Tomlin spent 12 years building a Minecraft city

https://www.nytimes.com/athletic/7648198/2026/10/01/mike-tomlin-minecraft-nfl-coach/
65•CoryOndrejka•1d ago•14 comments

Sites in ChatGPT

https://chatgpt.com/features/sites/
122•polvi•21h ago•135 comments

FLUX 3 Image

https://bfl.ai/models/flux-3-image
196•minimaxir•1d ago•45 comments

Blogging with Gleam, Org-Mode and Pandoc

https://byzantine-systems.github.io/blogging-with-gleam-org-mode-and-pandoc/
18•schonfinkel•8h ago•2 comments

One month coding with GLM 5.3 Flash

https://wagtail.org/blog/one-month-on-glm-53-flash/
33•ThibWeb•4h ago•26 comments

The Legend of von Neumann (1973) [pdf]

https://gwern.net/doc/math/1973-halmos.pdf
209•suopspaces•6h ago•121 comments

Our Project Suncatcher prototype satellite is in orbit

https://blog.google/innovation-and-ai/models-and-research/google-research/project-suncatcher-prot...
22•pantalaimon•8h ago•24 comments

The first packet sent via RFC1149 avian carrier is up for auction at Christie's

https://onlineonly.christies.com/s/fine-printed-books-manuscripts-science/carrier-pigeon-internet...
12•peter_hansteen•7h ago•1 comments

From the creator of Redis; run LLM locally with ds4

https://dwarfstar.sh/
23•fibo•1h ago•1 comments

How accurately calibrated is Jev?

https://maximumeffort.substack.com/p/jev-is-poorly-calibrated
20•dblack12705•4h ago•7 comments

STS-51-F Abort-to-Orbit (1985)

https://en.wikipedia.org/wiki/STS-51-F
11•schoen•3h ago•3 comments

Show HN: Giving Opus 5.5 a simulated paint canvas

https://stillwet.art/
138•alstonite•19h ago•46 comments

Venice’s failed war against Constantinople led to the first bond market

https://bigthink.com/books/a-fabulous-debt/
15•RickJWagner•6h ago•5 comments

100 years of student radio history in the DLARC college radio collections

https://blog.archive.org/2026/10/02/100-years-of-student-radio-history-in-the-dlarc-college-radio...
14•HieronymusBosch•7h ago•1 comments

What if we stopped using GPUs? [video]

https://www.youtube.com/watch?v=xc2FTBGRSJo
12•sandslash•1h ago•3 comments

GrapheneOS has fixed the Android 17 QPR1 kernel performance regression

https://discuss.grapheneos.org/d/42511-grapheneos-has-fixed-the-massive-android-17-qpr1-kernel-pe...
9•Cider9986•13m ago•0 comments

"The only intuitive interface is the nipple" (2012)

https://www.greenend.org.uk/rjk/misc/nipple.html
12•ibobev•40m ago•7 comments

Show HN: Pyxel – A Python retro game engine with built-in art and sound editors

https://github.com/kitao/pyxel
24•kitao•20h ago•4 comments

Anatomy of a Lean proof for software engineers

https://agostbiro.net/posts/2026-10-anatomy-of-a-lean-proof/
24•abiro•1d ago•0 comments

Supabase is acquiring Turso

https://supabase.com/blog/supabase-is-acquiring-turso
171•cvburgess•4h ago•90 comments

Updates to Full Disk Access in macOS

https://developer.apple.com/news/?id=p6zjojqw
11•notfirstpost•21m ago•1 comments

F.02 Decommission

https://www.figure.ai/news/f-02-decommission
18•ad_hockey•9h ago•5 comments
Open in hackernews

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

https://developers.redhat.com/articles/2026/10/02/benchmarking-ai-decision-models-against-traditional-guardrails
21•tomncooper•6h ago

Comments

dominotw•42m ago
prompts that these evaluations were done are too trivial
segmondy•29m ago
duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.

LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.

traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.

decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3

AnthusAI•23m ago
That was a pretty simple task they gave it, and sure you can use BERT with sequence classification for simple classification tasks.

In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.

https://hard-decisions.anth.us/models/

6thbit•22m ago
Shouldn't LLMs intuitively be better with a high number of available options? This article only does simple prompts with only options to block or not block.

What is openai doing for their decisions API, a finetuned luna?

reexpressionist•4m ago
The key properties for using such models for conditional-branching decisions in agentic stacks (and related) is that they should be well-calibrated (under the definition chosen for the task) and informative (e.g., always predicting the mean might be "well-calibrated" in a theoretical sense for some chosen quantities of interest, but isn't particularly useful in practice).

The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:

[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics.

[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269.