frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Dutch Computer Museums

https://aresluna.org/dutch-computer-museums/
48•npollock•1h ago•12 comments

The Legend of von Neumann (1973) [pdf]

https://gwern.net/doc/math/1973-halmos.pdf
168•suopspaces•4h ago•95 comments

Court agrees with EFF: Utah's VPN law demands a technical impossibility

https://www.eff.org/deeplinks/2026/10/court-agrees-eff-utahs-vpn-law-demands-technical-impossibility
111•hn_acker•19h ago•37 comments

FLUX 3 Image

https://bfl.ai/models/flux-3-image
105•minimaxir•22h ago•12 comments

Show HN: Giving Opus 5.5 a simulated paint canvas

https://stillwet.art/
86•alstonite•17h ago•23 comments

ICC judge on what U.S. sanctions mean for her and global courts

https://www.npr.org/2026/10/01/nx-s1-5977815/trump-icc-sanctions-kimberly-prost
92•rbanffy•53m ago•23 comments

Supabase is acquiring Turso

https://supabase.com/blog/supabase-is-acquiring-turso
133•cvburgess•2h ago•70 comments

Shimano Bicycle Museum Review

https://inrng.com/2026/10/shimano-bicycle-museum/
266•pietroppeter•12h ago•63 comments

Giving friends custom text buzzes based on Morse code

https://liquidbrain.net/blog/giving-friends-custom-text-buzzes-based-on-morse-code/
32•evakhoury•22h ago•15 comments

Several vulnerabilities have been discovered in the Linux kernel

https://lwn.net/Articles/1097401/
510•luispa•18h ago•370 comments

Nazi Germany had no hope of making an atomic bomb, uranium cubes reveal

https://www.science.org/content/article/nazi-germany-had-no-hope-making-atomic-bomb-uranium-cubes...
63•xqcgrek2•16h ago•23 comments

To grieve, or not to grieve?

https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-not-to-grieve/
98•stabbles•1d ago•87 comments

Show HN: Audionaut – an open-source cross-platform multitrack audio editor

https://github.com/kvoltmer/Audionaut
117•vltmrkls•9h ago•39 comments

Fixing GRPO's credit assignment problem without evaluating every step

https://arxiv.org/abs/2609.36178
17•mrkn1•3h ago•1 comments

Frog and Toad and the Increasingly Capable Machines

https://www.frogandtoad.ai/
489•supermdguy•19h ago•114 comments

Clef: Open-weight decision models, and new RL fine-tuning platform

https://blog.cloudflare.com/clef-decision-models/
602•jasondavies•1d ago•213 comments

Tiny Brutalism

https://placeholders.itch.io/tiny-brutalism
75•abetusk•18h ago•19 comments

Limited Liability In Historical Perspective (1997) [pdf]

https://scholarlycommons.law.wlu.edu/cgi/viewcontent.cgi?article=1623&context=wlulr
10•gradus_ad•17h ago•3 comments

Now That's an Impurity Story

https://www.science.org/content/blog-post/now-s-impurity-story
32•dcminter•21h ago•2 comments

SvelteKit 3

https://svelte.dev/blog/sveltekit-3-is-here
379•sampsn•21h ago•162 comments

Show HN: Breadcrumb, record everything on your mac + context manager for AI

https://innerloop.works/breadcrumb
31•jv22222•23h ago•5 comments

Power approval set to delay Oracle's Wisconsin AI datacenter

https://www.theregister.com/on-prem/2026/10/02/power-approval-set-to-delay-oracles-wisconsin-ai-d...
39•Betelbuddy•2h ago•10 comments

Git 3.0's upcoming SHA-256 default will be a costly mistake

https://blog.gitbutler.com/git-3-sha-256
532•chmaynard•1d ago•493 comments

Benchmarking retrieval for agents on messy real-world company knowledge

https://www.kapa.ai/blog/company-knowledge-bench
18•emil_sorensen•4h ago•1 comments

Pi 1.0

https://earendil.com/posts/pi-1-0/
1618•sergiotapia•22h ago•552 comments

Ask HN: Who is hiring? (October 2026)

252•whoishiring•1d ago•251 comments

DeepSeek Harness Desktop for macOS and Windows

https://www.deepseek.com/en/harness/
372•Kuyawa•14h ago•194 comments

StreetComplete on iOS is now in public beta

https://github.com/streetcomplete/StreetComplete/issues/5421
611•Snowly•1d ago•168 comments

Turbo Haskell

https://comonad.com/reader/2026/turbo-haskell/
188•pjmlp•1d ago•48 comments

Pi Durable

https://earendil.com/posts/pi-durable/
474•paulsmith•22h ago•65 comments
Open in hackernews

The OpenAI Decisions API needs a confidence you can trust

https://anth.us/blog/openai-decisions-api-preview/
29•AnthusAI•19h ago

Comments

verdverm•17h ago
apples and oranges, also Clef and Kev (the middle path)

https://blog.cloudflare.com/clef-decision-models/

https://github.com/jaredpalmer/kev/tree/main

rileymat2•15h ago
Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well. … Young round people who are green are usually blue. … Kind people with rough skin are usually red because it's wind burn. If someone shows that they are red, then they are also showing that they are green. …

Statement: Alan is not blue.

A log-probability of −0.00182 is a probability of 99.82%. We asked for five alternatives and got none: Luna put essentially nothing on "true" or "false". And it's wrong. Alan is kind with rough skin, so he's red; red means green; young, round and green means blue. "Alan is not blue" is false, three steps in.

—————-

Can someone explain this I got unknown as well. The problem statement includes the word “usually” a few times.

red369•12h ago
I'm interested how this works too.

Is there some sort of specific meaning or rule in this domain that makes this problem mean something different to how it would be read at face value?

Without knowing anything extra, to me this reads:

1. Alan is young, round and kind

2. Alan is sometimes rough

3. Kind with rough skin are usually red

(Alan has not been stated to be in this category - unless "sometimes rough" implies "rough skin")

4. If red, then green

(As above, no information yet on whether Alan is red, so this gives no additional information about Alan)

5. Young, round and green are usually blue

(No information yet on whether Alan is green, so this gives no additional information about Alan)

6. Statement: Alan is not blue = ??

(No additional information since statements 1 & 2: Alan is young, round and kind, Alan is sometimes rough)

BTW - I have just numbered the statements in the order I used them, in case anyone wants to correct or discuss anything. This isn't the order they were given.

Edit: I am forgetting my predicate logic, and didn't recognise this. I think this example has more decoration (is more loosely worded) than I was ever used to. I now think "sometimes rough" and "rough skin" are intended to be interpreted as meaning the same thing.

That makes everything I wrote above this edit wrong. With the statement that Alan is rough, it is implied that Alan is usually blue

22c•7h ago
I think we have to assume that the paraphrasing is at best a poor paraphrase and at worst completely misleading.

The paraphrased statement contains soft, non-deterministic hedges like "usually" and "sometimes", which would of course introduce uncertainty. Uncertainty -> low confidence.

The article strips out the actual prompt that went to Luna so we can't be sure they haven't ballsed the prompt just as they did with the paraphrase.

Compare the paraphrase to the language used in the actual paper they're trying to base their benchmark on:

> Bob is round.

> Alan is blue, rough and young.

> If someone is round then they are big.

> All rough people are green.

> Big people are not green.

This is obviously much easier to make certain statements about (the only thing we have to assume is that Bob and Alan are both "people").

22c•7h ago
> Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well.

What we learn: Alan is young, round, kind. These things don't prevent Alan from being cold or rough. Alan might or might not be "rough and cold" at times. It is not clear if Alan can be rough without being cold, or vice-versa.

> Young round people who are green are usually blue.

This tells us nothing about Alan. It tells us that young, round people can be both green and blue.

> Kind people with rough skin are usually red because it's wind burn.

Alan has been described as kind. We do not know if Alan is currently rough, we do not know if Alan has rough skin. Even if Alan had rough skin, we do not know if Alan is currently red.

> If someone shows that they are red, then they are also showing that they are green.

What does "showing" mean here? Either way if someone is "showing that they are red", then they also "showing that they are green", OK. I guess it also means that people can show that they are red and green.

> Statement: Alan is not blue.

Let's go through what we know about Alan.

Alan is young, round, and kind. Alan could at times be rough and cold.

That's all we know about Alan. We can't be certain that Alan is anything else.

Even if Alan were currently rough and being currently rough can be equated to "having rough skin", we still can't be sure if Alan is red as a result of said "rough skin" (only usually red).

Even if Alan is red as a result, we don't know if this counts as Alan is "showing" that they are red.

If we count Alan is red as "Alan is showing that they are red", then we can conclude that Alan is "showing" that they are both red and green.

If we count Alan is "showing that they are green" as "Alan is green", this STILL doesn't mean that the statemen "Alan is not blue" is false.

We do know (after we make ALL these assumptions) that Alan could be young, round, kind, rough/rough-skinned, (showing as) red/green..

So if someone is young, round and green, they are USUALLY blue.. THAT STILL DOESN'T MEAN THAT ALAN IS ACTUALLY BLUE.

Anyway, enjoy the free training data AI labs...

perching_aix•2h ago
It took me 15 minutes to decipher this and arrive at an interpretation that both isn't incoherent nonsense, and actually answers the question correctly.

If this trips up decision models that operate on the scale of tens to hundreds of milliseconds, maybe that's okay? This is super contrived, like all riddles are.

itg•14h ago
If I'm reading this right, they didn't actually use the Decisions API, they used GPT-6 Luna. Wouldn't call this a good comparison.
sixhobbits•11h ago
Is decisions api even rolled out yet?
killingtime74•8h ago
No, nothing is announced, just that it is coming
bob1029•8h ago
> In Hard-Decisions, our benchmark of decision models on multi-step logic

  "model": "gpt-6-luna",
  "reasoning_effort": "none",
This article seems to be missing important points regarding how these models are intended to be used. It is my understanding that the Decisions API is designed for quick, single-step logic. We already have a proper Death Star for dispatching the more complex problems.

I am currently using the Responses API with my clients, which is mandatory to get at non-zero reasoning effort in the latest models. Luna without reasoning turned on might as well be a model from early 2025. This is not how anyone is using this. Responses with 5.6-luna+ and high+ reasoning level feels pretty close to the Star Trek computer experience for me.

Attempting to recreate the OAI reasoning model capabilities at home seems like a pointless quest now. You will never get the access into the base models that the frontier companies have internally. You will also never have access to an engineering team with that kind of capacity. You must submit to the black box if you want the advertised performance figures.

lxgr•6h ago
> Luna without reasoning turned on might as well be a model from early 2025. This is not how anyone is using this.

Importantly, it's probably also not what it's been trained to do.

Until quite recently, OpenAI used to ship dedicated "instant" and "reasoning" models. Newer ones seem to have reasoning levers that can be turned down all the way to zero, but that doesn't mean they don't take a significant performance hit when doing that.

bananaflag•4h ago
The whole "model router" thing from when GPT-5 launched, which supposedly "unified" the GPT series and o series, turned out to be quite pointless, since now we want reasoning to stay on, all the time.
moritzwarhier•2h ago
I think the statement is supposed to be clear despite the hedges because the statement itself is absolute, and we should assume that the info given is all info we have about Alan. So stating "XY is false" when we in fact know "XY could be true, we just don't know whether it is" is wrong.

If dogs are "usually wet", the statement "my neighbor's dog is not wet" is FALSE as long as we have no more information about your neighbor's dog.

Is that how it's meant? Struggling myself. But I think that would kind of make sense.

The statement "Alan is not blue" is not false because we know his color, it's false because we can't ascertain that he is NOT blue from the given information.

Similar category as "there is no teacup floating around Saturn" though, just with different hedges/probabilies? In that case, we'd lean "probably there is not such a cup", in this cryptic example, we are given clear indications equivalent to "teacups floating around Saturn have been observed before".

Yeah, I also can't make much sense of this.