Statement: Alan is not blue.
A log-probability of −0.00182 is a probability of 99.82%. We asked for five alternatives and got none: Luna put essentially nothing on "true" or "false". And it's wrong. Alan is kind with rough skin, so he's red; red means green; young, round and green means blue. "Alan is not blue" is false, three steps in.
—————-
Can someone explain this I got unknown as well. The problem statement includes the word “usually” a few times.
Is there some sort of specific meaning or rule in this domain that makes this problem mean something different to how it would be read at face value?
Without knowing anything extra, to me this reads:
1. Alan is young, round and kind
2. Alan is sometimes rough
3. Kind with rough skin are usually red
(Alan has not been stated to be in this category - unless "sometimes rough" implies "rough skin")
4. If red, then green
(As above, no information yet on whether Alan is red, so this gives no additional information about Alan)
5. Young, round and green are usually blue
(No information yet on whether Alan is green, so this gives no additional information about Alan)
6. Statement: Alan is not blue = ??
(No additional information since statements 1 & 2: Alan is young, round and kind, Alan is sometimes rough)
BTW - I have just numbered the statements in the order I used them, in case anyone wants to correct or discuss anything. This isn't the order they were given.
Edit: I am forgetting my predicate logic, and didn't recognise this. I think this example has more decoration (is more loosely worded) than I was ever used to. I now think "sometimes rough" and "rough skin" are intended to be interpreted as meaning the same thing.
That makes everything I wrote above this edit wrong. With the statement that Alan is rough, it is implied that Alan is usually blue
The paraphrased statement contains soft, non-deterministic hedges like "usually" and "sometimes", which would of course introduce uncertainty. Uncertainty -> low confidence.
The article strips out the actual prompt that went to Luna so we can't be sure they haven't ballsed the prompt just as they did with the paraphrase.
Compare the paraphrase to the language used in the actual paper they're trying to base their benchmark on:
> Bob is round.
> Alan is blue, rough and young.
> If someone is round then they are big.
> All rough people are green.
> Big people are not green.
This is obviously much easier to make certain statements about (the only thing we have to assume is that Bob and Alan are both "people").
What we learn: Alan is young, round, kind. These things don't prevent Alan from being cold or rough. Alan might or might not be "rough and cold" at times. It is not clear if Alan can be rough without being cold, or vice-versa.
> Young round people who are green are usually blue.
This tells us nothing about Alan. It tells us that young, round people can be both green and blue.
> Kind people with rough skin are usually red because it's wind burn.
Alan has been described as kind. We do not know if Alan is currently rough, we do not know if Alan has rough skin. Even if Alan had rough skin, we do not know if Alan is currently red.
> If someone shows that they are red, then they are also showing that they are green.
What does "showing" mean here? Either way if someone is "showing that they are red", then they also "showing that they are green", OK. I guess it also means that people can show that they are red and green.
> Statement: Alan is not blue.
Let's go through what we know about Alan.
Alan is young, round, and kind. Alan could at times be rough and cold.
That's all we know about Alan. We can't be certain that Alan is anything else.
Even if Alan were currently rough and being currently rough can be equated to "having rough skin", we still can't be sure if Alan is red as a result of said "rough skin" (only usually red).
Even if Alan is red as a result, we don't know if this counts as Alan is "showing" that they are red.
If we count Alan is red as "Alan is showing that they are red", then we can conclude that Alan is "showing" that they are both red and green.
If we count Alan is "showing that they are green" as "Alan is green", this STILL doesn't mean that the statemen "Alan is not blue" is false.
We do know (after we make ALL these assumptions) that Alan could be young, round, kind, rough/rough-skinned, (showing as) red/green..
So if someone is young, round and green, they are USUALLY blue.. THAT STILL DOESN'T MEAN THAT ALAN IS ACTUALLY BLUE.
Anyway, enjoy the free training data AI labs...
If this trips up decision models that operate on the scale of tens to hundreds of milliseconds, maybe that's okay? This is super contrived, like all riddles are.
"model": "gpt-6-luna",
"reasoning_effort": "none",
This article seems to be missing important points regarding how these models are intended to be used. It is my understanding that the Decisions API is designed for quick, single-step logic. We already have a proper Death Star for dispatching the more complex problems.I am currently using the Responses API with my clients, which is mandatory to get at non-zero reasoning effort in the latest models. Luna without reasoning turned on might as well be a model from early 2025. This is not how anyone is using this. Responses with 5.6-luna+ and high+ reasoning level feels pretty close to the Star Trek computer experience for me.
Attempting to recreate the OAI reasoning model capabilities at home seems like a pointless quest now. You will never get the access into the base models that the frontier companies have internally. You will also never have access to an engineering team with that kind of capacity. You must submit to the black box if you want the advertised performance figures.
Importantly, it's probably also not what it's been trained to do.
Until quite recently, OpenAI used to ship dedicated "instant" and "reasoning" models. Newer ones seem to have reasoning levers that can be turned down all the way to zero, but that doesn't mean they don't take a significant performance hit when doing that.
If dogs are "usually wet", the statement "my neighbor's dog is not wet" is FALSE as long as we have no more information about your neighbor's dog.
Is that how it's meant? Struggling myself. But I think that would kind of make sense.
The statement "Alan is not blue" is not false because we know his color, it's false because we can't ascertain that he is NOT blue from the given information.
Similar category as "there is no teacup floating around Saturn" though, just with different hedges/probabilies? In that case, we'd lean "probably there is not such a cup", in this cryptic example, we are given clear indications equivalent to "teacups floating around Saturn have been observed before".
Yeah, I also can't make much sense of this.
verdverm•17h ago
https://blog.cloudflare.com/clef-decision-models/
https://github.com/jaredpalmer/kev/tree/main