frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Decisions API is in public beta

https://developers.openai.com/api/docs/guides/decisions
55•chiefstorm•2h ago

Comments

Imustaskforhelp•1h ago
This rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.

Interesting to see where all this leads us and if other major labs follow suit

Edit: decisions voice looks really interesting as well[1]

[0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev

[1]: https://developers.openai.com/api/docs/guides/decisions-voic...

lab14•1h ago
How is the pricing vs Jev?
jerrygenser•40m ago
$0.10/mm input vs. $0.042/mm input. Both free output.
esafak•1h ago
You knew it was going to happen! Benchmarks or it didn't happen.
peterson_lock•38m ago
Can we use this through subscription?
mritchie712•36m ago
it already supports image inputs, which was the first big gap I found in Jev.
mohsen1•36m ago
Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API
binlog•34m ago
One of the examples on the docs page is it playing a video game. Doubt it’ll be able to run anything complex though.
dvt•32m ago
I genuinely do not understand why anyone would pay OpenAI for this. Running something comparable to Jev is pretty trivial. The whole point of paying for ChatGPT is because OpenAI has a bunch of warehouses that can run a zillion-parameter model.

Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.

mediaman•29m ago
Why would I run it myself? It's $0.10 per million tokens. Dirt cheap. (Jev is even cheaper.)

You could ask the same question about why anyone would rent a VPS. I can just run my own hardware, it's just a computer!

Buy vs rent is not just about what's possible, it's about what's economic.

dvt•18m ago
Yeah and OAI is twice as expensive as Jev, which is kind of my point. And more expensive than open models, which you don't necessarily have to host yourself. Pure bandwaggoning.
TSiege•29m ago
There isn’t a moat in the sense of self hosting but you need a reason for people who don’t want that to stay on your platform. Customers save time and effort managing payments easier this way. However it’s a race to the bottom price wise.

Going to be all about branding and platform stickiness for OpenAI to make investors and creditors whole.

simonw•28m ago
Depends on the quality of the results. These things are driven by text prompts. If it turns out the OpenAI one returns better quality results than open weight variants they'll be rewarded by the market.

Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.

TSiege•31m ago
The response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.

Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.

If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.

gobdovan•22m ago
> If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.

Hopefully people will flock to whatever product is making its mission to be commodity and the easiest to replace. Really don't want another free ingress, 100$/TB egress Cloud situation.

simonw•30m ago

  curl https://api.openai.com/v1/decisions \
    -H "Authorization: Bearer $(llm keys get openai)" \
    -H "Content-Type: application/json" \
    --data '
  {
    "model": "gpt-6-luna",
    "input": [{
      "role": "user",
      "content": [
        {"type": "input_text", "text": "I am angry about the new product feature"}
      ]
    }],
    "questions": [{
      "type": "predicate",
      "name": "complaint",
      "instructions": "Is this a complaint?"
    }, {
      "type": "predicate",
      "name": "compliment",
      "instructions": "Is this a compliment?"
    }]
  }'
Returned:

  {
    "model": "gpt-6-luna",
    "answers": [
      {
        "type": "predicate",
        "name": "complaint",
        "probability": 0.91
      },
      {
        "type": "predicate",
        "name": "compliment",
        "probability": 0.06
      }
    ],
    "usage": {
      "input_tokens": 310,
      "input_tokens_details": {
        "cached_tokens": 0,
        "cache_write_tokens": 0
      },
      "output_tokens": 0,
      "output_tokens_details": {
        "reasoning_tokens": 0
      },
      "total_tokens": 310
    }
  }
That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.
sidcool•24m ago
Jev really shook up the industry. This seems obvious in hindsight
Topfi•24m ago
Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).

Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).

Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).

Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.

Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).

Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.

oh_no•11m ago
I think releasing something like this makes sense even if it's underbaked, it's still very cheap, and if you have existing enterprise OpenAI relationship it's a lot easier to onboard something like this than set up a new Jev contract.

How these models play out is an open question but existing provider contracts and T&C are important for enterprise.

tmhall•28m ago
For my use case it will cost like $11 a month and we already have OpenaAI keys and accounts with billing in place. I don't want to run my own model infra and I don't want to get permission to set up an account with typesafe.ai
super256•27m ago
Existing enterprise contracts? Data retention contracts (some have zero data retention contracts)? Staying with a single provider because it's easier to have everything in one place?

There are probably a lot more reasons.

jcims•26m ago
If you work for a company that has a 3 to 6 month onboarding period for new vendors and a lifetime commitment to maintain a whole bunch of vendor management horseshit for as long as that relationship exists, it makes a ton of sense.

Add in a bunch of model governance and oversight for anything you train yourself and it’s pretty much a slam dunk deal.

csharpminor•23m ago
If you're in an enterprise that already has a procurement agreement with OpenAI, this means you don't have to onboard another vendor. Bucket platform strategy.

Sharing AI Progress in Mathematics

https://openai.com/index/sharing-ai-progress-in-mathematics/
104•OfficialTurkey•40m ago•60 comments

Mistral Large 4

https://mistral.ai/news/mistral-large-4/\
1506•Philpax•9h ago•932 comments

Decisions API is in public beta

https://developers.openai.com/api/docs/guides/decisions
56•chiefstorm•2h ago•24 comments

EmbeddingGemma 2: An open, lightweight multimodal embedding model

https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/
169•ilreb•6h ago•25 comments

Nobel Prize in Physics 2026: Francis Halzen

https://www.nobelprize.org/prizes/physics/2026/
497•solarist•13h ago•166 comments

Paramount Skydance has completed its $111B merger with Warner Bros. Discovery

https://arstechnica.com/tech-policy/2026/10/paramount-completes-111b-warner-merger-creating-skyda...
120•Mgtyalx•2h ago•173 comments

OpenTPU – An open-source AI accelerator, developed by AI

https://github.com/FeSens/openTPU
198•fsbonetto•6h ago•259 comments

Claude Code’s suggested message feature: I think the real customer is the model

https://www.zohaib.cc/blog/smartest-claude-code-feature
50•zed_labs_dev•4h ago•26 comments

The Query Transformation Pipeline

https://readyset.io/blog/how-readyset-rewrites-your-sql-inside-the-query-transformation-pipeline
3•gvsg-rs•48m ago•0 comments

OpenSSH 10.6

https://www.openssh.org/releasenotes.html#10.6
64•torcete•2h ago•10 comments

Benchmark in Milliseconds

https://matklad.github.io/2026/10/05/benchmark-milliseconds.html
123•surprisetalk•1d ago•34 comments

Ask HN: Why is Ask HN only showing me 14 posts?

24•Gooblebrai•1h ago•24 comments

The Early History of Smalltalk (1993)

https://worrydream.com/EarlyHistoryOfSmalltalk/
106•_reza•7h ago•56 comments

Gleam doesn't compile to Erlang source anymore

https://gleam.run/news/gleam-doesnt-compile-to-erlang-source-anymore/
283•ingve•14h ago•120 comments

LLMs may have helped my RSI

https://vaughanhilts.me/2026/10/05/llms-immensely-helped-my-rsi.html
3•vaughands•19h ago•1 comments

Berthd

https://berthd.app/
37•handfuloflight•3h ago•55 comments

Erdosproblems.com Succumbs to the AI Onslaught

https://www.erdosproblems.com/forum/thread/blog:9
79•pfdietz•10h ago•37 comments

Show HN: I turned my iPhone and a $20 smart plug into an f-stop timer

https://peterszentkiralyi.eu/darkplug/
61•pentakkusu•8h ago•15 comments

Mathematics of Geothermal Energy

https://www.ebsco.com/research-starters/power-and-energy/mathematics-geothermal-energy/
63•srameshc•9h ago•36 comments

Subquadratic 3SUM and Subcubic APSP

https://arxiv.org/abs/2610.06783
90•mauriziocalo•10h ago•34 comments

Toronto-Based VPN Provider Plans to Quit Canada over Lawful-Access Bill

https://citizenlab.ca/toronto-based-vpn-provider-plans-to-quit-canada-over-lawful-access-bill/
67•speckx•4h ago•27 comments

Ask HN: How would you feel if we nationalized Google?

4•59percentmore•16m ago•4 comments

System-level ad-blocking in Android

https://kevinboone.me/adblock.html
35•birdculture•2h ago•28 comments

Beam: Reflection's 501B open-weight model

https://reflection.ai/blog/introducing-beam
538•Philpax•1d ago•167 comments

Polars 2.0

https://pola.rs/posts/release-polars-2/
399•simicd•10h ago•94 comments

Example.com just launched the biggest redesign in decades

https://www.debugbear.com/blog/example-dot-com-redesign-history
338•jgx0•1d ago•226 comments

Nature's capacity to 'bounce back' when species are lost is overestimated: study

https://phys.org/news/2026-10-nature-capacity-species-lost-vastly.html
301•pseudolus•11h ago•149 comments

What's Earth's dominant species by mass?

https://signoregalilei.com/2026/09/27/whats-earths-dominant-species-by-mass/
118•surprisetalk•10h ago•70 comments

Show HN: Parseable, an open observability datalake, handles 100M time-series/min

https://www.parseable.com
72•yashdotrv•9h ago•18 comments

Dust: Pretraining Transformers Without Backpropagation

https://qlabs.sh/research/dust
269•E-Reverance•1d ago•79 comments