frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

How accurately calibrated is Jev?

https://maximumeffort.substack.com/p/jev-is-poorly-calibrated
22•dblack12705•4h ago

Comments

aaryan__verma•4h ago
[flagged]
dang•53m ago
Can you please not post AI-generated or AI-edited comments to HN? It's not allowed here - see https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079.

Of course, it's impossible to know for sure what was LLM processed or not, but this post got classified that way.

popalchemist•50m ago
not very
singularity2001•39m ago
Also, it's a very small and dumb model. That's the only reason why it's so fast.
dvt•36m ago
> So Jev does actually know which distributions are correct, it just fails to produce them

Claims like this needs to be deeply analyized. Models do have some emergent capabilities[1], and I think there's a lot of evidence to show that semantics is actually learned (Word2Vec), and some math seems like it also might be learned (e.g. modular arithmetic). But it's hard to exactly say where there's some internal mechanism generating a true answer and where we're just getting lucky with some distribution so the answer just seems right.

[1] https://arxiv.org/pdf/2502.00873

cannedbread•34m ago
IMO by asking Jev underspecified questions like this, you're essentially using it as a random number generator (similar to the dice example). On actual NLP problems (including ones with uncertainty under human review) it does appear to be well calibrated: https://leonardgrazian.com/blog/jev-calibration/
amluto•26m ago
I’ve been contemplating this situation. I think that what the author wants out of a Jev-like model is not at all what I want out of it.

> I decided to check this on questions where the answer is well understood. For example:

> A classical particle of mass m is embedded in a system at thermodynamic equilibrium with temperature T. What is its velocity v?

If I feed that into a model, the answer I want is: “the combination of the model and the provided state has nothing useful to add to your prior”.

If I want to know the Maxwell-Boltzmann distribution, I can look it up or I can derive it or I can ask a fancy LLM to do it for me (at the cost of some reasoning tokens and some time - unless I’m using an ultraspeed inference system, I’m not getting this answer in 50ms). [0]

Similarly, if I want to know that 73% of incoming customer support requests are spam/fraud, I should measure that - it’s a property of my system, it takes some manual classification and a database query, and it will be a different percentage than your customer support system would see. I neither expect nor want my classifier to know this (unless I’m using a conventional classifier manually trained on my data, and the whole point of Jev is to avoid this).

What I want out of a system like Jev is to tell me how the probabilities change as a result of the per-sample data I provide. Which, is the case of this Boltzmann distribution question, is nothing: I provided no data and the classifier can infer nothing.

[0] A really good answer would observe that the answer depends on the dimension of the system (probably 3, but 2D systems are a thing) and also on whether the particles are hot enough for relativistic effects to matter (probably not). And maybe a good answer would check whether the material is a gas - the answer for a solid is not the same, but I suppose that’s not classical. Oh, and one shouldn’t forget drift: if you have a classical particle in a moving fluid or a classical charged particle in an electric field, you will again get a different answer.

Yes, I’m being pedantic. But if you want good answers you should be pedantic, and the Jev-like model is not where the pedantry should go.

Apple Pass Designer

https://developer.apple.com/pass-designer/
114•soheilpro•57m ago•55 comments

Around 2-6% of World Bank foreign aid got siphoned into crypto wallets

https://www.nber.org/papers/w35655
61•bko•1h ago•10 comments

Court agrees with EFF: Utah's VPN law demands a technical impossibility

https://www.eff.org/deeplinks/2026/10/court-agrees-eff-utahs-vpn-law-demands-technical-impossibility
310•hn_acker•21h ago•140 comments

With most information hidden, the game Stratego had stumped AI until now

https://arstechnica.com/science/2026/10/ai-finally-beat-the-best-stratego-player-in-history-and-d...
81•PaulHoule•5h ago•21 comments

A 12-year sequence of telescope images of a star and four planets orbiting

https://bsky.app/profile/theplanetaryguy.com/post/3mwucf5ert22f
52•mariuz•8h ago•10 comments

GrapheneOS has fixed the Android 17 QPR1 kernel performance regression

https://discuss.grapheneos.org/d/42511-grapheneos-has-fixed-the-massive-android-17-qpr1-kernel-pe...
16•Cider9986•19m ago•0 comments

Greg Kroah-Hartman – Security in the LLM Age [video]

https://www.youtube.com/watch?v=NnV_cWeoo5Q
69•usernomdeguerre•17h ago•8 comments

On social reality in China

https://www.lesswrong.com/posts/b5cSYh4emQb2qrGmK/on-social-reality-in-china
51•thicTurtlLverXX•8h ago•31 comments

Loss of cell identity drives human aging: Two new papers

https://erictopol.substack.com/p/loss-of-cell-identity-drives-human
75•bookofjoe•1d ago•11 comments

Mike Tomlin spent 12 years building a Minecraft city

https://www.nytimes.com/athletic/7648198/2026/10/01/mike-tomlin-minecraft-nfl-coach/
72•CoryOndrejka•1d ago•15 comments

Blogging with Gleam, Org-Mode and Pandoc

https://byzantine-systems.github.io/blogging-with-gleam-org-mode-and-pandoc/
21•schonfinkel•8h ago•2 comments

Sites in ChatGPT

https://chatgpt.com/features/sites/
126•polvi•21h ago•137 comments

Our Project Suncatcher prototype satellite is in orbit

https://blog.google/innovation-and-ai/models-and-research/google-research/project-suncatcher-prot...
25•pantalaimon•8h ago•24 comments

One month coding with GLM 5.3 Flash

https://wagtail.org/blog/one-month-on-glm-53-flash/
36•ThibWeb•4h ago•27 comments

From the creator of Redis; run LLM locally with ds4

https://dwarfstar.sh/
26•fibo•2h ago•2 comments

Updates to Full Disk Access in macOS

https://developer.apple.com/news/?id=p6zjojqw
16•notfirstpost•27m ago•2 comments

FLUX 3 Image

https://bfl.ai/models/flux-3-image
199•minimaxir•1d ago•48 comments

The Legend of von Neumann (1973) [pdf]

https://gwern.net/doc/math/1973-halmos.pdf
210•suopspaces•6h ago•123 comments

STS-51-F Abort-to-Orbit (1985)

https://en.wikipedia.org/wiki/STS-51-F
12•schoen•3h ago•4 comments

The first packet sent via RFC1149 avian carrier is up for auction at Christie's

https://onlineonly.christies.com/s/fine-printed-books-manuscripts-science/carrier-pigeon-internet...
13•peter_hansteen•7h ago•1 comments

How accurately calibrated is Jev?

https://maximumeffort.substack.com/p/jev-is-poorly-calibrated
22•dblack12705•4h ago•8 comments

Show HN: Giving Opus 5.5 a simulated paint canvas

https://stillwet.art/
143•alstonite•19h ago•46 comments

Venice’s failed war against Constantinople led to the first bond market

https://bigthink.com/books/a-fabulous-debt/
19•RickJWagner•6h ago•5 comments

"The only intuitive interface is the nipple" (2012)

https://www.greenend.org.uk/rjk/misc/nipple.html
13•ibobev•46m ago•11 comments

What if we stopped using GPUs? [video]

https://www.youtube.com/watch?v=xc2FTBGRSJo
14•sandslash•1h ago•3 comments

100 years of student radio history in the DLARC college radio collections

https://blog.archive.org/2026/10/02/100-years-of-student-radio-history-in-the-dlarc-college-radio...
16•HieronymusBosch•7h ago•1 comments

Anatomy of a Lean proof for software engineers

https://agostbiro.net/posts/2026-10-anatomy-of-a-lean-proof/
28•abiro•1d ago•0 comments

F.02 Decommission

https://www.figure.ai/news/f-02-decommission
22•ad_hockey•9h ago•5 comments

Show HN: Pyxel – A Python retro game engine with built-in art and sound editors

https://github.com/kitao/pyxel
25•kitao•20h ago•4 comments

Supabase is acquiring Turso

https://supabase.com/blog/supabase-is-acquiring-turso
175•cvburgess•4h ago•90 comments