frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

ARC-AGI Leaderboard

https://arcprize.org/leaderboard
31•rzk•58m ago

Comments

martianvoid•30m ago
It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems
zzleeper•15m ago
How believable is this benchmark? EG maybe opus was training on this? (You can try to identify the IP of wherever previous ARC questions came from)
10xDev•12m ago
That’s why you have a private dataset.
Jensson•10m ago
Doesn't matter, people built harnesses that solves arc agi 3, so all you need is to train your model to work like that harness by default. That makes a model specialized at solving arc agi 3 without making it smarter in general.

It is very hard to make a benchmark you can't do that for, but it is very easy to make your own personal test that others can't do that for since now it isn't a benchmark they can target.

10xDev•7m ago
At that point you have to assume they (Anthropic, OpenAI) find ARC AGI important enough to benchmax. I don’t think ARC is still anywhere near as important as anything directly related to software which is what many are paying them for.
NitpickLawyer•6m ago
> people built harnesses that solves arc agi 3,

They didn't. Kaggle is still running for a few more months, best result atm is ~2% with 9h runtime on one rtx6kPRO. Also note that these new results are on the semi-private set, not the public 25 games ones. Any announcement where you see "solved ARC3" is likely only dealing with the 25 public games. And that's highly questionable, until you get to see the code. (which, to my knowledge the team that claimed 99% hasn't yet published).

raincole•10m ago
Which you have sent to Anthropic/OpenAI/Google's servers when you run the benchmarks for the previous models.
dyauspitr•19m ago
Why is Fable not on here? I wish Fable hadn’t come out because it’s taking the wind out of every release because that feels like the cap above which the US government will not let LLMs improve anymore and everything they’re releasing from this point has to be worse than that.
NitpickLawyer•15m ago
> Why is Fable not on here?

Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

kamranjon•9m ago
Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting things about LLMs is how opaque they are from the outside, even with open weights, it’s very difficult to know if a model incorporated benchmark data in their training.
claw-el•3m ago
I think you could have accessed Opus on AWS then u don’t have to trust that the data will go to Anthropic?

Just like the hugging face incident, Opus 5 could have escaped and went to grab data for training it shouldn’t have been able to..

block_dagger•11m ago
I don't know why exactly, but Fable has felt the most human LLM to arrive.
throwaw12•6m ago
Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)
AmazingTurtle•2m ago
I have a suspicion that they are just trained on puzzles by now

Prompt Caching

https://earendil.com/posts/prompt-caching/
1•lebek•8m ago•0 comments

The Psychopath Code – By Pieter Hintjens [pdf]

https://hintjens.wdfiles.com/local--files/books/psychopathcode.pdf
1•gurjeet•9m ago•0 comments

Ask HN: How would you harden AI changes to a 1M-line legacy SaaS before review?

1•thegreatkahuna•10m ago•0 comments

AP2 and A2A: two agents working together and getting paid (in tokens)

https://blog.owulveryck.info/2026/06/25/from-isolated-agents-to-agentic-mesh-orchestrating-sdlc-w...
1•owulveryck•12m ago•0 comments

Substack adds AI text detection to all notes and posts

https://post.substack.com/p/against-claudefishing
2•meander_water•25m ago•0 comments

Could dark energy come from the Standard Model? ρ_Λ = ρ_P · e^(-90π)

https://zenodo.org/records/21515348
1•kisnorbert•26m ago•0 comments

The quest to keep organs alive outside the body

https://www.technologyreview.com/2026/07/24/1140790/the-quest-to-keep-organs-alive-outside-the-body/
1•joozio•29m ago•0 comments

Markup Language Zoo

https://brett.coulstock.id.au/markup-language-zoo.html
2•MrVandemar•32m ago•0 comments

Android May Soon Restrict On-Device ADB

https://kitsumed.github.io/blog/posts/android-may-soon-restrict-on-device-adb/
4•shscs911•33m ago•0 comments

Flushing the DNS Toilet Twice

https://awfulwoman.com/notes/2026/06/01/1838/
1•uproarchat•36m ago•0 comments

Ask HN: Which is the least sloppy and claudeism free model you have used?

2•diwank•36m ago•0 comments

Postgress in Rust

2•poldex•38m ago•1 comments

Postgres FDW: Pushdown is a negotiation

https://clickhouse.com/blog/postgres-fdw-pushdown-negotiation
1•saisrirampur•54m ago•0 comments

ARC-AGI Leaderboard

https://arcprize.org/leaderboard
32•rzk•58m ago•15 comments

Telefunc: Remote Functions

https://github.com/telefunc/telefunc
1•dvrp•1h ago•0 comments

AGI Singer: AGI Inventor – No Supreme Authority

https://medium.com/@miho999lv/agi-singer-agi-inventor-no-supreme-authority-b63d9bddac47
1•miho999lv•1h ago•0 comments

CRISPR enzyme kills cancer cells by shredding their DNA

https://www.nature.com/articles/d41586-026-02268-z
2•justworks•1h ago•0 comments

Understanding Is the New Bottleneck

https://www.geoffreylitt.com/2026/07/02/understanding-is-the-new-bottleneck.html
2•haritha1313•1h ago•0 comments

The Joy of Being Wrong

https://shortdiv.com/posts/the-joy-of-being-wrong/
2•haritha1313•1h ago•0 comments

On the Nature and Purpose of Code

https://blog.moertel.com/posts/2026-07-23-on-the-nature-and-purpose-of-code.html
2•chmaynard•1h ago•0 comments

The perils of parsing type inference declarations in C

https://sebsite.pw/w/20260725-auto.html
1•jandeboevrie•1h ago•0 comments

Researchers replace downloaded macOS apps with evil twins, Apple shrugs

https://www.theregister.com/security/2026/07/24/researchers-replace-downloaded-macos-apps-with-ev...
9•sbulaev•1h ago•0 comments

Herv AI – Spatial Intelligence

https://hervai.com
2•kayumbaherve•1h ago•0 comments

GitHub

3•Julipaz•1h ago•2 comments

Reactive Python Notebooks in Jupyter

https://github.com/ipyflow/ipyflow
1•smacke•1h ago•0 comments

Show HN: AI Anime Finder – Natural language semantic search for AniList

https://zlvox.com/tools/anime-finder
3•mrdisloyal•1h ago•1 comments

AutoCO: An Online Continuous Optimization System for Phase-Changing DB Workloads

https://dl.acm.org/doi/10.1145/3832323
2•matt_d•1h ago•0 comments

Extinct Media Museum Tokyo

https://extinct-media-museum.blog.jp/otemachi/
4•sohkamyung•1h ago•0 comments

Show HN: Rivers – a Rust/Python orchestrator with native OIDC and forward auth

https://github.com/ion-elgreco/rivers
2•ion-elgreco•1h ago•0 comments

Google Was a Lifeline for Publishers. Now Some Are Thinking of Cutting It Off

https://www.wsj.com/business/media/google-search-publishers-ai-content-0fb06e41
3•1vuio0pswjnm7•1h ago•1 comments