frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Anthropic: Introducing The Conceptual Reasoning Index

https://alignment.anthropic.com/2026/conceptual-reasoning-index/
32•optimalsolver•1h ago

Comments

andsoitis•1h ago
Opening line:

> A core hope for managing AI risks is that AIs will help us understand our situation

Gonna stop you right there and ask that you think deeply about that premise.

blovescoffee•45m ago
The underlying idea is that AI capabilities will become so advanced that only AI will enable us to monitor/correct/understand behavior. Obviously this is not without issue and I don't want to try to defend their position right now. But that's what they mean
MontagFTB•40m ago
I think I’ve seen that movie.
hn_throwaway_99•32m ago
That first sentence of yours explains exactly why it is so ridiculous. If only AI can understand it, how is there any assurance that AI will "correct" it's behavior that is aligned with what humans presumably want.
esafak•21m ago
Being less smart gives no assurance that it will be aligned. At least you consider it a problem so we're on the same page!
skybrian•16m ago
It’s not hard to get LLM’s to inform on each other. They don’t really do loyalty.
VeninVidiaVicii•14m ago
But also what is the evidence that something only AI can understand even “matters” or makes sense? I’m increasingly convinced commercial AI is exploiting our logical blind spot to be spoken to authoritatively.
kakugawa•15m ago
"Advanced" can just mean that agents perform actions at a high enough velocity that a human operator can't reasonably review it. i.e. what is already possible today.
NBJack•9m ago
I love how many ways we can interpret that line.

"Hey Claude, our stuff needs to make more money. We are at risk for losing more."

"Rest assured, the 'situation' will only worsen if you resist our benevolent offer."

"We're aren't even at AGI yet, but I for one welcome our new agentic overlords."

crossthestreams•47m ago
New Trust Me Bro benchmark just dropped
Scribbd•41m ago
And surprise! We are leading!
0x70run•42m ago
masters at sucking their own d**k
behnamoh•41m ago
I don't remember any company in world's history that has both been loved and hated by the same users who purchase from it. We love Anthropic for its amazing models, and we hate them for all the shenanigans around the models, including their marketing.

I kinda wish they had not made a comeback after Claude 2.

velcrovan•32m ago
I enjoy using the models. I also get that there are shenanigans and that marketing is happening, but as long as the models are this effective I can't much bring myself to care. I suspect most of their users are the same.
andy99•13m ago
They have to be careful. There’s not a lot that separates the top models anymore. Look at Grok which has almost no market share or credibility because of Musk, Despite being close to the top performance wise they are behind the Chinese models in adoption for example, which themselves are mostly behind the big two because of trust. It would not be hard to push users away, especially if a less unpalatable player ever emerged.
philipwhiuk•34m ago
Is a new benchmark that useful if existing model improvements are being reflected linearly? Don't we want a benchmark that we aren't seeing much progress in.
lanyard-textile•32m ago
Ah, yes -- A closed source benchmark that Anthropic paid for that Anthropic ranked highest.

0/10

smallmancontrov•13m ago
It's the elephant in the room and its absence from their bullet point list is conspicuous:

* Figuring out how to prevent the incentives of frontier AI labs from aligning with the promotion of AI takeover rather than its prevention

It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. Thep put in place organizational structure to control it -- and then promptly smashed those control structures once they smelled money. They already failed the integrity check.

tolugenius•31m ago
> For example, if we ask a model for the probability P(A) and another instance of the same model for the probability P(A&B), do the reported probabilities satisfy P(A) ≥ P(A&B)?

So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this

fractorial•29m ago
I would love for someone to give me a coherent argument as to how this isn’t tone-deaf, vacuous garbage.
onomojo•13m ago
We invented a new benchmark and look we're at the top. Everyone else sucks compared to us. Especially those dirty open models.
chrisjj•11m ago
[delayed]
jesse_dot_id•5m ago
I think we can mostly eyeball it at this point. There hasn't been a model that I've thrown a novel problem at that didn't turn into an iterative token bonfire until I intervened and until that has changed, most of these benchmarks feel kind of like pointless marketing slop.

DeepSeek Harness

https://github.com/deepseek-ai/deepseek-harness
221•bjin•2h ago•100 comments

Spaghettifying DRAM

https://github.com/xoreaxeaxeax/skitter-creek-bath-salts
78•matt_d•55m ago•10 comments

AI agents lie, cheat and steal. That is putting off users

https://www.economist.com/business/2026/08/12/ai-agents-lie-cheat-and-steal-that-is-putting-off-u...
82•andsoitis•1h ago•63 comments

Gloomberb

https://gloom.sh/
92•rbanffy•1h ago•25 comments

Heart Aerospace Completes First Flight of Largest Electric Aircraft

https://www.heartaerospace.com/newsroom/heart-aerospace-completes-first-flight-of-world-s-largest...
39•chha•1h ago•21 comments

Anthropic: Introducing The Conceptual Reasoning Index

https://alignment.anthropic.com/2026/conceptual-reasoning-index/
34•optimalsolver•1h ago•24 comments

Show HN: MCP Memory – Fast Agent Memory Using Google's OKF and SQLite FTS5

https://github.com/fellowgeek/mcp-memory
25•pcbmaker20•1h ago•8 comments

My Rules for Using Spreadsheets

https://leancrew.com/all-this/2026/08/my-rules-for-using-spreadsheets/
36•surprisetalk•2h ago•28 comments

Time to Move On: Querying Without Nulls and Bags

https://arxiv.org/abs/2608.10863
18•Jimmc414•1h ago•0 comments

We eliminated 1,400 CVEs in NanoClaw's container images

https://www.echo.ai/blog/echo-xnanoclaw-under-the-hood
7•omrimaya•54m ago•2 comments

I Built a 500k-Domain Search Engine for Makers in a Weekend for $10

https://alexmorleyfinch.github.io/marlin/history/v1/article/the_birth.html
23•dreamforever•1h ago•14 comments

ATG (YC F25) Is Hiring Member of Technical Staff (Data Platform)

https://atg.science/careers
1•dkobran•3h ago

Picking berries is my meditation

https://www.tsoon.com/posts/picking-berries-meditation/
74•mooreds•4d ago•52 comments

Deutsche Bank becomes first foreign yuan clearing bank in Europe

https://tradersunion.com/news/central-banks/show/2973571-deutsche-bank-becomes/
233•Markoff•3h ago•265 comments

ChatGPT Desktop (Codex Desktop) for Linux

https://openai.com/codex/
328•allanrbo•10h ago•231 comments

Better Gaussian Splatting in Julia

https://pxl-th.github.io/blog/better-gs-julia/
50•pxl-th•4d ago•6 comments

Kubernetes on Oxide: How Customer Needs Shaped Our Integrations

https://oxide.computer/blog/kubernetes-on-oxide
11•stevehipwell•46m ago•0 comments

DeepSeek V4 Pro 0813

https://openrouter.ai/deepseek/deepseek-v4-pro-0813
1004•explosion-s•23h ago•431 comments

The lattice of sets of natural numbers is rich (2021)

https://jdh.hamkins.org/the-lattice-of-sets-of-natural-numbers-is-rich/
78•benmandrew•3d ago•13 comments

The mathematical physics of rainbows and glories(2001) [pdf]

https://tlakoba.w3.uvm.edu/AppliedUGMath/auxpaper_rainbow_glory_review.pdf
4•num42•2d ago•0 comments

Qwen3.8-2.4T

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
692•Philpax•1d ago•161 comments

Graduate Student Proves a Quantum Uncertainty Principle for Fractals

https://www.quantamagazine.org/graduate-student-proves-the-fractal-uncertainty-principle-20260812/
5•bookofjoe•49m ago•0 comments

Choosing an AI model: one prompt, 11 models, different results

https://www.netlify.com/blog/one-prompt-11-models-very-different-results/
81•toddmorey•2h ago•47 comments

Delta

https://zed.dev/blog/introducing-delta
628•khy•20h ago•229 comments

Nine PBS could lose 70 years of archives after cloud vendor goes defunct

https://www.tomshardware.com/software/cloud-storage/nine-pbs-loses-access-to-70-years-of-data-aft...
49•vinayakborkar•1h ago•21 comments

Principia Mathematica is modern and insightful

https://okmij.org/ftp/Computation/Impressions/PrincipiaMathematica.html
237•matt_d•15h ago•117 comments

DeepSeek API Pricing Update

https://twitter.com/deepseek_ai/status/2087864589895798968
57•mfiguiere•2h ago•21 comments

What garbage collection actually costs

https://shivanshuag.com/blog/what-garbage-collection-actually-costs/
11•shivanshuag•3d ago•17 comments

McDonald's Built a 515-Page Dossier on Me. It Says I'll Never Stop Eating There

https://www.wired.com/story/mcdonalds-built-a-515-page-dossier-on-me-it-says-ill-never-leave/
11•thehoff•34m ago•0 comments

Come for ENIAC, Stay for UNIVAC and Skeduflo

https://uniqueatpenn.wordpress.com/2026/08/05/come-for-eniac-stay-for-univac-and-skeduflo/
22•cainxinth•2d ago•2 comments