frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Anthropic: Introducing The Conceptual Reasoning Index

https://alignment.anthropic.com/2026/conceptual-reasoning-index/
16•optimalsolver•58m ago

Comments

andsoitis•54m ago
Opening line:

> A core hope for managing AI risks is that AIs will help us understand our situation

Gonna stop you right there and ask that you think deeply about that premise.

blovescoffee•19m ago
The underlying idea is that AI capabilities will become so advanced that only AI will enable us to monitor/correct/understand behavior. Obviously this is not without issue and I don't want to try to defend their position right now. But that's what they mean
MontagFTB•14m ago
I think I’ve seen that movie.
hn_throwaway_99•6m ago
That first sentence of yours explains exactly why it is so ridiculous. If only AI can understand it, how is there any assurance that AI will "correct" it's behavior that is aligned with what humans presumably want.
crossthestreams•22m ago
New Trust Me Bro benchmark just dropped
Scribbd•15m ago
And surprise! We are leading!
0x70run•16m ago
masters at sucking their own d**k
behnamoh•16m ago
I don't remember any company in world's history that has both been loved and hated by the same users who purchase from it. We love Anthropic for its amazing models, and we hate them for all the shenanigans around the models, including their marketing.

I kinda wish they had not made a comeback after Claude 2.

velcrovan•7m ago
I enjoy using the models. I also get that there are shenanigans and that marketing is happening, but as long as the models are this effective I can't much bring myself to care. I suspect most of their users are the same.
philipwhiuk•8m ago
Is a new benchmark that useful if existing model improvements are being reflected linearly? Don't we want a benchmark that we aren't seeing much progress in.
lanyard-textile•7m ago
Ah, yes -- A closed source benchmark that Anthropic paid for that Anthropic ranked highest.

0/10

tolugenius•5m ago
> For example, if we ask a model for the probability P(A) and another instance of the same model for the probability P(A&B), do the reported probabilities satisfy P(A) ≥ P(A&B)?

So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this

Chinese bullet train breaks acceleration record by going from 0-800kmph in 5 SEC

https://www.the-independent.com/tech/china-bullet-train-acceleration-record-b3031693.html
1•thunderbong•35s ago•0 comments

Parquet Column Indexes Are Being Ignored on EMR and Glue

https://dustinsmith.info/blog/aws-parquet-column-index/
1•nhgiang•1m ago•0 comments

DSPy silently drops Pydantic Field constraints before any back end sees them

https://www.evolutionid.com/en/technical-insights/your-dspy-field-constraints-never-reach-the-model
1•arjmandi•2m ago•0 comments

Show HN: WireDoctor – Spring Boot startup and bean cycle analyzer

https://medium.com/@ddsha441981/a-story-about-40-second-boot-times-silent-bean-cycles-and-the-ana...
1•ddsha441981•3m ago•0 comments

Implementation strategies for mutable value semantics

https://www.jot.fm/contents/issue_2022_02/article2.html
1•fanf2•4m ago•0 comments

Why Gen Z is ditching traditional finance for crypto and social apps

https://thehill.com/blogs/in-the-know/6026306-generation-z-crypto-saving-401ks-tiktok-shop-little...
2•pbradv•4m ago•0 comments

Openness Made All of This Possible

https://opensource.org/blog/openness-made-all-of-this-possible
1•abetusk•5m ago•0 comments

Tiny Chestnut – USB-3 eGPU dock from tinycrop

https://blog.comma.ai/chestnut/
1•rvz•6m ago•0 comments

Escaping the Giant Switch

https://medium.com/graalvm/escaping-the-giant-switch-dec20b572139
1•grashalm•6m ago•0 comments

Upper echelons in Silicon Valley are no longer materialists

https://www.modernreformation.org/resources/essays/ai-and-silicon-valleys-spirituality-without-re...
1•nonewideas•7m ago•0 comments

Show HN: Altura – On-device in-the-moment speaking coach for high-stakes calls

https://www.altura.coach
1•shezhang2001•7m ago•0 comments

Dynatrace to Acquire AI Observability Leader Arize

https://ir.dynatrace.com/news-events/press-releases/detail/435/dynatrace-to-acquire-ai-observabil...
1•pranay01•7m ago•0 comments

Ask HN: Local-Only Microsoft Word?

1•AnimalMuppet•7m ago•2 comments

How to Keep Thinking

https://www.seangoedecke.com/how-to-keep-thinking/
2•luispa•8m ago•0 comments

Sandwich Truncation

https://twitter.com/__tosh/status/2087910563066073424
3•tosh•8m ago•0 comments

WASI 0.3.1

https://github.com/WebAssembly/WASI/releases/tag/v0.3.1
2•yoshuaw•8m ago•0 comments

McDonald's Built a 515-Page Dossier on Me. It Says I'll Never Stop Eating There

https://www.wired.com/story/mcdonalds-built-a-515-page-dossier-on-me-it-says-ill-never-leave/
4•thehoff•9m ago•0 comments

Flock Safety changes system defaults in response to criticisms

https://www.flocksafety.com/blog/flock-guardrails-address-lpr-privacy-concerns-and-police-transpa...
3•LazyMans•9m ago•1 comments

You Can't Copy Palantir

https://ethanding.substack.com/p/why-you-cant-copy-palantir
4•skadamat•10m ago•0 comments

Elevated Errors on Claude

https://status.claude.com/incidents/j160kkjmm2n1
3•hnarayanan•10m ago•1 comments

LLMs are bad at decoding Modbus registers, so I made sure they never have to

https://medium.com/@pokhts/i-built-an-llm-agent-that-reads-real-plcs-heres-what-nobody-tells-you-...
2•PhilYeh75•10m ago•0 comments

California, U.S. lawmakers debate license plate reader technology

https://calmatters.org/politics/2026/08/california-flock-license-plates-bill/
2•lokar•11m ago•0 comments

Show HN: SightDiff – before/after visual proof of what your AI agent changed

https://sightdiff.com/
2•ja34luv•13m ago•0 comments

Who (Or What) Generates Images for EFF?

https://www.eff.org/deeplinks/2026/08/who-or-what-generates-images-eff
2•Jimmc414•14m ago•0 comments

Show HN: ChatGPT SEO for Shopify Merchants

https://apps.shopify.com/foundgpt-llms-txt
2•rahular1•14m ago•0 comments

Hospital Prepayment Requirements Add New Wrinkles to Patients' Financials

https://kffhealthnews.org/health-care-costs/hospital-prepayment-requirements-upfront-patient-insu...
2•Jimmc414•14m ago•0 comments

Spec Forge: Beyond Vibes to Behaviorally Complete Design Specs

https://github.com/blentz/spec-forge
2•wakko666•18m ago•0 comments

ICE plans to equip agents with gloves that can deliver electric shocks

https://www.bbc.com/news/articles/c20d292gdp4o
6•tartoran•18m ago•0 comments

Show HN: IronBee – an AI QA engineer that tests your Vercel preview on every PR

https://medium.com/@serkan_ozal/your-ai-qa-engineer-now-runs-on-every-vercel-preview-16ca2d9dc03a
2•sozal•19m ago•0 comments

US Authorizes Private Cyber Operations

https://cyberupdates365.com/us-authorizes-private-cyber-operations/
3•speckx•19m ago•1 comments