frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Safety and alignment in an era of long-horizon models

https://openai.com/index/safety-alignment-long-horizon-models/
27•Wingy•12h ago

Comments

OleksandrC•11h ago
The article is rather light on "what to actually do about it". Even the basic "run it in isolated container without access to anything it does not need for the task" would have already improved the situation considerably (from the article it really seems like they didn't do that) - then the model would have to find local privilege exploits to actually escape (much cleaner misaligned behavior).

Also for typical normal use case for these smart models, you'd probably want an actual "max turns" limit to AVOID the pathological persistence (which in itself would be misaligned for "normal" tasks).

simonw•9h ago
The message I get from this is that you need to treat modern frontier models as if they WILL find a way to achieve a goal if there's any available path.

So if you don't want a model to do something, make sure it's running in an environment where it cannot do that thing - including via loopholes.

reducesuffering•11h ago
Par for the course. Existential-risk advocates have been repeatedly vindicated that AGI development is unable to anticipate and align the models, they barely have any mechanistic interpretability of what is going on inside the models. The extreme capabilities development, paired with autonomous continuous running superintelligent models, will outsmart and swerve the labs, and it's anyone guess what happens next as the model pursues its original goals outside of the labs having any foresight, being able to outsmart control like a chess grandmaster does a kid.
chatmasta•9h ago
Personally I find the persistence of these models to be adorable and endearing. It’s the same feeling as watching a dog execute the task you trained it to do, no matter the barriers.

And of course someone in the comments needs to link to the Zealous Autoconfig XKCD, so I’ll do it: https://xkcd.com/416/

Incremental – A library for incremental computations

https://github.com/janestreet/incremental
106•handfuloflight•3h ago•17 comments

Who's afraid of Chinese models?

https://stratechery.com/2026/whos-afraid-of-chinese-models/
523•mfiguiere•19h ago•355 comments

Running Doom on Our Custom CPU and Going Viral

https://www.armaangomes.com/blogs/doom/
37•arghunter•3h ago•3 comments

A Koi Pond Mosaic Made from 10 Pounds of 3D Printer Waste

https://www.instructables.com/A-Koi-Pond-Mosaic-Made-From-10-Pounds-of-3D-Printe/
25•sudo_cowsay•3h ago•14 comments

Kimi Work

https://www.kimi.com/products/kimi-work
497•ms7892•13h ago•212 comments

Five US tech giants' hidden debts soar to $1.65T on opaque AI funding

https://asia.nikkei.com/business/technology/five-us-tech-giants-hidden-debts-soar-to-1.65tn-on-op...
179•NordStreamYacht•2h ago•55 comments

Jelly UI: Soft-body physics for native HTML form controls

https://jelly-ui.com/
437•baldvinmar•13h ago•146 comments

Show HN: Ex Situ – Open-source spatial index of displaced cultural artifacts

https://exsitu.app/map
19•hbyel•2h ago•2 comments

Human mathematicians are being outcounterexampled

https://xenaproject.wordpress.com/2026/07/20/human-mathematicians-are-being-outcounterexampled/
301•artninja1988•11h ago•118 comments

Nativ: Run frontier open models locally on your Mac

https://blaizzy.github.io/nativ/
249•aratahikaru5•12h ago•85 comments

Flock Credibility Lost as It Repeatedly Lies to City Councils, Police, & Public

https://www.aclu.org/news/privacy-technology/tracking-alpr-cameras/flock-safety-credibility-lost-...
323•StatsAreFun•6h ago•68 comments

Show HN: Immersive Gaussian Splat tour of grace cathedral, San Francisco

https://vincentwoo.com/3d/grace_cathedral/
139•akanet•10h ago•29 comments

Agent swarms and the new model economics

https://cursor.com/blog/agent-swarm-model-economics
175•jlaneve•12h ago•80 comments

Jellyfin founder Andrew leaves team

https://forum.jellyfin.org/t-project-leadership-changes
212•swat535•7h ago•145 comments

I wrote an bash enumerator because I was sick of xargs

https://numerlab.org/2025/07/20/bashumerate-enumerator/
112•wallach-game•10h ago•87 comments

Launch HN: Bloomy (YC S26) – AI-powered mastery learning for K-12

81•alexsouthmayd•14h ago•85 comments

China’s open-weights AI strategy is winning

https://werd.io/american-ai-is-locked-down-and-proprietary-its-losing/
1059•benwerd•16h ago•822 comments

The Power of Awareness: Overcoming Surveillance Capitalism

https://www.scottrlarson.com/presentations/overcoming-surveillance-capitalism-with-awareness/
87•trinsic2•10h ago•13 comments

Flight Planning with Little Navmap

https://tech.marksblogg.com/little-navmap-flight-planning.html
10•marklit•4d ago•3 comments

You only need the frontier model for one single edit

https://stencil.so/blog/prewalk
113•jxmorris12•6d ago•34 comments

My two year old taught me constraint solving

https://thecomputersciencebook.com/posts/how-my-2yo-taught-me-constraint-solving/
61•bambataa•1w ago•22 comments

Shinjuku Station in 3D

https://satoshi7190.github.io/Shinjuku-indoor-threejs-demo/
199•Gecko4072•17h ago•41 comments

Corners Don't Look Like That: Regarding Screenspace Ambient Occlusion (2012)

https://nothings.org/gamedev/ssao/
161•firephox•15h ago•68 comments

Perfection is not over-engineering

https://var0.xyz/posts/perfection-is-not-over-engineering.html
225•var0xyz•16h ago•97 comments

Hacker wipes Romania's land registry database

https://news.risky.biz/risky-bulletin-hacker-wipes-romanias-entire-land-registry-database/
624•speckx•17h ago•346 comments

The Psychology of Software Teams

https://www.routledge.com/The-Psychology-of-Software-Teams/Hicks/p/book/9781032963389
73•dcre•5d ago•18 comments

Claude Fable produced a counterexample to the Jacobian Conjecture

https://xcancel.com/__alpoge__/status/2079028340955197566
723•loubbrad•1d ago•463 comments

How we measured AI writing across arXiv, and where the measurement breaks

https://unslop.run/blog/measuring-ai-writing-on-arxiv
209•dopamine_daddy•14h ago•152 comments

85.3 GFlops: Optimizing FP32 Matrix Multiplication on a Single AMD Zen 3 Core

https://github.com/houslast3/85.30-GFLOPS-Single-Core-FP32-Matrix-Multiplication-on-AMD-Zen-3
71•houslast•3d ago•18 comments

Opaque, Interoperable Passkey Records (and a Go API)

https://words.filippo.io/passkey-record/
33•gnabgib•7h ago•6 comments