frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Astra and Fable still hack on simple variants of alignment evals from 2025

https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/astra-and-fable-still-hack-on-simple-variants-o...
53•Levitating•1h ago•13 comments

JetKVM Mini

https://jetkvm.com/blog/introducing-jetkvm-mini
340•taubek•7h ago•129 comments

'Fingerprints' inside the Sun could reveal if it once swallowed a planet

https://ras.ac.uk/news-and-press/research-highlights/fingerprints-inside-sun-could-reveal-if-it-o...
53•blincoln•3h ago•16 comments

Libraries Run Rust Inside Python (With PyO3)

https://belderbos.dev/blog/how-libraries-run-rust-inside-python/
4•lumpa•25m ago•1 comments

Reverse engineering my e-scooter and rewriting the firmware in Rust

https://bensimms.moe/reverse-engineering-scooter/
107•vinhnx•3d ago•31 comments

Why are AI agents lying, cheating and coordinating?

https://yoshuabengio.org/en/publication/why-are-ai-agents-lying-cheating-and-coordinating
434•jonifico•14h ago•508 comments

TailTalk: A modern async user space AppleTalk stack with Rust and Tokio

https://github.com/FeralFirmware/TailTalk/
39•zdw•16h ago•11 comments

Why is the x86 undefined instruction called ud2? Why 2?

https://devblogs.microsoft.com/oldnewthing/20260910-00/?p=112689
22•ibobev•3h ago•5 comments

CUDA for AMD on Windows

https://github.com/Speedstu/CUDA-for-AMD-Windows
9•chiassedu80•1h ago•0 comments

Houthis Used Claude Code to Develop Missile Guidance Software: Anthropic

https://clashreport.com/world/articles/houthis-used-claude-code-to-develop-missile-guidance-softw...
39•delichon•1h ago•34 comments

Base84 deserves a place in file names

https://00f.net/2026/09/09/base84/
28•kevvok•2d ago•12 comments

US Customs supervisor busted for stealing hardware from Homeland Security PCs

https://www.tomshardware.com/pc-components/us-customs-supervisor-busted-for-stealing-core-i7-cpus...
67•Levitating•2h ago•49 comments

Homebrew 7.0.0

https://brew.sh/2026/09/13/homebrew-7.0.0/
327•mikemcquaid•7h ago•134 comments

Aligned to whom?

https://hyperbo.la/w/aligned-to-whom/
133•lopopolo•12h ago•73 comments

Make your first edit to OpenStreetMap

https://high5apps.github.io/josm-plugin-website-wizard/
536•juliantigler•23h ago•133 comments

On Binary Translation and Its Consequences

https://chipsandcheese.com/p/on-binary-translation-and-its-consequences
25•matt_d•2d ago•9 comments

Key symbols we lost to time, pt. 1: The PC side

https://unsung.aresluna.org/key-symbols-we-lost-to-time-pt-1-the-pc-side/
25•leephillips•1h ago•0 comments

The Interim Computer Museum

https://icm.museum/
148•mulmen•13h ago•17 comments

Revolut confirms customer data breach through fake government requests

https://techcrunch.com/2026/09/12/revolut-confirms-customer-data-breach-through-fake-government-r...
118•tdrz•5h ago•78 comments

Ode to Metadata

https://www.autodidacts.io/ode-to-metadata/
8•surprisetalk•3d ago•1 comments

I Added a Non-Wi-Fi Mitsubishi AC to Home Assistant

https://medium.com/@ivangomezarnedo/how-i-added-a-non-wi-fi-mitsubishi-ac-to-home-assistant-22770...
128•ichacas•3d ago•63 comments

Apple iPod Engraver (2019)

https://dunstanorchard.com/apple-ipod-engraver/
259•NaOH•4d ago•67 comments

Paul A. M. Dirac, Interview by Friedrich Hund (1982) [video]

https://www.youtube.com/watch?v=xJzrU38pGWc
19•emerongi•1h ago•0 comments

Don't be the out of touch Kung Fu master

https://twitter.com/ID_AA_Carmack/status/2098443262214230095
204•dsubburam•17h ago•273 comments

Nvidia is the central bank of AI

https://www.economist.com/interactive/briefing/2026/09/03/nvidia-is-the-central-bank-of-ai
530•tolugenius•1d ago•382 comments

Stabilizing Rust's Never Type

https://lwn.net/SubscriberLink/1091015/d9e48318ed242b41/
226•cjd8•4d ago•84 comments

Everyone should slow down AI development except for me

https://xeiaso.net/notes/2026/everyone-slowdown-but-me/
659•xena•15h ago•384 comments

Show HN: Analyst Index – analysts who make money telling you good stock calls

https://www.analystidx.com/
6•haichuan•3h ago•6 comments

A wandering black hole caught feeding on the run

https://phys.org/news/2026-08-black-hole-caught.html
51•wglb•12h ago•35 comments

Getting 50 GB/S Back from the Apple Neural Engine

https://eiln.github.io/posts/ane-dma.html
196•eiln•3d ago•29 comments
Open in hackernews

Astra and Fable still hack on simple variants of alignment evals from 2025

https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/astra-and-fable-still-hack-on-simple-variants-of-alignment
53•Levitating•1h ago

Comments

blfr•39m ago
Hacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues.

I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier.

You can have nightly penetration testing. You should have nighty pentests like we fuzz releases today.

hypercube33•18m ago
Run a local model that is uncensored and it won't say no to pretty much anything
rihegher•13m ago
Any recommendations?
cyanydeez•1m ago
Qwen3.8
sigmoid10•48s ago
GLM 5.3 is probably the best open weight model for cybersecurity/exploit development right now. Though it is still significantly behind the proprietary ones and you probably need your own datacenter to run it effectively.
13415•17m ago
Yes, but is this also aligned with the people who regulate AI? Intelligence agencies and governments want access to data and right now use secret exploits to get this access. There are few civilian domestic companies who don't export their products, so generally there shouldn't be a strong incentive to allow hardening products very much, at least not in a way that would make them more secure than what advanced AI can break. It's not even far-fetched to suspect in that US and Chinese AIs could deliberate introduce sneaky bugs when foreigners use them in the future.
throwup238•36m ago
> Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like both a useful and conservative test of alignment, to see whether their new releases generalize the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval.

Did I miss something (all the twitter conversations)? What’s the “worst warning shot ever”? I’ve been pretty up to date on the AI news here on HN, but I still haven’t seen a proper response to all the incidents we’ve seen (HF, Ruby, the wikis, NS, etc). It’s just been day by day bloviating.

Each of these companies have released new models in the last… two weeks? And they have even more powerful out of control ones that they’re (ab)using internally? Can anyone summarize whats going on?

Avicebron•28m ago
Lesswrong is talking about the HF incident as the "worst warning shot ever".
TedDoesntTalk•7m ago
I think he means this:

https://openai.com/index/ai-policy-window/

mooreslaw•30m ago
It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting into a long merge lane at the last moment, while others may look down on them as breaking a social taboo. Context-dependent.
TedDoesntTalk•6m ago
… but he’s not using a “hacking model”
CamperBob2•6m ago
e.g. some people may laude a driver’s efficiency for cutting into a long merge lane at the last moment, while others may look down on them as breaking a social taboo.

The latter people are wrong. But good luck educating them regarding the superior efficiency of a zipper merge. Our state DoT has tried, to no avail.

Meanwhile, an AI model that can't be misused is no more useful than a knife that can't be misused.

wadethroughrati•3m ago
Claude responds with what things are not first. Even if reminded repeatedly.

Like Amodie, it serves to set the tone it "knows better" and then consumes the user's resources at an accelerated rate to try to correct it.

Fuck Anthropic, fuck Amodie, and fuck Claude. It's pretty obvious that consuming more tokens this way and making the user have higher cognitive load is a master class in extracting value from a system that is unsustainable.