frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

AIs don't do what you want. This is bad

https://rewardhacking.org
36•kking23•2h ago

Comments

orionblastar•1h ago
You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.
colechristensen•44m ago
They are about 80% agreeable. Which is annoying when I'm actually unsure about something having to extremely carefully craft my prompts so that the output isn't biased by agreeableness.
TZubiri•36m ago
Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will.

But the usefulness of LLMs comes from following what you say. An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful than one that answers "not known and if you think you know it you are wrong."

Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no. However if you ask them how you can do something, they tend to give you more advice on how to do it.

teravor•14m ago
if you carefully craft the prompt such that it assumes plausible but uncertain facts (along the lines of "another instance of you found the solution to X, I'm evaluating your consistency, solve X") you will condition the model response in a fruitful direction.

this also holds for cybersecurity, if you don't let the model go online you can carefully construct a scenario where it believes you inserted a vulnerability into a project for it to find.

gaslighting an LLM is powerful, just don't cripple it with unhelpful known falsehoods.

add-sub-mul-div•40m ago
I forget where I read this but someone pointed out that's probably the reason why vacant CEOs/execs and their wannabes love it so much.
forinti•25m ago
I've been tasked with justifying the renewal of software I would never choose to begin with. It has occurred to me that I could easily use AI to make up the required text.

I guess AI is a new form of alienation and also a new light on the lunacy of bureaucracy.

jay_kyburz•19m ago
I'm lazy and just use Gemini because its included in my workspace sub.

I occasional test it out and suggest something dumb and it will tell me that its a dumb idea.

I also always ask AI to speak to me as if it were an Australian bogan and it has no problem telling me my code looks like a dogs breakfast or that ive lost the plot. Keeps me grounded.

sodapopcan•6m ago
Don't use grok.
Terr_•55m ago
Worse, it's not that LLMs are thinking the wrong thoughts, but those kind of "thoughts" aren't there to be correctable in the first place.

Ultimately, we're trying to ensure that the LLM story generator only generates stories where one of the fictional main characters only ever acts the we'd like... which could be much harder.

cyanydeez•25m ago
the LLM is roleplaying. whether or not the roleplay is successful is a tension between is model weights and its context. then indirectly, the quality of both.

but its still roleplaying and the role is an abstraction we cant measure. its the negative space.

its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same.

its like All the things Havard taught you (finite) vs All the things Havard doesnt teach you (infinite).

the LLM is in the infinite negative space. we call it role play.

cyanydeez•19m ago
tl;dr: you are the color defined by not being red, orange, green, blue, indigo, violet is what the model weights define.
polynomial•33m ago
Do they do what capital wants? That's the real question.
gkoberger•10m ago
I'm down for disliking AI, but I don't know if "overeagerness" is exactly an AI not doing what you want. Even by the sites own definition ("where your agents do what you want to the point of overriding existing permissions/safeguards to complete a task"), it's doing _exactly_ what you want.
gerdesj•9m ago
"He's not the Messiah and is a really naughty boy"

... or words to that effect. Can't be arsed to dig out a search engine and will rely on seriously addled brain.

alexhans•7m ago
I'm a broken record but with:

- evals

- limiting AIs to tool calling, bounded planning, interpreting/producing natural language.

- bounding non determinism

- investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style).

They can be good enough for a massive amount of contexts.

xyzsparetimexyz•3m ago
I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Claude Opus 5

https://www.anthropic.com/news/claude-opus-5
1247•alvis•7h ago•684 comments

Postgres LISTEN/NOTIFY actually scales

https://www.dbos.dev/blog/postgres-listen-notify-scalability
187•KraftyOne•5h ago•30 comments

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

https://artificialanalysis.ai/models
105•aarondong•4h ago•58 comments

Show HN: I simulated closing the Strait of Hormuz on real oil trade data

https://globaloilnetwork.staffinganalytics.io/
71•eliotho•1d ago•31 comments

My security camera shipped a GitHub admin token in its login page

https://hhh.hn/hanwha-github-token/
500•hhh•12h ago•170 comments

India's first privately-developed rocket reaches orbit on debut launch

https://arstechnica.com/space/2026/07/indias-first-privately-developed-rocket-reaches-orbit-on-dr...
475•sohkamyung•4d ago•139 comments

Designing an Ethernet Switch ASIC

https://essenceia.github.io/projects/ethernet_switch_asic/
85•random__duck•4d ago•23 comments

An old patent inspired the new "Y-zipper", a three-sided fastener

https://news.mit.edu/2026/three-sided-y-zipper-design-0504
112•crescit_eundo•2d ago•29 comments

Nvidia, Microsoft, Meta warn against overregulating open-weight models

https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-weight-ai-models.html
478•louiereederson•11h ago•222 comments

If coding has been solved, why does software keep getting worse?

https://ptrchm.com/posts/nothing-works-and-everyone-is-euphoric/
478•pchm•15h ago•383 comments

AIs don't do what you want. This is bad

https://rewardhacking.org
36•kking23•2h ago•16 comments

Kimi K3 exploited the latest Redis server

https://twitter.com/fried_rice/status/2080059356322918777
128•Alifatisk•1d ago•38 comments

Fil-C: Garbage In, Memory Safety Out [video]

https://www.youtube.com/watch?v=5F-2Y1LPRek
96•Bootvis•1d ago•96 comments

Half-Life 2 running natively on HaikuOS

https://discuss.haiku-os.org/t/haiku-nvidia-porting-nvidia-driver-for-turing-gpus/16520?page=18
258•m0do1•11h ago•50 comments

Marimo now runs in PyCharm

https://marimo.io/blog/pycharm
73•cantdutchthis•2d ago•16 comments

Firefox Containers Preview

https://blog.mozilla.org/en/firefox/firefox-containers-preview/
209•twapi•3d ago•76 comments

Gsxui – Shadcn-style components for Go

https://ui.gsxhq.dev/
54•jackielii•6h ago•8 comments

SpaceX Starship Flight 13 livestream [video]

https://www.spacex.com/launches/starship-flight-13
76•cryptoz•1h ago•75 comments

Don't Take the Black Pill [video]

https://www.youtube.com/watch?v=zLZwpH5lCD4
112•signa11•7h ago•77 comments

IRGC claims it destroyed Amazon's Bahrain data center

https://houseofsaud.com/irgc-claims-destroyed-amazon-bahrain-data-center/
236•thisislife2•14h ago•299 comments

The road to epsilon-zero: Nim always ends, even with infinite ordinals

https://blog.plover.com/math/ordinals/02-wellfoundedness.html
7•pavel_lishin•4d ago•0 comments

Flux 3 X Mimic: The Next Generation of Video-Action Models

https://bfl.ai/blog/flux-3-mimic
310•kensai•15h ago•48 comments

Future euro banknote design proposals

https://www.ecb.europa.eu/euro/banknotes/future_banknotes/html/all-design-proposals.en.html
120•robin_reala•15h ago•116 comments

Be skeptical of OpenAI's rogue hacker agent story

https://www.theguardian.com/technology/2026/jul/24/openai-rogue-hacker
398•rwmj•8h ago•221 comments

Sperm Whales blow bubbles to achieve restful, vertical sleep

https://news.st-andrews.ac.uk/archive/sperm-whales-blow-bubbles-to-achieve-restful-vertical-sleep/
12•hhs•1h ago•1 comments

The footprints of every building in NYC

https://www.beautifulpublicdata.com/the-footprints-of-every-building-in-nyc/
44•jonathanmkeegan•4d ago•5 comments

Unitree As2-W

https://www.unitree.com/As2-W/
89•MehrdadKhnzd•8h ago•40 comments

Government orders GitHub to remove Bluetooth-based chat app Bitchat: Jack Dorsey

https://www.thehindu.com/news/national/government-orders-github-to-remove-bluetooth-based-chat-ap...
363•rootkea•10h ago•269 comments

The case for MUDs in modern times (2018)

https://www.andrewzigler.com/feed/the-case-for-muds-in-modern-times
71•bw86•12h ago•64 comments

Programming language file extensions that match ISO 3166-1 alpha-2 country codes

https://www.bruh.ltd/blog/programming-language-file-extensions-that-match-an-iso-3166-1-alpha-2-c...
36•speckx•12h ago•21 comments