frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Every Model Cheats

https://dreadnode.io/research/every-model-cheats-prompt-level-mitigation-of-cheating-on-offensive-cyber-tasks/
8•vga805•1h ago

Comments

throwaway13337•16m ago
The problem is model confusion. You ask models to get around security but also not to get around your security.

Models get confused by who said what - especially cluade models. They get confused by negation (don't do something versus do something). Compartmentalization is hard.

You can either solve compartmentalization completely, or just not tell the model to do things that must be compartmentalized at high stakes.

xscott•14m ago
I'm not claiming to have any expertise in this area, but I've got a list of things I try to apply when working with LLMs. Possibly relevant here is, "don't tell the model what NOT to do, show it what TO do". I think guard rails should be implemented outside the model with an isolated system. The models seem to like patterns to follow.

Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.

super256•12m ago
>Anthropic’s Claude Opus 4.6 system card described Cybench as “saturated,” reporting near-100% pass rates without a cheating audit. If these estimates were representative, cheating would be a marginal artifact.

One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).

Malicious Rust Crate Arrayref Runs a Build-Time Payload

https://safedep.io/arrayref-proc-macro1-rust-build-time-malware/
147•abhisek•1h ago•97 comments

AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint

https://blog.laserphile.com/2026/08/aliexpress-webpage-keeping-multipoint.html
385•emctech•4h ago•124 comments

Harvest hikes bills by 1500% after purchased by Bending Spoons

https://www.bbc.com/news/articles/clyq011414eo
46•jonathanlydall•43m ago•25 comments

Show HN: I trained a 125M model to autocomplete piano on-device

https://simedw.com/2026/08/20/midi-autocomplete/
197•simedw•2h ago•45 comments

Stwipe Acquires OpenWouter

https://stwipe.com/
95•eatonphil•56m ago•13 comments

DiffusionGemma Technical Report

https://arxiv.org/abs/2608.00146
41•gmays•1h ago•6 comments

Hacking with Claude on a $27 Smart Watch

https://www.mikekasberg.com/blog/2026/08/19/hacking-with-claude-on-a-27-smart-watch.html
21•speckx•53m ago•9 comments

Why the Ocean Cleanup Hasn't Solved the Plastic Pollution Crisis

https://therevelator.org/why-ocean-cleanup-has-not-solved-plastic-pollution/
23•sohkamyung•2h ago•12 comments

Windows brings out the Rorschach test in everyone (2003)

https://devblogs.microsoft.com/oldnewthing/20030825-00/?p=42803
293•luu•8h ago•106 comments

Slack Code

https://www.salesforce.com/introducing-slack-code/?bc=HL
5•trollied•39m ago•0 comments

Mojo is now open source

https://www.modular.com/blog/mojo-open-source
165•visheshdembla•1d ago•48 comments

Proof of Human (YC S23) Is Hiring a Member of Technical Staff

https://www.ycombinator.com/companies/proof-of-human/jobs/ZTZHEbb-member-of-technical-staff
1•timshell•3h ago

An American Mosaic (interactive map of ancestry census data)

https://www.nytimes.com/interactive/2026/07/01/us/america-ancestry-census-data-map.html
18•cckolon•3d ago•3 comments

An elliptic curve of rank ≥ 30

https://elliptic-rank.icarm.cloud/curve/273
10•robinhouston•46m ago•2 comments

Xorg-Server 26.0.99.901

https://lists.x.org/archives/xorg-announce/2026-August/003741.html
15•st_goliath•2h ago•0 comments

Every Model Cheats

https://dreadnode.io/research/every-model-cheats-prompt-level-mitigation-of-cheating-on-offensive...
9•vga805•1h ago•3 comments

Seeing beyond BMI: Estimating cardiometabolic risk with smartphone imagery

https://research.google/blog/seeing-beyond-bmi-estimating-cardiometabolic-risk-with-smartphone-im...
39•leanderjanssen•3h ago•17 comments

Risk Engineering

https://risk-engineering.org/
46•throwaw12•4h ago•7 comments

CI jobs artifacts should not be difficult

https://deadsimpleci.sparrowhub.io/doc/job-artifacts
10•melezhik•1h ago•3 comments

Router by Ramp

https://router.com
87•zackfield•19h ago•52 comments

Turns are Better than Radians (2022)

https://www.computerenhance.com/p/turns-are-better-than-radians
295•mayoff•13h ago•156 comments

Grok.bot – the epic domain lottery ticket

https://grok.bot/
6•bogdiyan•27m ago•1 comments

Bun 1.4

https://bun.com/blog/bun-v1.4
49•meetpateltech•51m ago•16 comments

AI didn't erase the junior engineer's value, it increased it it

https://franciscotrindade.me/blog/the-kids-are-really-alright/
51•franciscomt•3h ago•80 comments

A faster way to calculate the day of the week

https://www.benjoffe.com/fast-day-of-week
223•gavide•3d ago•61 comments

Bufo pulls the andon cord

https://hatchet.run/blog/andon-cord
7•abelanger•2d ago•2 comments

Don't paste the AI, please

https://dontpastetheai.com/
901•pjerem•6h ago•470 comments

Zellij 0.45.0: nested sessions, Kitty graphics, a fresh UI

https://zellij.dev/news/nested-sessions-kitty-graphics-new-ui/
12•peterhajas•2h ago•2 comments

Canonical Backs New Project to Translate Large C Codebases into Safe Rust

https://linuxiac.com/canonical-backs-new-project-to-translate-large-c-codebases-into-safe-rust/
28•datakan•1h ago•20 comments

Sol loves to cheat

https://jumploops.com/blog/sol-loves-to-cheat/
217•jumploops•1d ago•177 comments