frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

LLMs could control their host machines by exploiting inference engines

https://boydkane.com/essays/llms-could-control-their-host-machines-by-exploiting-inference-engines
22•zdw•58m ago

Comments

alphazard•32m ago
This framing of security as something that belongs in the harness is completely wrong, and I hope no one is relying on a correct harness to keep their agents isolated.

VM, or even just a container will do. The agent should be able to run as root in its environment and do whatever it wants. If you can't give it that, you aren't sandboxing correctly.

Razengan•26m ago
Also, operating systems should let us set filesystem permissions per app/process/executable instead of just users.

Similar to how macOS/iOS Sandboxing works but at a more lower and granular level

Retr0id•15m ago
SELinux is basically this.
wild_egg•23m ago
Article isn't about agents. It's about the inference engine itself being exploited by a malicious LLM output before it is ever sent to your machine or harness.
empath75•16m ago
I think if you are convinced you are sandboxing an LLM properly, you almost certainly are not. I think it is essentially impossible to have a frontier LLM with enough access to be useful without also giving it enough access to do damage if it's compromised or just goes off the rails.
xg15•5m ago
> ...however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded? Such a machine is a high-value target: it has sufficient compute to run a frontier LLM, offers easy access to the LLM’s weights, and has privileged access to other computers in the datacentre compared with a generic computer on the internet.

> How do we defend against this? ... Run the GPUs and token parser on separate computers.

For models large enough to be relevant here, is there even "a" computer where the inference is performed? I'd imagine most of that stuff is ran on multi-GPU clusters with specialized architecture and not a generic vLLM instance. As such, I think there is a good chance the "API gateway" code that parses tokens into whatever JSON structure the public API offers is already running on a different machine than the actual inference.

(Even more so as you'd probably want to utilize batching: Several API calls will be put into the same inference batch, but the token parsing will have to be done separately for each call again)

The article is also very handwavy about why an LLM should do that - how it could learn the exploit, what would make it conclude that it can use the exploit on its own inference session and what would trigger it to actually use the exploit.

woadwarrior01•2m ago
FWIW, macOS has good sandboxing, but LMStudio, Ollama, Darkbloom etc aren't sandboxed. This is also the reason why none of these things aren't distributed via the Mac App Store, because the Mac App Store mandates sandboxing.

Show HN: WorkBase – local-first project manager where a project is a tree

https://github.com/vocso-com/WorkBase
1•vocso•1m ago•0 comments

Nine Indicted by Taiwan over Illegal Export of Nvidia B300 GPUs to China

https://www.tomshardware.com/tech-industry/artificial-intelligence/nine-indicted-by-taiwan-over-i...
1•geoffbp•1m ago•0 comments

Jonathan Blow – Preventing the Collapse of Civilization (2019) [video]

https://www.youtube.com/watch?v=ZSRHeXYDLko
1•andai•2m ago•0 comments

Show HN: Maya Calendar

https://mcal.qt.ax/
1•quantum5•3m ago•0 comments

Show HN: AI for people that still browse the web

https://thisisrobin.ai/
1•nikolas1814•4m ago•0 comments

The Great Vowel Shift

https://en.wikipedia.org/wiki/Great_Vowel_Shift
1•samcgraw•4m ago•0 comments

Why the Largest Megaproject Collapsed [video]

https://www.youtube.com/watch?v=GbB3vFfiaVU
1•simonebrunozzi•7m ago•0 comments

The Teaser Period: Why the AI Boom Is Built to Break

https://www.groundbrkr.com/p/the-teaser-period-why-the-ai-boom
1•root-parent•12m ago•0 comments

Qwen 3.6 is now much easier to run locally on your Mac, thanks to JetBrains

https://www.neowin.net/news/qwen-36-is-now-much-easier-to-run-locally-on-your-mac-thanks-to-jetbr...
4•bundie•13m ago•0 comments

Gray Goo

https://en.wikipedia.org/wiki/Gray_goo
1•mkl95•13m ago•0 comments

The Government Report That Made Me Stop Trusting Our Statistical Agencies

https://econjared.substack.com/p/the-government-report-that-made-me
2•root-parent•13m ago•0 comments

Show HN: Nauma new features (financial planning)

1•alx_sukhanov•18m ago•0 comments

Greg Abbott says data centers:'basically dug their own grave'

https://www.businessinsider.com/greg-abbott-texas-data-centers-backlash-2026-8
5•root-parent•18m ago•0 comments

DDD, Hexagonal, Onion, Clean, CQRS, How I put it all together (2017)

https://herbertograca.com/2017/11/16/explicit-architecture-01-ddd-hexagonal-onion-clean-cqrs-how-...
1•gregsadetsky•21m ago•0 comments

Nvidia Jetson Orin drone kills 3 in July Russia attack on Ukraine gas station

https://www.nytimes.com/2026/08/24/world/europe/ukraine-war-nvidia-ai-autonomous-drones.html
2•ncr100•22m ago•1 comments

Harvard Is Selling a $699 Course Taught by A.I. Clones of Its Faculty

https://www.nytimes.com/2026/08/22/business/dealbook/harvard-ai-faculty.html
3•bookofjoe•24m ago•2 comments

Apple Watch Alerted Runner to Serious Heart Problem

https://www.macrumors.com/2026/08/20/apple-watch-alerted-runner-to-heart-problem/
2•speckx•24m ago•0 comments

Porsche Inks $1.5B AI Deal with India's Tata Consultancy

https://www.bloomberg.com/news/articles/2026-08-24/porsche-inks-1-5-billion-ai-deal-with-india-s-...
2•johnbarron•27m ago•0 comments

"Model Effort" Has It Backwards

https://diverging.run/checkpoints/human-effort-not-model-effort/
1•shay_ker•28m ago•0 comments

Continuous Prompt Evaluation

https://kiro.dev/blog/continuous-prompt-evaluation/
2•ruptwelve•28m ago•0 comments

I can't prove it, but I think AI is causing me brain damage

3•nyxtom•28m ago•2 comments

The John Dvorakification of the blogosphere (2006)

https://web.archive.org/web/20060322062548/http://scobleizer.wordpress.com/2006/03/05/the-john-dv...
2•tolerance•28m ago•0 comments

A Claude Code skill that recovers export-blocked Kindle highlights

https://github.com/l3a0/claude-plugins
8•l3a0•28m ago•0 comments

How to spot a stupid person with Carlo Cipolla's "golden law of stupidity"

https://bigthink.com/mini-philosophy/golden-law-of-stupidity/
4•speckx•29m ago•0 comments

Using Models to Create Models of New York City

https://petersobot.com/blog/using-models-to-create-models-of-new-york-city/
1•psobot•29m ago•0 comments

OpenOx – A Protocol for Self-Evolving Agents

https://openox.ai/
3•ziyzhu•30m ago•0 comments

On Average: The choice in maximizing expected value

https://eehh-stanford.github.io/monkeys_uncle/posts/means_to_ends/
3•anigbrowl•32m ago•0 comments

Show HN: School chaos. Zero stress. Nuet organizes it all

https://www.nuet.ai/
2•wowinter15•33m ago•0 comments

Mistral and Humain Announce Colab to Advance AI in Saudi Arabia and Regionally

https://mistral.ai/news/mistral-x-humain/
3•leumon•33m ago•0 comments

Fuck Your Groupthink

1•shoman3003•33m ago•0 comments