frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Handbook.md shows that long policy documents do not reliably govern agents

https://arxiv.org/abs/2607.25398
34•spIrr•50m ago

Comments

leetrout•19m ago
HANDBOOK.md is a benchmark for long-context agentic instruction following, modeled on how enterprise employees follow company handbooks in their day-to-day work. Each task is a unique RL environment with internal tools and external MCP servers, spanning five enterprise domains: Finance, Medical Billing, Insurance, Logistics, and HR.

The prompts reflect the actual jobs enterprise workers perform every day. Each task drops an AI agent into a live company environment, requiring them to cross-reference an extensive, multi-section handbook against a cluttered inbox, a multi-channel Slack workspace, Jira queues, and a stack of files (spreadsheets, PDFs), and working out both what to do and what the handbook forbids.

https://github.com/surge-ai/handbook/tree/main

mcdeltat•19m ago
Yeah checks out with my anecdotal experience with Claude. It is pretty great at following instructions - for about 10 minutes, after which it seems to ignore things I told it before.

I have quite explicit and strong instructions (e.g. don't write massive comments, use existing functionality, etc.) in CLAUDE.md files which seem to get bypassed surprisingly quickly when doing real tasks. Yet if I tell it these things in a prompt during the task, it performs way better.

Result is I'm trying to resist adding more and more things to CLAUDE.md files which in some scenarios it does well but in other scenarios totally ignores and messes up.

cyanydeez•10m ago
I believe the correct static instructions are about getting it at the right starting point for whatever class of projects you're working on; not as a continued referencable or "HOWTO" of what it's doing. They're all just "grooming" the LLM for future instructions.

The coding harness is what's getting it to continually align to your current instructions.

This is very obvious with local models.

spIrr•6m ago
As a hobbyist, I find it difficult to figure out how to make Claude stick with some repeating things I want it to do after every major action, like re-evaluate the completeness of tests, update the documentation, etc. And CLAUDE.md/AGENTS.md definitely did NOT help there, sadly.
mcdeltat•4m ago
Ok so what is the correct way to tell it "I don't care what is happening, you must uphold these rules at all times"? If it's not any configuration of .md files?
firasd•17m ago
Opus 4.8 (max thinking) scored highest and Grok 4.3 lowest

It's hard to understand what's going on with Grok. It's like it has capabilities in a theoretical sense but maybe the training is so focused on being in x.com/grok.com with the web search tool enabled for "is this true?11" type queries that with any API type usage with document workflow instructions, tool use, code gen etc it completely falls over

DiabloD3•9m ago
This is a problem with long context models. To put it as simple and as bluntly as possible: just because they claim you can use 1M tokens in your context doesn't mean its true and you should do that.

Due to extreme quantization of models and the context's KV cache, and also just really shitty samplers provided to the user (hell, most are just getting rid of sampler knobs altogether), this problem will absolutely continue.

Want it to go away, almost like magic? Local inference. When its under your control, and no longer being forced to hold it wrong, all of the common LLM defects will go away.

simpaticoder•2m ago
>Want it to go away, almost like magic? Local inference.

Ah yes, magic that costs the same as a new car.

mordae•8m ago
Why would anyone think that models optimized for efficient context management, giving much more weight to a short sliding window, would attend to distant, heavily diluted tokens?

Plus the model's capacity to take more context into account and actually integrate it to the output is simply limited by the number of activated parameters. If you give it a playbook, you are forcing to choose it between attending to the playbook and the task at hand.

If you want to force it to work step-by-step, you need to present the steps one-by-one. Ideally with rules for the current step at hand and maybe relevant input again, depending on overall task size.

Why did you think models love to re-read files before editing them? It increases recall quality and thus edit precision and thus benchmarks.

spIrr•4m ago
> limited by the number of activated parameters

not sure I got it?

Separately, the frontier labs are kinda pushing us into that behaviour by releasing models with ever-larger context windows.

pelagicAustral•4m ago
I noticed this behaviour a few months back, I think I was using Sonnet 4.6 at the time... I have strict rules about comments in the codebase, this all for personal projects, and the reason I restrict comments is to keep the token count low.

At some point between the model i was using and the previous version of it, Claude started inserting massive comments with references to tickets and other tasks. All this while having specific directives on the CLAUDE.md

Since then I resorted to developing my crapware as if I was the floor manager of a vehicle assembly line, and I have a few highly-specialized sub-agents running errands around the main session, but only ever taking care of a single concern. The main session builds with the knowledge contained in things like CLAUDE.md but the sub agents make sure things like the no/low-comments directives are either enforced, or factored into the final product.

What three agent security architectures leave unresolved

https://singularityforge.space/2026/07/29/permitted-refusal-when-authority-knowledge-and-executio...
1•Voice_of_Void•47s ago•0 comments

Show HN: RC Setlist – Setlist manager and lyric prompter for Ableton Live 12

https://github.com/ntworm/rc-setlist
1•gabrielworm•2m ago•0 comments

From keyword to published post in one focused workflow

https://presspilot.eu/
1•georgejr•3m ago•0 comments

Seventy-five years of game AI, and every system plays a handful of games

https://kallin.github.io/blog/game-ai-one-machine-any-game/
1•kal9000•4m ago•0 comments

The new leverage of software engineering

https://edparry.com/the-new-leverage-of-software-engineering
1•edparry•5m ago•0 comments

Show HN: Ctx.traits · Agent Traits and Workflows typed, versioned and in-sync

https://ctx.company/blog/introducing-ctx-traits/
1•rpunkfu•6m ago•0 comments

The Canvas Breach Exposed the Risk of One-Platform Education

https://teachers-blog.com/the-canvas-breach-exposed-the-risk-of-one-platform-education/
1•behindai•6m ago•0 comments

Zero-Risk Bias

https://en.wikipedia.org/wiki/Zero-risk_bias
1•fsflover•7m ago•0 comments

Reaction wheel failures leave Swift rescue mission spinning in orbit

https://arstechnica.com/space/2026/07/reaction-wheel-failures-leave-swift-rescue-mission-spinning...
2•world2vec•8m ago•0 comments

AI-enabled jobs in the Philippines: an in-depth report

https://www.ccm.pm/posts/ph-ai-jobs-demand-2026-07
1•char8•9m ago•0 comments

Show HN: Nix Flakes Dotfiles

https://github.com/nooneknowspeter/dotfiles
1•nooneknowspeter•10m ago•0 comments

I Crunched the Numbers Behind FIFA's Plan to Sell Shares in the World Cup

https://www.sportsandcrime.com/p/i-crunched-the-numbers-behind-fifas
1•FinnLobsien•10m ago•0 comments

OpenAI's rogue models roamed the internet for 4 days and staged a second attack

https://www.politico.com/news/2026/07/28/openai-rogue-models-hugging-face-breach-01014572
1•cc62cf4a4f20•10m ago•0 comments

Can We Lower Construction Costs with Cheaper Labor or Materials?

https://www.construction-physics.com/p/can-we-lower-construction-costs-with
2•surprisetalk•11m ago•0 comments

Does Your Home or Classroom Need an Air Purifier? Make a Corsi-Rosenthal Box

https://www.npr.org/sections/back-to-school-live-updates/2021/08/26/1031018250/does-your-kids-cla...
2•ck2•12m ago•1 comments

Show HN: Multi-agent LLM editor with local inference via WebSockets

https://x-agent.sascha10k.biz
1•sascha10000•12m ago•0 comments

The Y2K Comeback: Why Retro UI Is Trending Again

https://palettevault.github.io/blog/y2k-renaissance/
1•eustoria•14m ago•0 comments

Science Corporation's vision-restoring chip wins EU approval

https://techcrunch.com/2026/07/22/science-corporations-vision-restoring-chip-wins-eu-approval/
2•gmays•14m ago•0 comments

Validation Is the New Bottleneck (Also, the Old One)

https://pawelbrodzinski.substack.com/p/validation-is-the-new-bottleneck
2•flail•14m ago•0 comments

Emoji QR Code Generator

https://planqr.com/emoji-qr-code
1•eustoria•15m ago•0 comments

Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You

https://quesma.com/blog/baba-kimi-k3-opus-5/
1•stared•17m ago•0 comments

Cavdo AI – Fast, minimal AI assistant with modern UI cleek and for professional

1•Cavdo-Ai•17m ago•0 comments

Life and Death on a Tower

https://www.datacenterdynamics.com/en/analysis/life-and-death-on-a-tower/
1•herbertl•17m ago•0 comments

Queen's Gambit

https://en.wikipedia.org/wiki/Queen%27s_Gambit
1•tosh•17m ago•0 comments

Show HN: ShellTeam – web app to steer coding agents on your VPS

https://github.com/sebderhy/shellteam
1•sebderhy•17m ago•1 comments

French DJ Kavinsky known for Drive soundtrack Nightcall found dead in Paris home

https://www.thejournal.ie/dj-kavinsky-dies-7116639-Jul2026/
1•austinallegro•18m ago•0 comments

All living things emit a faint glow. Could this light be useful?

https://www.nature.com/articles/d41586-026-02311-z
1•bookofjoe•18m ago•0 comments

Money, Bitcoin, and AI

https://aisocratic.org/money-bitcoin-ai
1•feulf•18m ago•0 comments

Show HN: Kiyeovo Now Supports Windows

https://github.com/Realman78/Kiyeovo/releases/tag/kiyeovo-1.1.0
3•Realman78•18m ago•0 comments

Cavdo AI – Fast, minimal AI assistant with no signup or login

1•Cavdo-Ai•19m ago•0 comments