frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: DFAH-Bench – same agent decision, different tool paths

https://github.com/ibm-client-engineering/output-drift-financial-llms
1•raffisk•54m ago
Hey - released this PyPi library after getting annoyed most agent evals only focus on outcomes, and the tool calling path is just as important in recognizing the outcome decision. So defined decision agreement rates, trajectory agreement rates etc etc

Most tell us whether the final answer was right, but not whether the agent actually did the same thing twice.

In prospective replays of financial services tasks, agents reached the same decision about 94–95% of the time. Exact tool-name paths matched only 67–69%. Once the tool arguments were included, agreement fell to 45–52%.

I don’t think every different path is automatically a failure. Sometimes it’s harmless flexibility. But if an agent can touch files, APIs, email, or financial controls, I’d at least like to know that it changed its path before calling the system stable.

The package repeatedly runs your agent, stores the traces, and compares the decision, tool order, arguments, and result identities.

Quick local test - no key or paid model needed:

dfah check-agent --agent dfah.demo → dfah check-agent --agent dfah.demo:toy_agent --episode-timeout-s 5 Paper: https://arxiv.org/abs/2607.20491 PyPI: https://pypi.org/project/dfah-bench/

It’s still early. I’m especially curious about two things:

How would you separate harmless path diversity from instability that should block or flag a run?

And which agent framework should I build the next adapter for? Or is there other ways folks are thinking about tracking tool paths and decision agreements

August Is Here: Time Has Come to Label AI Deepfakes

https://read.misalignedmag.com/august-is-here-time-has-come-to-label-ai-deepfakes-2181ad0435ac
1•lcubw•1m ago•1 comments

TikTrack

https://einzzcookie.org/
1•kekseesser•2m ago•0 comments

Show HN: Replaybook, an Infrastructure Agent Evaluation Framework

https://github.com/ducks/replaybook
1•ducks_•5m ago•0 comments

Path of Exile 2 devs convinced Nvidia to fix a bug by sending it PC

https://www.pcgamer.com/games/rpg/after-a-year-and-a-half-of-rigorous-testing-path-of-exile-2-dev...
1•EvgeniyZh•6m ago•0 comments

Ask HN: What free browser tools do you use daily?

1•tooly_work•7m ago•0 comments

List of United States presidential assassination attempts and plots

https://en.wikipedia.org/wiki/List_of_United_States_presidential_assassination_attempts_and_plots
1•loughnane•7m ago•0 comments

Mesh2Motion Update 12: Horses and Offline Application

https://mesh2motion.org/news
1•wertyk•8m ago•0 comments

Tokenmaxxing Is Dead. Now Comes the Belt Tightening

https://www.bloomberg.com/news/articles/2026-07-31/corporate-america-cracks-down-on-ai-spending-a...
3•mancerayder•14m ago•0 comments

How the words we teach English language learners changed

https://pudding.cool/2026/07/essential-words/
2•c-oreills•16m ago•0 comments

A discarded SpaceX rocket is on a high-speed collision course with the moon

https://apnews.com/article/spacex-rocket-moon-crash-512c4dd708b4cda1160d30b764f9fdb5
1•marc__1•17m ago•0 comments

Why Sheep Need Pigs in Sheepdog's Clothing

https://www.overcomingbias.com/p/why-sheep-need-pigs-in-sheepdogs
1•jger15•18m ago•0 comments

Show HN: I spent 2.5 years developing my 3D portfolio website, all under 5 MB

https://shees.dev
3•sheesdev•19m ago•2 comments

The Case for Bad Coffee

https://www.seriouseats.com/the-case-for-bad-coffee
1•colinprince•20m ago•1 comments

Decadence Without Pleasure: Why Nothing Feels Joyful Anymore

https://www.theguardian.com/us-news/ng-interactive/2026/jul/26/age-of-decadence-pleasure-ai
2•yarapavan•20m ago•0 comments

Ask HN: Will Stripe Buy PayPal?

1•cyanregiment•22m ago•0 comments

Hungary to shut down sole nuclear plant as Danube falls to record low

https://www.reuters.com/business/energy/hungarys-paks-nuclear-plant-could-be-powered-down-this-we...
2•geox•24m ago•0 comments

Counter-Based random number generator

https://en.wikipedia.org/wiki/Counter-based_random_number_generator
1•binyu•29m ago•0 comments

Show HN: Interactive Heatsink Area Calculator by Geometric Profile

https://specled.com/en/w/heatsink-profile-area-calculator-by-geometric-profile/
1•Specled•29m ago•0 comments

Demobscene

https://demobscene.com/
1•fatliverfreddy•29m ago•0 comments

Write Amplification Explained

https://everything.explained.today/Write_amplification/
1•ankitg12•29m ago•0 comments

Mosh in a Lift (2012)

https://mosh.org/elevator.txt
1•gavide•30m ago•0 comments

Having fun with oh my pi, DeepSeek-V4-Flash, GPT-5.6 Luna and Antigravity CLI

https://flashblaze.xyz/posts/having-fun-with-omp-deepseek-luna-and-agy/
2•flashblaze•33m ago•0 comments

Yes you can measure engineering

https://www.rubick.com/valuesum-metric/
5•meetpateltech•34m ago•1 comments

Security Flaw Placed 30 Years of DNA Evidence at Risk of Hacking

https://www.wsj.com/tech/cybersecurity/security-flaw-placed-30-years-of-dna-evidence-at-risk-of-h...
1•impish9208•38m ago•1 comments

The Impending, Inescapable Deluge of A.I

https://www.nytimes.com/interactive/2026/07/29/technology/ai-chips-data-center-boom.html
3•bookofjoe•38m ago•1 comments

Show HN: A fixed harness for comparing LLM agent memory systems

https://github.com/AML-memory/agent-memory-leaderboard
1•IreneAI•39m ago•0 comments

VC-backed startups commit more fraud, and researchers think they know why

https://techcrunch.com/2026/07/31/vc-backed-startups-commit-more-fraud-and-researchers-think-they...
2•richardchilders•39m ago•0 comments

Mozilla's Inaugural 'State of Open Source AI' Report Is Here

https://blog.mozilla.org/en/mozilla/mozilla-state-of-open-source-ai-report/
3•nlpnerd•41m ago•0 comments

Stack Trace for Distributed Systems

https://github.com/leandromoreira/distributed-stack-trace/tree/main
1•dreampeppers99•47m ago•1 comments

Claude on Political Compass

https://utopiagov.com/blog/political-compass
2•makosst•50m ago•0 comments