frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Don't let AI kill the author

https://www.economist.com/leaders/2026/09/24/dont-let-ai-kill-the-author
1•edward•52s ago•0 comments

Cost of Delay

https://www.votito.com/methods/cost-of-delay/
2•gojkoa•1m ago•0 comments

RevExOS – JEV

https://revexos.com/jev-playground
1•revexos•1m ago•0 comments

JEV-as-a-Judge

https://academy.dair.ai/papers/jev-as-a-judge-accept-when-confident-escalate-when-unsure-2609.26550
1•omarsar•5m ago•0 comments

DOJ: Uncle Sam bought forensics software from same Russian operation supplying

https://www.theregister.com/software/2026/09/24/doj-uncle-sam-bought-forensics-software-from-same...
2•sbulaev•8m ago•0 comments

Internationalization (Internationalization) and Localization (L10n)

https://www.kashyapsuhas.com/blog/i18n-l10n
1•kashyapS07•8m ago•0 comments

Why the FBI investigated archbishop sheen during WWII

https://www.osvnews.com/why-the-fbi-investigated-archbishop-sheen-during-wwii/
1•MarcoDewey•8m ago•0 comments

Ask HN: Does spec-driven development still pay off with frontier coding models?

1•sarangk90•9m ago•0 comments

AI is accelerating and amplifying threats: immediate action is necessary

https://english.aivd.nl/latest/news/2026/09/24/ai-is-accelerating-and-amplifying-threats-immediat...
1•HelloUsername•10m ago•0 comments

Marvel Has Been Wearing a Horror Movie Under That Spandex Since 1962

https://medium.com/theentertainmentbreakdown/marvel-has-been-wearing-a-horror-movie-under-that-sp...
1•raynchad•12m ago•0 comments

Show HN: Rig – Open-source cloud desktops for AI agents

https://github.com/ShadowWalker2014/rig
2•fengjiabo2400•13m ago•0 comments

Russian Cosmism

https://en.wikipedia.org/wiki/Russian_cosmism
3•PaulHoule•15m ago•0 comments

'Palantir is a great company', EU defence chief tells Euronews

https://www.euronews.com/2026/09/24/palantir-is-a-great-company-eu-defence-chief-tells-euronews
3•g-b-r•15m ago•1 comments

JEV-Star: Low-Cost StarCraft II Control with Language-Model Planning

https://github.com/sc2musa/Jev_Star
1•tndl•15m ago•0 comments

Run-assert-eval: Find the risk, fix it, prove it

https://commandline.microsoft.com/run-assert-eval-responsible-ai-agent-risk-discovery-at-runtime/
1•doomroot13•16m ago•0 comments

Digital Cash Payoff (2001)

https://www.technologyreview.com/2001/12/01/235339/digital-cash-payoff/
1•ronfriedhaber•18m ago•0 comments

Show HN: DrawCMS – An open-source animated diagramming tool for AI Agents

https://github.com/drawcms/drawcms
1•dimas3399•18m ago•0 comments

Logs you, your AI and cron jobs write to that you can prove weren't altered

https://freshjots.com/
2•arsphy•21m ago•1 comments

Show HN: bananabread AI – Find what customers love and want from your products

https://bb-product-api-docs.web.app/
1•aditya314159•22m ago•0 comments

GoDaddy receives takeover offer from maker of Norton antivirus software

https://www.ft.com/content/6cc7149f-a9f1-47f6-9faf-b98a28f9aedc
1•JumpCrisscross•22m ago•0 comments

If you don't have the factories, you lose the expertise

https://lemire.me/blog/2026/09/24/if-you-dont-have-the-factories-you-lose-the-expertise/
3•ibobev•23m ago•0 comments

Show HN: Headwire – WireGuard with NAT traversal via Tailscale's magicsock

https://github.com/brofranks/headwire
1•brof•23m ago•0 comments

Namernut – domain name generator that scores how crowded a name is

https://namernut.com
1•coreinch•25m ago•0 comments

Show HN: MEF LLM Studio

https://www.mef-llm-studio.com/de/
1•walti1972•25m ago•0 comments

Fragments: September 24

https://martinfowler.com/fragments/2026-09-24.html
1•ibobev•26m ago•0 comments

Double-entry bookkeeping and paper and tokens

https://honza.pokorny.ca/2026/09/double-entry-bookkeeping-and-paper-and-tokens/
1•ibobev•27m ago•0 comments

Pummelvision Back

https://www.pummelvision.ai/
1•SerKnight•28m ago•0 comments

Legal and General to Cut 10% of Jobs by Middle of 2027

https://www.wsj.com/business/legal-general-to-cut-10-of-jobs-by-middle-of-2027-0109ef4c
1•frank_clover•28m ago•0 comments

Beware Overreliance on Metaphor

https://evnm.substack.com/p/beware-overreliance-on-metaphor
2•Mongoose•30m ago•0 comments

AI Workers' Inquiry 2026

https://techworkersinquiry.org/ai/
1•utiiiD•30m ago•0 comments
Open in hackernews

For Computer Use, the harness matters as much as the model

https://www.stagehand.dev/evals
20•MiguelG719•1h ago

Comments

MiguelG719•1h ago
Hi HN,

Over the last 2 years, we observed computer use models improving at a rapid pace and saturating benchmarks. This new benchmark replaces Online-Mind2Web with our own Browserbase Benchmark v2 that better represents the complex tasks that browser agents face in the real world. It runs against 23 models (frontier and open-weight) and 9 harnesses (Claude Code to LangChain Deep Agents) on accuracy, speed, and cost.

This new benchmark confirmed our belief that the choice of an harness is becoming as important as the choice of a model. For example: claude-opus-5 runs 74% at $1.50/task on LangChain deep agents but 71% at ~$10/task on fx.

The eval harness is a CLI you can run yourself (pick harness + tools/mcps + model, pass high-level tasks, grades with LLM verifiers, has trials/concurrency/OTEL tracing): https://github.com/browserbase/stagehand/tree/main/packages/...

Happy to get into methodology, and if you want your model or harness added, just let me know.

alyssamaru•45m ago
Yes, would love to hear the methodology!
peeet•54m ago
I love the website, it would be 10/10 if I could go to chrome://dino
MiguelG719•50m ago
Try https://stagehand.dev/dino ;)
mhykim•41m ago
Why use Stagehand when agents can write CDP / Playwright on the fly for browser use
MiguelG719•30m ago
Token efficiency, performance, observability, and most importantly: permissions/security policies
starlightttt•40m ago
Love the design
smpandya•33m ago
Are there results comparing agents running different tools (agent-browser, playwright MCP, browse CLI), or is this mostly Stagehand focused?
MiguelG719•26m ago
You can swap the driver/tool yourself with the evals CLI! Just use `—tool-surface <one-of-the-supported-tools>`. Or define your own using the interface
pranaygup12•28m ago
What model family do you find is the best for browser use overall? or does it change pretty regularly
MiguelG719•22m ago
It changes so frequently, and the world wild web is vast so it’s dependent on your use case. In general the leaderboard reflects what we see working across a broad range of domains, but the best way to tell is to define your tasks and run the evals yourself; that’s what this is for
dericdinudaniel•23m ago
Cost difference across harnesses is interesting to me. Would love to see more info about more optimizations in the harnesses to trim down costs.
devk03•21m ago
Hyper personalized harnesses are the edge that the labs cannot beat startups on. There will be a whole entire era of new harnesses coming out soon.
dericdinudaniel•19m ago
a whole new era of custom harnesses with unnecessary stuff trimmed out sounds like the future
MiguelG719•17m ago
It still feels like the harness and the model need to co-evolve together