frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Memory Ordering in CPUs

https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-cpus/
1•Tomte•17m ago•0 comments

Placemark is now open source

https://macwright.com/2026/08/25/placemark-fully-open-source
1•Tomte•17m ago•0 comments

Run Open Models on Claude Desktop

https://ollama.com/blog/claude-desktop
1•ttunguz•36m ago•0 comments

Pilot program offering free bus rides to Detroit students to become permanent

https://www.detroitnews.com/story/news/local/detroit-city/2026/08/14/sheffield-makes-free-ddot-tr...
2•thunderbong•36m ago•0 comments

Ask HN: How do you explain system design to engineers who weren't in the room?

1•pandurang90•39m ago•1 comments

Kraftapp AI – Describe it. We build it. Customers find it

https://kraftapp.ai
1•ssurse•42m ago•0 comments

Show HN: Hands-on Docker security labs, from CIS checks to AI context poisoning

https://github.com/opscart/docker-security-practical-guide
1•opscart•45m ago•0 comments

I ran my code intelligence engine on 1k GitHub repositories

https://github.com/thecolourfoundation/rune�
1•malixp•48m ago•0 comments

I'm an Amazon SVP who hadn't coded in 25 years

https://www.aboutamazon.com/news/workplace/amazon-svp-ai-tools-employees-careers
2•dhruv3006•49m ago•0 comments

Show HN: MageCDN – Image CDN with unlimited bandwidth

https://magecdn.com/
1•shubhamjain•50m ago•0 comments

Drop JD and get customized interview course

https://www.interviewbasecamp.com/
1•balkrishnajha•54m ago•0 comments

Semantic caching has a structural gap that no threshold fixes

https://github.com/KushagraKanaujia/throttle/blob/main/scripts/negation_check.py
1•JohnScheuer•1h ago•0 comments

Show HN: A proxy that makes Forgejo speak the GitHub API

https://github.com/ThatXliner/anvil
1•thatxliner•1h ago•0 comments

The AI founders who walked away from Bezos-backed Prometheus to model universe

https://www.reuters.com/business/ai-founders-who-walked-away-bezos-backed-prometheus-model-univer...
2•petethomas•1h ago•0 comments

Extension 'ms-VSCode-remote.remote-SSH' CANNOT use API proposal

https://github.com/microsoft/vscode/issues/329735
2•themostunique•1h ago•0 comments

Students prefer Gemini over ChatGPT and Claude for AI essays in blind tests

https://studyarena.com/blog/chatgpt-vs-claude-vs-gemini-college-essays-2026
2•pasharayan•1h ago•1 comments

BetterReads: Read/Import/Track/Understand

https://www.betterreads.online
1•zlu•1h ago•1 comments

Data Centers Are Driving an Alarming Gas Power Expansion in the US

https://www.wired.com/story/us-data-centers-drive-gas-power-expansion/
3•newsomix9xl•1h ago•1 comments

Neuromancer Official Teaser

https://www.youtube.com/watch?v=g79GPZSQHBk
3•wslh•1h ago•0 comments

Ask HN: What is one simple thing LLMs are insanely bad at?

22•davidest•1h ago•38 comments

Plans for nuclear-powered merchant ships must confront risks

https://www.nature.com/articles/d41586-026-02388-6
1•mmooss•1h ago•0 comments

What platform has the best AI CMO?

1•probablygrillin•1h ago•0 comments

ILSpy in the Browser

https://osenkov.com/ilspy/
1•l33t_d0nut•1h ago•0 comments

Visual Analysis of Binary Files

https://binvis.io/#/
2•vismit2000•1h ago•0 comments

Show HN: Implementation of Kimi K3 in PyTorch

https://www.youtube.com/watch?v=U6sobPCsdaY
1•prasoon21•1h ago•0 comments

Darkbloom (AI inference on idle Macs) – security audit with PRs submitted

https://gist.github.com/mudiam/1ffc898333ac3d5bdc5d7fac96d33360
2•mudiam•1h ago•0 comments

RL Environments are all you need

https://twitter.com/madiator/status/2084657077637746957
2•gmays•1h ago•0 comments

Libracy – An ad-free, minimalist book tracker without social feeds

https://play.google.com/store/apps/details?id=com.libracy.app&hl=en_US
2•cehnzzdev•1h ago•0 comments

AWS Activate Credits

https://aws.amazon.com/
5•m4sk1994•1h ago•0 comments

Scottish photographer shot portraits of Alabama gingers to find American unity

https://www.al.com/news/2026/08/scottish-photographer-shot-stunning-portraits-of-alabama-gingers-...
2•thunderbong•1h ago•1 comments
Open in hackernews

A simple heuristic for agents: human-led vs. human-in-the-loop vs. agent-led

1•fletchervmiles•1y ago
tl;dr - the more agency your agent has, the simpler your use case needs to be

Most if not all successful production use cases today are either human-led or human-in-the-loop. Agent-led is possible but requires simplistic use cases.

---

Human-led:

An obvious example is ChatGPT. One input, one output. The model might suggest a follow-up or use a tool but ultimately, you're the master in command.

---

Human-in-the-loop:

The best example of this is Cursor (and other coding tools). Coding tools can do 99% of the coding for you, use dozens of tools, and are incredibly capable. But ultimately the human still gives the requirements, hits "accept" or "reject' AND gives feedback on each interaction turn.

The last point is important as it's a live recalibration.

This can sometimes not be enough though. An example of this is the rollout of Sonnect 3.7 in Cursor. The feedback loop vs model agency mix was off. Too much agency, not sufficient recalibration from the human. So users switched!

---

Agent-led:

This is where the agent leads the task, end-to-end. The user is just a participant. This is difficult because there's less recalibration so your probability of something going wrong increases on each turn… It's cumulative.

P(all good) = pⁿ

p = agent works correctly n = number of turns / interactions

Ok… I'm going to use my product as an example, not to promote, I'm just very familiar with how it works.

It's a chat agent that runs short customer interviews. My customers can configure it based on what they want to learn (i.e. why a customer churned) and send it to their customers.

It's agent-led because

→ as soon as the respondent opens the link, they're guided from there → at each turn the agent (not the human) is deciding what to do next

That means deciding the right thing to do over 10 to 30 conversation turns (depending on config). I.e. correctly decide:

→ whether to expand the conversation vs dive deeper → reflect on current progress + context → traverse a bunch of objectives and ask questions that draw out insight (per current objective)

Let's apply the above formula. Example:

Let's say:

→ n = 20 (i.e. number of conversation turns) → p = .99 (i.e. how often the agent does the right thing - 99% of the time)

That equals P(all good) = 0.99²⁰ ≈ 0.82

So if I ran 100 such 20‑turn conversations, I'd expect roughly 82 to complete as per instructions and about 18 to stumble at least once.

Let's change p to 95%...

→ n = 20 → p = .95

P(all good) = 0.95²⁰ ≈ 0.358

I.e. if I ran 100 such 20‑turn conversations, I’d expect roughly 36 to finish without a hitch and about 64 to go off‑track at least once.

My p score is high. I had to strip out a bunch of tools and simplify but I got there. And for my use case, a failure is just a slightly irrelevant response so it's manageable.

---

Conclusion:

Getting an agent to do the correct thing 99% is not trivial.

You basically can't have a super complicated workflow. Yes, you can mitigate this by introducing other agents to check the work but this then introduces latency.

There's always a tradeoff!

Know which category you're building in and if you're going for agent-led, narrow your use-case as much as possible.