frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Why Are Coding Agents So Dumb?

https://mtlynch.io/why-are-coding-agents-so-dumb/
39•mtlynch•7h ago

Comments

cbrake•3h ago
Enjoyed this article, lots I can relate to.

One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.

While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:

https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...

imimayj1337•2h ago
It's true! It feels like we've been talking about harness optimisation and 'cool features' available in the cli tools for months at this point, but ostensibly there has not really been any significant upgrades to these harnesses since at least Claude Code imo. It does feel like a contrived way to harvest more and more information and test each conversation/action tool, to the detriment of those of us actually using them!
mstank•48m ago
I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.

I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.

ilamont•41m ago
the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”

This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.

Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?

TeMPOraL•38m ago
OTOH, would you trust the vendor to pick the best model for you? Would you trust them not to prioritize their own load-balancing concerns first?

The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).

wrs•21m ago
Actually, the last time I asked Claude Code about itself, it located and read its own minified source and told me something that wasn’t even in the docs.
asdff•14m ago
So that's how model distillation is done. Just ask for the source code directly.
brandtcormorant•19m ago
They are as dumb as their instructions.

Have you tried telling models about your dream agent environment?

They can build it.

polyterative•15m ago
I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.

A lot can be improved, but this is already so much speed.

kgeist•10m ago
AI models can multitask/use parallel subagents just fine; the issue is with harnesses that don't make it a priority via the default system prompt, etc.

I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.

So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)

chrisjj•6m ago
[delayed]
Kuyawa•5m ago
Perhaps is not the agent that is dumb?

I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!

Sooo, which agent?

YouTuber Says Cops Visited Him After He Built a Flock-Style Camera to Track Cops

https://gizmodo.com/youtuber-says-cops-paid-him-a-visit-after-he-built-flock-style-camera-to-trac...
46•gumby•20m ago•1 comments

No Man Is an Island

https://borretti.me/article/no-man-is-an-island
130•zetalyrae•1h ago•47 comments

Cloudflare acquires Deno

https://deno.com/blog/cloudflare
934•ilreb•8h ago•493 comments

Triple-A Minesweeper

https://minesweeper.mikelacher.com/
358•robin_reala•5h ago•78 comments

Show HN: Carrier-Explode: iPhone, Pixel and Galaxy carrier settings decoded

https://carrierexplode.com/
127•simplyalec•3h ago•13 comments

Our $445M Series D

https://oxide.computer/blog/our-445m-series-d
521•ahlCVA•8h ago•220 comments

Typesafe AI raises $870M at $7.5B

https://typesafe.ai/blog/series-ai
180•tosh•4h ago•142 comments

Ideas aren't getting harder to find, anyone who tells you otherwise is a coward

https://www.experimental-history.com/p/ideas-arent-getting-harder-to-find
64•rafaelc•3h ago•21 comments

Sorry, I'm in a meeting

https://iminafleeting.com/
658•splintersio•12h ago•215 comments

Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

https://jessewaites.com/blog/post/i-pointed-ai-at-400-years-of-archives/
70•piratebroadcast•9h ago•33 comments

Show HN: Let your AI agents paint big arrows, boxes and text on your screen

https://github.com/franzenzenhofer/big-arrow-on-the-screen
346•franze•10h ago•147 comments

Nobel Peace Prize for 2026 to Navanethem Pillay

https://www.nobelprize.org/prizes/peace/2026/press-release/
403•Anon84•11h ago•205 comments

M7.6 Earthquake in Panama

https://earthquake.usgs.gov/earthquakes/eventpage/us6000u18k/executive
97•gslin•3h ago•33 comments

The reciprocal sum of the prime-prefix-free numbers converges [pdf]

https://jdb19937.github.io/prime-prefix-free/ppf.pdf
12•jdb1729•4d ago•12 comments

'Wallace and Gromit,' 90% Alone

https://animationobsessive.substack.com/p/wallace-and-gromit-90-alone
101•vinhnx•7h ago•14 comments

Why isn't the industry freaking out about DeepSeek 4.1 Flash?

https://www.dgt.is/blog/2026-10-07-deepseek-freek-out/
1059•jonotime•1d ago•933 comments

A statement on the Tor Project's relationship with Mullvad

https://blog.torproject.org/on-tor-relationship-with-mullvad/
78•runtimewire•5h ago•158 comments

Microsoft-Decision-1, our model for fast decision-making

https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
74•lisajaloza•2h ago•26 comments

Whistle: Speech to Text in 16.9 MB

https://cactuscompute.com/blog/whistle
910•gmays•1d ago•176 comments

Germany transforms former coal mines into Europe's largest lake landscape

https://www.euronews.com/2026/04/14/almost-like-lake-como-germany-transforms-former-coal-mines-in...
162•ohjeez•6h ago•89 comments

Training Text-to-Image Models Without a VAE

https://www.linum.ai/field-notes/pyramid-jit
41•schopra909•3d ago•14 comments

Yes, and

https://htmx.org/essays/yes-and/
709•Michelangelo11•1d ago•280 comments

You might want to try being less creative

https://blog.bawolf.com/p/you-might-want-to-try-being-less
41•bryantwolf•2h ago•23 comments

Keyboard differences between Windows and Macs

https://unsung.aresluna.org/deeper-dive-keyboard-differences-between-windows-and-macs/
308•sohkamyung•18h ago•251 comments

OpenAI fires three safety researchers for "mishandling research information"

https://techcrunch.com/2026/10/08/fired-openai-safety-researchers-dispute-misconduct-claims-warn-...
280•trakkstar•11h ago•182 comments

MXC - a sandboxed code execution system

https://github.com/microsoft/mxc
163•nreece•15h ago•78 comments

Programming Isn't Special

https://blog.glyph.im/2026/10/programming-isnt-special.html
146•ingve•13h ago•163 comments

Why Are Coding Agents So Dumb?

https://mtlynch.io/why-are-coding-agents-so-dumb/
40•mtlynch•7h ago•11 comments

Theranos.world

https://www.theranos.world/
538•kbyatnal•1d ago•199 comments

Once: Cache CLI commands

https://github.com/alex0ptr/once
83•baquero•11h ago•37 comments