frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: LLM is useless without explicit prompt

4•revskill•1y ago
After months playing with LLM models, here's my observation:

- LLM is basically useless without explicit intent in your prompt.

- LLM failed to correct itself. If it generated bullshits, it's an inifinite loop of generating more bullshits.

The question is, without explicit prompt, could LLM leverage all the best practices to provide maintainable code without me instruct it at least ?

Comments

ben_w•1y ago
Your expectations are way too high.

> - LLM is basically useless without explicit intent in your prompt.

You can say the same about every dev I've worked with, including myself. This is literally why humans have meetings rather than all of us diving in to whatever we're self-motivated to do.

What does differ is time-scales of the feedback loop with the management:

Humans meetings are daily to weekly.

According to recent research*, the state-of-the-art models are only 50% accurate at tasks that would take a human expert an hour, or 80% accurate at tasks that would take a human expert 10 minutes.

Even if the currently observed trend of increasing time horizons holds, we're 21 months from having an AI where every other daily standup is "ugh, no, you got it wrong", and just over 5 years from them being able to manage a 2-week sprint with an 80% chance of success (in the absence of continuous feedback).

Even that isn't really enough for them to properly "leverage all the best practices to provide maintainable code", as archiecture and maintainability are longer horizon tasks than 2-week sprints.

* https://youtu.be/evSFeqTZdqs?si=QIzIjB6hotJ0FgHm

revskill•1y ago
It's not as high as you think.

LLM failed at the most basic things related to maintainable code. Its code is basicaly a hackery mess without any structure at all.

It's my expectation is that, at least, some kind of maintainable code is generated from what's it's learnt.

ben_w•1y ago
Given your expectation:

> It's my expectation is that, at least, some kind of maintainable code is generated from what's it's learnt.

And your observation:

> LLM failed at the most basic things related to maintainable code. Its code is basicaly a hackery mess without any structure at all.

QED, *your expectations* are way too high.

They can't do that yet.

Show HN: AI Security Leaderboard – comparing cyber and CBRN safeguards

https://leaderboard.far.ai/
1•AdamGleave•1m ago•0 comments

Engineers have stopped reviewing PRs

https://aq.dev/guides/how-to-review-an-ai-coding-session/
1•knighthacker•3m ago•1 comments

Husk – a desktop workspace for terminal AI agents

https://github.com/DorShaer/Husk
1•DorShaer•6m ago•0 comments

Show HN: Play ROMs inline in your terminal

https://github.com/jhickner/rom
1•jhickner•7m ago•0 comments

NanoClaw and Echo launch agent runtime that secures browsers, tools and libs

https://thenewstack.io/nanoclaw-echo-agent-runtime/
1•four_fifths•7m ago•0 comments

AI's top startups are barely publishing their research

https://www.science.org/content/article/ai-s-top-startups-are-barely-publishing-their-research
6•YeGoblynQueenne•9m ago•0 comments

A Terrarium iPhone Case That Houses Actual Growing Plants and Moss

https://www.designboom.com/technology/transparent-phone-case-moss-plants-inside-terrarium-daniel-...
1•karakoram•9m ago•0 comments

Casio Moflin

https://www.casio.com/uk/moflin/
2•austinallegro•12m ago•0 comments

Models and theories of lambda calculus (2009)

http://inria.hal.science/inria-00380183v1/document
3•sargstuff•13m ago•0 comments

There is no way to know if an LLM API is manipulating you

https://twitter.com/lafalcemateo/status/2082250304330809738
1•lafalce•15m ago•0 comments

Dr. Dobb's Journal (articles from 1988 – 2009)

https://jacobfilipp.com/thedoctor/
2•slu•17m ago•1 comments

Expansion of Flock Camera Systems Is Leading to More False Positives

https://www.techdirt.com/2026/07/29/expansion-of-flock-camera-systems-is-leading-to-even-more-fal...
2•cdrnsf•19m ago•0 comments

Show HN: Dev-like – turn public engineering practice into agent skills

https://mrbro.dev/dev-like/
1•MarcusRBrown•19m ago•0 comments

AI in Education: Using ChatGPT Without Losing Critical Thinking

https://www.notion4teachers.com/blog/the-ai-learning-divide-when-technology-supports-thinking-and...
1•behindai•20m ago•0 comments

Modest Advice

https://stearnslab.yale.edu/modest-advice
1•geneticdrifts•20m ago•0 comments

Sequence Is Not Structure: Getting Lost in Long LLM Conversations

https://msg.samsonov.io/2026-07-28-sequence-is-not-structure/
2•kryzhovnik•23m ago•0 comments

Ask HN: What do you use for privacy focused analytics?

1•garyhbutton•25m ago•1 comments

Claude Outage – Mac Menu Bar Item

https://github.com/titojankowski/claude-status
3•tito•26m ago•1 comments

Sun Path and PPFD Simulator – Calculate Natural Sunlight and DLI for Plants

https://specled.com/en/w/sun-path-ppfd-dli-simulator/
1•Specled•27m ago•0 comments

The Cold Email

https://zachholman.com/posts/cold-email
1•holman•28m ago•0 comments

PyCuTe: Reference implementation and examples of the CuTe Layout

https://github.com/NVlabs/CuTe
1•matt_d•28m ago•0 comments

Word worm crawls into Copilot, spreads chaos

https://www.theregister.com/security/2026/07/29/word-worm-crawls-into-copilot-spreads-chaos/5280588
1•Bender•29m ago•0 comments

Are Yelp Reviews Harsher in Your City? (2019)

https://web.archive.org/web/20201127091203/https://nycdatascience.com/blog/student-works/are-yelp...
1•jerlam•29m ago•1 comments

Moscow slaps Telegram founder on wanted list, Durov responds one-finger salute

https://www.theregister.com/applications/2026/07/29/moscow-slaps-telegram-founder-on-wanted-list-...
2•Bender•29m ago•0 comments

When the Agent Shouldn't Run on Your Laptop

https://jasonrobert.dev/blog/2026-07-29-when-the-agent-shouldnt-run-on-your-laptop/
1•cebert•29m ago•0 comments

Who wins and who loses after US bans foreign robots?

https://arstechnica.com/ai/2026/07/who-wins-and-who-loses-after-us-bans-foreign-robots/
1•Bender•30m ago•0 comments

Show HN: Email domain health check with shareable reports

https://shipmail.to/tools/email-health
1•jcoulaud•30m ago•0 comments

Show HN: Rebuilding ARKit 3D Reconstruction

https://twitter.com/pablovelagomez1/status/2082256328966422542
3•pablovelagomez•30m ago•1 comments

Anthropic's Lonely Island

https://www.axios.com/2026/07/29/anthropic-claude-open-models-ban-china
1•theanonymousone•30m ago•0 comments

RCade: The Arcade Cabinet with CI/CD Deployment, Custom Graphics Card for CRT [video]

https://www.youtube.com/watch?v=W-OpIbLUOU0
1•evakhoury•31m ago•0 comments