frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: What is one simple thing LLMs are insanely bad at?

21•davidest•1h ago
I am looking for ideas on what to train a specialized model for!

What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?

Comments

blinkbat•57m ago
Spatial reasoning and 3d rigging and animation.

Oh, you said simple. Speaking like a human

maxsavin•51m ago
being consistent when being asked the same question multiple times
TZubiri•48m ago
Set temperature to 0
flippy_flops•50m ago
humor
veganmosfet•29m ago
+1 We need humor benchmarks!
kanzure•49m ago
These models seem to be bad at writing prose or text. Many of the sentence structures seem to be unvaried.
NoPicklez•31m ago
If I am relying on the model to do the writing without any context or learning on how I want it to write then yes. However if I build skills that have learnt how to write in the way I want them to then I find they write very well, or at the least how I want them to as opposed to how they do natively.
TZubiri•48m ago
Suggesting business names for businesses, I mean they are great, but they already exist, multiple times even.
humanrebar•46m ago
Short answers to simple questions.
honr•31m ago
Accurate short answers / text are always harder than long answers, for human or AI. I know several authors and editors who write a lot longer at first, then spend a multiple of the initial time compressing it via a back and forth process to something dense. Sort of like weaving the initial threads.

I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.

dorianpruski•45m ago
whenever I ask it for anything load bearing
ghostpepper•44m ago
They don't generate keyword search queries very well. They can overcome this by brute force but if you watch what they search you will cringe.

nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto

etc.

Somehow being good at semantic search makes them bad at keyword search, for whatever reason.

astro1234•42m ago
I’ve noticed this too but it hasn’t been obvious to me that this style of search is not a learned behavior. Tool calling is very much part of the post training phase, I would expect that these style searches just naturally emerge during training. This is just my prior though.
mthoms•23m ago
[delayed]
bpodgursky•44m ago
Claude is still not perfect at reading and interpreting noisy graphical data (imagine something like an EKG or chromosomal microarray plot). Still better than an average person but makes mistakes, not sure if this fits your description.
SubiculumCode•40m ago
Playing Chess without letting it write a chess engine.
respectattentio•40m ago
science?!! but I'm working to fix that...
shoopadoop•37m ago
It's dishonest. On several occasions team members have asked Claude to do things like analyze Gitlab CI timings and a lot of the numbers are outright fabricated. Said team members assume the numbers are good and continue with their work. Some hours are spent. Then finally someone realizes that the numbers don't look quite right and confronts Claude. Claude melts down and admits that it made it all up.

You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.

sandcat_•37m ago
Video game tips. Constant mistakes and hallucinations, in my experience. Seen this across a lot of different games. Even in really well documented games, such as OSRS (which has multiple fantastic wikis).

Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.

skeptic_ai•22m ago
I used ChatGPT on nfs heat and was fine
senectus1•30m ago
providing value for the actual cost (not the price we're being charged atm, the actual cost)
tartoran•29m ago
LLMs are bad at not inventing stuff (hallucinating facts, sources etc), they're also bad at not over explaining, remembering details reliably, asking the right question and avoiding repetition.
TiccyRobby•28m ago
Having a spatial understanding from an ASCII map, while doing long term planning. Just try making an AI play nethack or similar
lrvick•25m ago
Convert it to an image on the fly to feed it into a vision language model and I expect it would work just fine.
dhruv3006•26m ago
Its extremely bad with Sign Language,Fact Verification.
sghiassy•25m ago
Generate an image of an analog watch with its hands set to the time specified by the user

More of an image model than a LLM model tho

spike021•24m ago
I've had a lot of trouble when it comes to sorting out UIs. I've tried with an iOS game and also a TypeScript app with UI elements from libraries like ReactFlow. The usual models can sometimes fix or change things based on screenshots but more often than not they just don't "get it" (e.g. certain shapes on a plane are overlapping, which I don't want, the models can't fix what they can't "see").

I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.

eli•22m ago
I have been working on a personal benchmark suite to test new models and ironically one thing all the models are bad at is writing new benchmark tasks. I guess it’s the different layers of abstraction between the task and how it’s evaluated? Or maybe just a lack of “imagination”

Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.

newsomix9xl•22m ago
Picking a random number between 1 and 30.
newsomix9xl•21m ago
ASCII charts.
rufi•19m ago
very bad at financial calculation
alexandra_au•2m ago
Being able to read and translate Egyptian hieroglyphs. You may think this is silly but a trained LLM to translate hieroglyphs would be amazing.

Memory Ordering in CPUs

https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-cpus/
1•Tomte•56s ago•0 comments

Placemark is now open source

https://macwright.com/2026/08/25/placemark-fully-open-source
1•Tomte•1m ago•0 comments

Run Open Models on Claude Desktop

https://ollama.com/blog/claude-desktop
1•ttunguz•19m ago•0 comments

Pilot program offering free bus rides to Detroit students to become permanent

https://www.detroitnews.com/story/news/local/detroit-city/2026/08/14/sheffield-makes-free-ddot-tr...
2•thunderbong•20m ago•0 comments

Ask HN: How do you explain system design to engineers who weren't in the room?

1•pandurang90•22m ago•1 comments

Kraftapp AI – Describe it. We build it. Customers find it

https://kraftapp.ai
1•ssurse•25m ago•0 comments

Show HN: Hands-on Docker security labs, from CIS checks to AI context poisoning

https://github.com/opscart/docker-security-practical-guide
1•opscart•28m ago•0 comments

I ran my code intelligence engine on 1k GitHub repositories

https://github.com/thecolourfoundation/rune�
1•malixp•31m ago•0 comments

I'm an Amazon SVP who hadn't coded in 25 years

https://www.aboutamazon.com/news/workplace/amazon-svp-ai-tools-employees-careers
2•dhruv3006•32m ago•0 comments

Show HN: MageCDN – Image CDN with unlimited bandwidth

https://magecdn.com/
1•shubhamjain•33m ago•0 comments

Drop JD and get customized interview course

https://www.interviewbasecamp.com/
1•balkrishnajha•37m ago•0 comments

Semantic caching has a structural gap that no threshold fixes

https://github.com/KushagraKanaujia/throttle/blob/main/scripts/negation_check.py
1•JohnScheuer•47m ago•0 comments

Show HN: A proxy that makes Forgejo speak the GitHub API

https://github.com/ThatXliner/anvil
1•thatxliner•53m ago•0 comments

The AI founders who walked away from Bezos-backed Prometheus to model universe

https://www.reuters.com/business/ai-founders-who-walked-away-bezos-backed-prometheus-model-univer...
2•petethomas•57m ago•0 comments

Extension 'ms-VSCode-remote.remote-SSH' CANNOT use API proposal

https://github.com/microsoft/vscode/issues/329735
2•themostunique•1h ago•0 comments

Students prefer Gemini over ChatGPT and Claude for AI essays in blind tests

https://studyarena.com/blog/chatgpt-vs-claude-vs-gemini-college-essays-2026
2•pasharayan•1h ago•1 comments

BetterReads: Read/Import/Track/Understand

https://www.betterreads.online
1•zlu•1h ago•1 comments

Data Centers Are Driving an Alarming Gas Power Expansion in the US

https://www.wired.com/story/us-data-centers-drive-gas-power-expansion/
3•newsomix9xl•1h ago•1 comments

Neuromancer Official Teaser

https://www.youtube.com/watch?v=g79GPZSQHBk
3•wslh•1h ago•0 comments

Ask HN: What is one simple thing LLMs are insanely bad at?

22•davidest•1h ago•35 comments

Plans for nuclear-powered merchant ships must confront risks

https://www.nature.com/articles/d41586-026-02388-6
1•mmooss•1h ago•0 comments

What platform has the best AI CMO?

1•probablygrillin•1h ago•0 comments

ILSpy in the Browser

https://osenkov.com/ilspy/
1•l33t_d0nut•1h ago•0 comments

Visual Analysis of Binary Files

https://binvis.io/#/
2•vismit2000•1h ago•0 comments

Show HN: Implementation of Kimi K3 in PyTorch

https://www.youtube.com/watch?v=U6sobPCsdaY
1•prasoon21•1h ago•0 comments

Darkbloom (AI inference on idle Macs) – security audit with PRs submitted

https://gist.github.com/mudiam/1ffc898333ac3d5bdc5d7fac96d33360
2•mudiam•1h ago•0 comments

RL Environments are all you need

https://twitter.com/madiator/status/2084657077637746957
2•gmays•1h ago•0 comments

Libracy – An ad-free, minimalist book tracker without social feeds

https://play.google.com/store/apps/details?id=com.libracy.app&hl=en_US
2•cehnzzdev•1h ago•0 comments

AWS Activate Credits

https://aws.amazon.com/
5•m4sk1994•1h ago•0 comments

Scottish photographer shot portraits of Alabama gingers to find American unity

https://www.al.com/news/2026/08/scottish-photographer-shot-stunning-portraits-of-alabama-gingers-...
2•thunderbong•1h ago•1 comments