frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Apple introduces M6 and M5 Ultra

https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-perform...
1027•interpol_p•15h ago•948 comments

FDA authorizes first wearable device that monitors ketone and blood sugar levels

https://www.fda.gov/news-events/press-announcements/fda-authorizes-first-wearable-device-continuo...
323•sunnynagra•9h ago•158 comments

OpenAI Jalapeño: Better than Nvidia Blackwell

https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
386•bmulholland•14h ago•257 comments

New Mac Studio with M5 Max and M5 Ultra

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
725•interpol_p•15h ago•481 comments

Queryable Executables

https://fzakaria.com/2026/08/24/actually-queryable-executables
53•rguiscard•4h ago•7 comments

Black hole singularity is a surface not a point

https://arxiv.org/abs/2608.21590
213•raattgift•11h ago•152 comments

Maiao: Gerrit-style code review workflow for GitHub, GitLab, Gitea, others

https://github.com/runetes/maiao
48•zdw•6h ago•19 comments

New Mac mini, featuring M6 and M5 Pro

https://www.apple.com/newsroom/2026/08/apple-unveils-a-more-powerful-mac-mini-featuring-the-all-n...
457•runako•15h ago•281 comments

When str.lower() is a security vulnerability in Python – Seth Larson

https://sethmlarson.dev/when-str-lower-is-a-security-vulnerability
89•rbanffy•7h ago•38 comments

C2PA Cameras Do Not Survive Contact with Reality

https://www.da.vidbuchanan.co.uk/blog/android-c2pa.html
106•Retr0id•9h ago•56 comments

Agentic Context Management: Memory and Cost as Architecture Problems

https://arxiv.org/abs/2607.21503
9•gdad•2h ago•3 comments

Nitter and XCancel receive cease and desist notices

https://github.com/zedeus/nitter/issues/1442
725•Banditoz•11h ago•620 comments

Show HN: TeXbrain, a LaTeX editor that runs pdfTeX in the browser via WASM

https://github.com/swimmingbrain/texbrain
70•swimmingbrain•6h ago•11 comments

Run OpenBSD on DigitalOcean for $4/month

https://nil.wallyjones.com/run-openbsd-on-digitalocean-for-4month/
138•speckx•11h ago•65 comments

Building a backyard office, the build and cost breakdown

https://www.imkylelambert.com/articles/building-a-backyard-office-the-build-and-cost-breakdown
296•surprisetalk•14h ago•202 comments

Bomb fishing is wreaking havoc on Indonesia's coral reefs

https://e360.yale.edu/digest/bomb-fishing-coral-reefs
289•speckx•14h ago•147 comments

Ask HN: What is one simple thing LLMs are insanely bad at?

22•davidest•1h ago•38 comments

Tooltips need a delay, and then they need to skip it

https://blog.master.dev/tooltips-need-a-delay-and-then-they-need-to-skip-it/
133•ibobev•12h ago•34 comments

Show HN: LatticeDB – Like SQLite but for graph databases

https://github.com/jeffhajewski/latticedb
128•smiths1999•11h ago•36 comments

Don't Wordle

https://dontwordle.com/
324•Hbruz0•16h ago•117 comments

Dolly Parton has died

https://www.theguardian.com/music/2026/aug/25/dolly-parton-country-singer-dead
1328•helsinkiandrew•10h ago•201 comments

Show HN: I made a Raspberry with Qwen my local car AI

https://github.com/ThinkOffApp/CarWatch
117•petruspennanen•13h ago•30 comments

My Friend Aaron

https://rorz.io/writing/my-friend-aaron
489•sarreph•11h ago•133 comments

A brief history of federal lift ticket regulation

https://zakpodmore.substack.com/p/a-brief-history-of-federal-lift-ticket
50•CGMthrowaway•9h ago•6 comments

Firefox 157 will include JPEG XL by default on all platforms

https://groups.google.com/a/mozilla.org/g/dev-platform/c/3YMV4MS34KA?pli=1
313•yboris•10h ago•86 comments

Visualizing Binary Files

https://movq.de/blog/postings/2026-08-05/0/POSTING-en.html
92•zdw•1d ago•17 comments

Tracking Costco gas prices

https://www.jack.bio/blog/costco-gas-tracking
96•lafond•1d ago•98 comments

Clara (YC P26) is hiring a growth engineer to bring AI doctors to market

https://www.ycombinator.com/companies/clara-2/jobs/8snci6k-founding-full-stack-growth-engineer
1•gfavvas•11h ago

Show HN: I built self-hosted deployment automation tool for Windows and IIS

https://fdeploy.com/
27•dt3ft•21h ago•13 comments

The brain may be about to have its Ozempic moment

https://www.economist.com/science-and-technology/2026/08/11/the-brain-may-be-about-to-have-its-oz...
64•Anon84•4h ago•47 comments
Open in hackernews

Ask HN: What is one simple thing LLMs are insanely bad at?

22•davidest•1h ago
I am looking for ideas on what to train a specialized model for!

What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?

Comments

blinkbat•1h ago
Spatial reasoning and 3d rigging and animation.

Oh, you said simple. Speaking like a human

maxsavin•1h ago
being consistent when being asked the same question multiple times
TZubiri•1h ago
Set temperature to 0
flippy_flops•1h ago
humor
veganmosfet•47m ago
+1 We need humor benchmarks!
kanzure•1h ago
These models seem to be bad at writing prose or text. Many of the sentence structures seem to be unvaried.
NoPicklez•49m ago
If I am relying on the model to do the writing without any context or learning on how I want it to write then yes. However if I build skills that have learnt how to write in the way I want them to then I find they write very well, or at the least how I want them to as opposed to how they do natively.
TZubiri•1h ago
Suggesting business names for businesses, I mean they are great, but they already exist, multiple times even.
humanrebar•1h ago
Short answers to simple questions.
honr•49m ago
Accurate short answers / text are always harder than long answers, for human or AI. I know several authors and editors who write a lot longer at first, then spend a multiple of the initial time compressing it via a back and forth process to something dense. Sort of like weaving the initial threads.

I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.

FriedFishes•18m ago
If I had more time I'd write a shorter letter, a la Pascal.

Editing is generally hard work, at the current token price I don't mind spending multiple passes of high effort to get down to a reasonable noise/signal ratio. I've seen some people pass off output to a weaker/cheaper model but that makes me a bit nervous when I don't have intimate knowledge of the subject.

dorianpruski•1h ago
whenever I ask it for anything load bearing
ghostpepper•1h ago
They don't generate keyword search queries very well. They can overcome this by brute force but if you watch what they search you will cringe.

nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto

etc.

Somehow being good at semantic search makes them bad at keyword search, for whatever reason.

astro1234•59m ago
I’ve noticed this too but it hasn’t been obvious to me that this style of search is not a learned behavior. Tool calling is very much part of the post training phase, I would expect that these style searches just naturally emerge during training. This is just my prior though.
mthoms•41m ago
[delayed]
jedbrooke•18m ago
that and always putting the “current year” at the end of the search term (so the results are more recent, I guess?), except that “current year” consistently ends up being 2-3 years ago since I guess that’s what’s in the training data (even on a harness that injects the current date)
areoform•17m ago
This behavior is a learned adaptation. It's most likely a feature not a bug.

Based on personal usage, I think it reflects search engine functionality degradation. I've found LLM keyword combinations are more likely to find the results I want with most search engines than mine. Including the big one.

The big one had solved this issue a long time ago by generating those associated keywords based on your input keywords, but somehow, something, somewhere has degraded that system to the point of inanity. And so here we are.

bpodgursky•1h ago
Claude is still not perfect at reading and interpreting noisy graphical data (imagine something like an EKG or chromosomal microarray plot). Still better than an average person but makes mistakes, not sure if this fits your description.
SubiculumCode•58m ago
Playing Chess without letting it write a chess engine.
respectattentio•57m ago
science?!! but I'm working to fix that...
shoopadoop•54m ago
It's dishonest. On several occasions team members have asked Claude to do things like analyze Gitlab CI timings and a lot of the numbers are outright fabricated. Said team members assume the numbers are good and continue with their work. Some hours are spent. Then finally someone realizes that the numbers don't look quite right and confronts Claude. Claude melts down and admits that it made it all up.

You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.

sandcat_•54m ago
Video game tips. Constant mistakes and hallucinations, in my experience. Seen this across a lot of different games. Even in really well documented games, such as OSRS (which has multiple fantastic wikis).

Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.

skeptic_ai•40m ago
I used ChatGPT on nfs heat and was fine
senectus1•48m ago
providing value for the actual cost (not the price we're being charged atm, the actual cost)
tartoran•47m ago
LLMs are bad at not inventing stuff (hallucinating facts, sources etc), they're also bad at not over explaining, remembering details reliably, asking the right question and avoiding repetition.
TiccyRobby•46m ago
Having a spatial understanding from an ASCII map, while doing long term planning. Just try making an AI play nethack or similar
lrvick•43m ago
Convert it to an image on the fly to feed it into a vision language model and I expect it would work just fine.
dhruv3006•44m ago
Its extremely bad with Sign Language,Fact Verification.
sghiassy•43m ago
Generate an image of an analog watch with its hands set to the time specified by the user

More of an image model than a LLM model tho

spike021•42m ago
I've had a lot of trouble when it comes to sorting out UIs. I've tried with an iOS game and also a TypeScript app with UI elements from libraries like ReactFlow. The usual models can sometimes fix or change things based on screenshots but more often than not they just don't "get it" (e.g. certain shapes on a plane are overlapping, which I don't want, the models can't fix what they can't "see").

I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.

eli•40m ago
I have been working on a personal benchmark suite to test new models and ironically one thing all the models are bad at is writing new benchmark tasks. I guess it’s the different layers of abstraction between the task and how it’s evaluated? Or maybe just a lack of “imagination”

Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.

newsomix9xl•39m ago
Picking a random number between 1 and 30.
newsomix9xl•39m ago
ASCII charts.
rufi•37m ago
very bad at financial calculation
elliotto•34m ago
They aren't funny. The jokes they come up with are extremely lame and the sort of thing you would expect a company HR manager to tweet.

I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.

alexandra_au•20m ago
Being able to read and translate Egyptian hieroglyphs. You may think this is silly but a trained LLM to translate hieroglyphs would be amazing.
alexeldeib•18m ago
What issues do you see in practice? This seems pretty easily "fixable"
znnajdla•18m ago
Editing a document without mixing edit instructions into the final document. Claude and ChatGPT do this all the time: I tell them to change X in a planning document or email draft, and instead of just changing X they also frequently add the edit instruction to “change X” into the document itself. They seem unable to take a step back and look at the document without “becoming” the document somehow. I do believe that dedicated subagents for editing may fix this but I am not sure.