What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
More of an image model than a LLM model tho
I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.
blinkbat•57m ago
Oh, you said simple. Speaking like a human