As soon as you add LLM fuzziness to a codebase, it seems to propagate like a virus. You can’t get anything reliably useful for machines downstream of a prompt—even with structured output & friends there’s always the possibility it’ll be flat-out wrong. Reminds me of the old “function coloring” issue except it’s probabilistic instead of async. Obviously it opens up incredible possibilities but I do miss the days where you could predict exactly what would happen by reading code.
justinweiss•28m ago
Yes! This is one of the major challenges I still have. It's easy to say "fix model errors in AGENTS.md / the system prompt." But finding a way to prove the effect of that change is the hard part.
It's so frustrating when a model says "no, your guidance was fine, I just didn't follow it, I'll do better next time..." It just makes me want to yell "No! You won't do better next time! This is the problem!"
> In that respect, working with an LLM is much more like working with another person than working with traditional software.
Yes, but if you work with another person, at least they'll have a chance of remembering suggestions you make.
The simple answer is "Have the AI do a refinement -> eval -> refinement loop." That pushes the problem somewhere else, into the "what goes in the eval" question ("and make sure you don't overfit" -- something AI-driven prompt refinements are not good at).