Just a simple reactor, my laptop's only little. But still.
That's not a boast, I don't think I was particularly good at that back then, e.g. I didn't really get how to think about automated tests until much later.
It's just to say that no, coding and software engineering are not the same thing. "Code Monkey" is a dead (or perhaps "undead") role now, but it wasn't always so.
Long term planning in LLMs has not been solved.
GitHub's Copilot cloud agent offering is suffering with a case of some of the worst corporate ADHD I've seen. We built a cloud agentic development pipeline on it, and it seems like almost every other week they silently change something with zero public announcement or documentation that creates real disruption for our team.
That's real, breaking changes to the platform that clearly aren't being tested/reviewed before being pushed to prod. Again with zero public announcement or documentation.
Support is useless – we're paying customers in the 4-5 figures and our tickets go unanswered.
I have no doubt that if you provide any AI system with an oracle with expected behavior that it can match that oracle with some amount of $ and tokens. I haven't seen any demonstration of anything else. Rewriting a codebase was always a challenge for humans not because of complexity, but because of the time and effort involved in matching the old version's prior behavior. It doesn't have anything to do with the serious level of work required to build something truly new from scratch in a performant way.
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
Brings to mind this classification https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...
"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""
Instead, I think what's closer to solved and what we're in the process of solving is product development.
Story: A while ago, I had a few programmers who were really, really fast almost always missed the mark on the assignment wrong. I loved having them on projects because in the time my senior precise engineers could deliver a MVP, the fast engineers would build the wrong thing, collect feedback, reiterate, build the wrong thing, collect feedback, eventually inching closer and closer to a product people would pay for, and it would almost always get delivered faster than my seniors.
I feel AI does the same thing.
I got lazy around claude fable and astra, and asked them to work in loop (pick specified issue, develop it, qa it ...) have a separate CTO checking on arch.
at the end both models swore that the code is perfect and well designed and nothing is lacking.
I ran the software and it suddenly started writing large amount of data to CSV files instead of the typical DB usage.
AI decided to use csv for testing, and just drifted away. 0 regards to the actual project, 0 regards to common sense.
anecdotal but really weird, the project category is rather standard, I wouldn't accept such a mistake from a junior developer.
I don't like the feeling being judged and tested by the author (missing number 5 point in the list).
This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.
There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.
So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.
The difference seems to be that some companies are above the law apparently.
Dear lord. Is that supposed to reflect the average thoughts and motivation of a person you want to hire? Or that of their employer?
Doesn't matter what you think about AI, "it isn't perfect" is clearly a nonsense reason not to object to it.
It's for example impossible to have a discussion with an LLM where you both learn something which you can apply tomorrow. The LLM doesn't learn until the next model is released and by then your discussion is just a tiny fraction of the training data (if present at all). AGENTS.md, skills and so on are just a proxy for what we actually want, an agent that listens and understands. A proxy mind you, that requires constant tweaking with no sign of generalisation in sight.
I'm also not sure what humans being non-deterministic even means here. The point is if you're comparing results with NFR, pure agentic coding falls short.
AI can write CRUD API endpoints almost perfectly now. It can also write quicksort, a heap, whatever much quicker than I can.
It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).
I suspect this will remain a problem for the models for a long time. All the things that the models are currently good at are the low hanging fruit of reinforcement learning for coding.
Think about the kind of reinforcement learning environment that needs to be created to train a model to become good at building and designing large scale software end to end. It would be a slog because you need to build the large scale software up front and then break it down to train the model to construct it in a systematic manner that allows for the software to evolve. And then you need enough of these training environments for it to generalize. I think they will eventually figure it out though but it may take a while.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
That concept might work a lot of the time but you will definitely run into situations where that'll never produce a correct or working response. To actually learn something you need an environment/playground to apply what you think you know and observe the results. Without that you're not really learning, you're jus regurgitating what people want to hear.
Especially expensive when you take into account the amount of that code which must have been boilerplate & meta-code in nature, meaning it should have been straightforward to move.
Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).
> I hope you can see the stupidity here if you expect to see any deterministic results at all.
Are you expecting humans to be deterministic in the code they produce?
This is all that's needed to actually use LLMs nowadays. How is it a "multiplier" rather than an "equalizer"?
Because without the responsible human engineer in the loop, it'll all gradually decay in a cascade of edge-cases. This happens with human written code as well (every "we'll replace this prototype before we ship" you've ever worked on), but with LLMs it happens at 10-100x the rate.
You will be surprised how many times, catches errores made by the AI coding agent. However,as you point, isn't deterministic. And you can guarantee the end results is 100% fine code
What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check.
hanifbbz•32m ago
If anyone has counter-arguments or cares to make me smarter, I'm all ears.
boxed•25m ago
Insanity•11m ago
But maybe I'm misremembering how fragile GH was in the 2010s.
automatic6131•23m ago