What really bothers me is that after the model decides a task is large it starts reducing scope, or splitting it, or avoiding parts of the implementation that may be critical or add a lot of value.
So my guess is that we've ended up with unnecessarily cautious agents. But I don't want that, especially if I'm paying extra usage credits for it. I want it to be ambitious about what it can take on.
So my question is: should we consider “hours” and “days” the wrong units for an agent? Maybe it should estimate itself in some other way that reflects how much it can actually take on and deliver.