From their abstract: "Our findings reveal that despite continuous development activity and growing codebase complexity of the agent harnesses, there is no statistically significant improvement in SWE-bench benchmark score (i.e., resolve rates of bugs) across releases for a given fixed LLM version. Worse, later agent harness versions consume nearly double the computational tokens and tool calls without corresponding quality gains. We explain this paradox from two angles: at the project level, we identify the development patterns (e.g., feature additions, fix-heavy releases, scattered small changes) correlating with quality fluctuations, while at the architecture level, we localize regressions to specific high-risk architectural components. Our findings call for a new practice of quality assurance in the development of agent harnesses."
wek•51m ago