I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.
Perhaps DS works better where targets have low-hanging fruits than can be found fast?
Given how fast and cheap DS is, it's just an ideal model with enough "IQ" to let it loose. Another thing they left out of the article, DS becomes really good with if provide custom tools for the task, on it's own it's mediocre.
“If we ban CFCs now the Chinese will win!”
“If we ban chemical weapons, nuclear weapons, etc etc our enemies will triumph! They won’t stop!”
“If we switch to biodegradeable plastic then our rivals will have an advantage.”
“If we dont externalize the costs to our population, then they will, and then will win!”
I think workflows can do the job agents do, 20x cheaper and more predictably and safely. They can completely displace agents, just as HFCs displaced CFCs and then we were able to ban CFCs and phase them out through international COOPERATION. The language of COOPERATION is what saves us vs COMPETITION is all about cutting corners and externalizing costs. Google the Montreal Protocol, Geneva Conventions, Nuclear Non Proliferation Treaty, Unleaded Gasoline etc etc.
Agents have got to be marginalized. They are just popular because the labs need to make a ton of money for their investors and recoup their massive spending on training models.
fwip•44m ago
tyingq•33m ago
> . The accepted runs cost $4.65. Failed attempts and replacement runs increased the complete cost to $5.14.
fwip•25m ago
Aldipower•32m ago