I can't complain about Opus 5.5, more than extracted my money's worth out of it but the final stages require pen-testing and Claude's masters stomp on it every time it tries to help do security reviews [cyber]. Ever so often it sneaks in a security fix. Even found a race condition in haproxy by mistake and Claude took it upon itself to find the crash string. That went horribly bad. Each time they try to up-sell Mythos and say I have to go through a verification program that I am not permitted to go through.
Aside from the uncensored Qwen forks, which frontier models can do extensive code security reviews, security fixes? Ideally something close to the quality of the NCC Group. This is for my own hobby craft. Maybe this does not exist and that is fine too.
bigyabai•59m ago
> Like Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse. We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests.
> We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts.
Bender•44m ago
Grok is telling me it can do code security reviews but I am not sure I believe it.