Phenomenal model, not sure what they did, but I have been able to do so much with my $20 plan!
233mhz•28m ago
If enough people keep saying it I'm sure they will nerf it
amelius•6m ago
I think what makes it great is that they trained it to write harnesses for the code it writes, so it can test stuff even if the supplied code is not complete.
rdli•24m ago
It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.
9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
chewchewchew•19m ago
9 hours?!
rdli•14m ago
Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
danbrooks•8m ago
Agreed on Opus 5.5 being a great model. It's the first one that I trust for long running (>1 hour) tasks.
Handy-Man•33m ago
233mhz•28m ago
amelius•6m ago