1. choose a good model
2. choose the appropriate reasoning effort
3. choose a prompt to nudge it even further
Then it comes down to understanding the intuition of what kind of task deserves what effort?
(Couldn't even get this to happen in a 30 person team...)
I use these models all day long, experiment, and have no clue how to choose that.
It just showed up one day in the interface with no explanation or guidance. Its use is mysterious, its effects unclear except through intensive experimentation, and to this day it mostly seems "how many bad decisions will Claude go forward with when it finally dumps out screenfuls of text instead of getting better guidance early on" though it's certainly not a guarantee on anything.
These models are being released at breakneck speed even before their creators know how to use them. It's a big project of collective discovery to figure out what they are doing and how to use them.
Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.
This is getting stupid folks.
IMO, human in the loop is the only serious usage of AI (I know, I know, "software factories bro"). Everything else is a hope and a prayer and a big bill.
We're sorta in the age of alchemy. Lots of cranks out there, but there are real recipes.
FinnLobsien•25m ago
I believe that we're in a scenario where usage is unlikely to go down and neither are frontier AI costs.
I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
A marketer doesn't need to default to Opus 5.5 to upload a blog article with MCP, which could be done by a model 10% of the price.
awesan•5m ago
After we started excessively documenting and extracting skills everyone has stopped complaining about running out of tokens, because agents stopped having to reconstruct the full context each time from scratch. Harnesses like claude code also push the model to aggressively keep this documentation in sync so there's little concern about drift.
The worst thing to do from a token usage pov is to give a model a vague open ended prompt because they are so scared to be wrong that they'll waste a ton of tokens "thinking" through the issue and verifying everything. Whereas they almost trust skills blindly and skip all this unnecessary work.
skeptic_ai•3m ago