What I had success with (although benchmarks are older) is to index the codebase with a dedicated local code embedding model. It's a bit expensive on the CPU side but in my benchmarks it reduced token use and wall clock time significantly. Of course, it's always dependent on statistical noise + host system load, and running sufficiently large benchmarks is simply too expensive, so take em with a grain of salt.
Why does it work you may ask? Well, LLMs basically brute force words/phrases and pipe that into find/grep/pgrep/whatever (or as recently discussed here write a python script for it - https://news.ycombinator.com/item?id=49654229). Semantic search looks for similarities so you have to do less brute forcing. Comes of course at the cost of indexing everything first.
You can find the project here: https://github.com/ory/lumen
``` Average cost per attempt, without → with RTK:
Claude/Fable: $1.72 → $1.64 (~5% cheaper) DeepSeek: $0.115 → $0.121 (~5% more expensive)
Almost all Claude savings came from a single task. Excluding it, savings were under 1%. ```
It took me a few rereads to parse out the top-line. This article really buries the lede.
vrighter•1h ago
nextaccountic•22m ago