frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Open-source model routing for coding agents at Astra-level performance

27•adchurch•1d ago
A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.

First of all, a quick explanation: the Weave Router (https://github.com/weave-os/router) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.

What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at https://weaveos.com/router!)

It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.

1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.

Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.

In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but how it got there. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.

Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!

2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.

3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.

We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.

Our router is open source (https://github.com/weave-os/router) so anyone can try it out. Or if you prefer you can use our hosted version (https://weaveos.com/router).

Comments

redrove•58m ago
Is the model you trained available as open weights?
thefourthchime•20m ago
Interesting work, and thanks for describing how your router works internally. It's definitely a fascinating subject. How would you say this compares to Cursor's auto mode?
aschla•13m ago
And similarly, Copilot’s Auto mode?

Show HN: Vote on which of Hacker News' challenges for AI have been met

https://stoppels.ch/goalposts/
25•stabbles•1h ago•27 comments

Show HN: Open-source model routing for coding agents at Astra-level performance

30•adchurch•1d ago•3 comments

Show HN: Breadcrumb, record everything on your mac + context manager for AI

https://innerloop.works/breadcrumb
5•jv22222•1h ago•0 comments

Show HN: I built an app that scours the internet to create a daily news briefing

https://github.com/yogthos/newsroom
2•yogthos•1h ago•0 comments

Show HN: MCP server for editing Excel with real-time calc and diffs

https://github.com/pixelsmasher13/gridpath
4•escapingsingula•1h ago•0 comments

Show HN: Winamp-style native audio player for macOS

https://vibeplayer.app/
4•cmicali•1h ago•0 comments

Show HN: Ledge.sh – Runnable Markdown Notes

https://ledge.sh
189•dancablam•1d ago•85 comments

Show HN: I made a computer vision tool for evaluating deadlift form

https://github.com/jeremyipark/vision-demos/tree/main/deadlift
3•dr_blueberry•1h ago•0 comments

Show HN: TrueScribe – Linux local AI transcription studio

https://truescribe.app/
5•tinykerneldev•3h ago•0 comments

Show HN: Gutsy, a 0.8B Jev-compatible decision model that runs on your CPU

https://github.com/kouhxp/gutsy
5•mrkn1•3h ago•0 comments

Show HN: OpenC6 v2.0 – Bare-metal BIOS and RISC-V microkernel for ESP32-C6

https://github.com/Rompass/openc6-bios
4•Rompass•3h ago•0 comments

Show HN: Q-DeflateRFC1951 compatible hyper-density compression SaaS

https://microforce.dev
2•Gen-N•3h ago•0 comments

Show HN: CoIsland – Your whole engineering stack in a notch

https://coisland.app/
4•LeDonT•3h ago•1 comments

Show HN: A working 3D model of an Enigma machine

https://enigma.design
105•primitivesuave•2d ago•40 comments

Show HN: The Zoomist – looking and listening to ordinary things more closely

https://www.youtube.com/@TheZoomist
2•mmdtdev•1h ago•0 comments

Show HN: JBR-001 – An open-source 3D printable desktop robot

https://projecthub.arduino.cc/syntheticaidata/jbr-001-a-desktop-companion-robot-powered-by-arduin...
130•gvuksic•2d ago•31 comments

Show HN: Yantra – an LALR(1) parser generator for C++

https://github.com/TantrixAuto/yantra
30•renjipanicker•16h ago•15 comments

Show HN: Dental Scope – Interactive 3D dental anatomy

https://dental-scope.com/
127•Zeruxe•2d ago•66 comments

Show HN: NSL – WSL for Linux

https://frostyard.github.io/nsl/
162•bketelsen•2d ago•107 comments

Show HN: The Kio Programming Language

https://jdevuyst.github.io/kio/
5•jdevuyst•6h ago•0 comments

Show HN: Using 2D DFT, dithering, etc. to maximize eInk manga image quality

https://github.com/ciromattia/kcc
82•seam_carver•3d ago•20 comments

Show HN: Real-time Solar System with 526k asteroids and all tracked satellites

https://space.bl2.net/
393•wanick•1d ago•107 comments

Show HN: Walk the Endless Dungeon

https://the-endless.pages.dev/
4•olup•9h ago•1 comments

Show HN: Perspica – A semantic diff for reviewing code

https://github.com/sshah03/perspica
14•sshah03•22h ago•5 comments

Show HN: Parrot – Open-Source Smart Meeting Recorder with Co-Pilot on Mac

https://openparrot.app
34•turantekin•1d ago•29 comments

Show HN: Fast browser agent (using Jev) with deep reasoning as needed

https://github.com/ironbee-ai/ironbee-express
6•berkay•3h ago•0 comments

Show HN: Strata – an expressive semantic layer that can say no to your LLM

https://strata.do/
21•ajoski9•1d ago•14 comments

Show HN: HN.watch – Videos of all Hacker News posts

https://hn.watch/
219•mrborgen•3d ago•97 comments

Show HN: Destroy Any Website with Stickman

https://destroy.spritefusion.com/
192•HugoDz•3d ago•41 comments

Show HN: Classican, listen + watch Classical music in a fun way

https://classican.vercel.app/
7•jaequery•12h ago•2 comments