frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Open-source model routing for coding agents at Astra-level performance

2•adchurch•46m ago
A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.

First of all, a quick explanation: the Weave Router (https://github.com/weave-os/router) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.

What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at https://weaveos.com/router!)

It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.

1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance. In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. This significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!

Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.

2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.

3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.

We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.

Our router is open source (https://github.com/weave-os/router) so anyone can try it out. Or if you prefer you can use our hosted version (https://weaveos.com/router).

Show HN: Let's Chat – A retro mIRC-style web chat app

https://letschat.exepad.app/
1•umut-neo•1m ago•0 comments

OpenResearch: A local-first workspace for research agents

https://github.com/alphaXiv/OpenResearch
1•criexe•2m ago•0 comments

Chompi portable sampler instrument is now open-source (hardware and software)

https://www.chompiclub.com/opensource
1•lashkari•3m ago•0 comments

Abbott launches Freenome's colorectal cancer screening tes

https://www.medtechdive.com/news/abbott-launches-freenomes-colorectal-cancer-screening-test/831771/
1•brandonb•3m ago•0 comments

Chinese-owned science journal created to rival the best

https://www.nature.com/articles/d41586-026-02759-z
1•bookofjoe•5m ago•1 comments

Cognition exploiting privileged info from competitors

https://twitter.com/matansf/status/2105335179502064038
1•AznHisoka•6m ago•0 comments

Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engine

https://modal.com/blog/quail-billion-tpm
1•gmays•6m ago•0 comments

Lichen – A local, BOY model system1 (Jev) server with image support

https://github.com/Mushroom-Systems/lichen
1•jeff_ciesielski•7m ago•1 comments

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

https://github.com/magnitudedev/magnitude
6•anerli•7m ago•1 comments

Show HN: Kyno – Direction as a first-class primitive for AI agents

https://cizambra.github.io/kyno/
2•cizambra•9m ago•0 comments

So what's happening with Russia's new, long-delayed crewed spacecraft?

https://arstechnica.com/space/2026/09/so-whats-happening-with-russias-new-long-delayed-crewed-spa...
1•rbanffy•10m ago•0 comments

Helheim-Emacs: Helheim to Helix is what Spacemacs and Doom are to Vim

https://github.com/helheim-emacs/helheim
1•philonoist•11m ago•0 comments

Show HN: I built a free Burp/Caido alternative that needs zero setup

https://apiaxess.dev
11•coding-maniac•13m ago•3 comments

Claude Says

https://ohhfishal.net/Posts/claude
2•speckx•13m ago•0 comments

Why Hawaii Is the Only US State with a Foreign Flag Inside Its Own

https://www.thecollector.com/hawaii-us-state-foreign-flag-inside/
2•Tomte•14m ago•0 comments

OpenAI is accused in lawsuit over rogue AI agents

https://www.fastcompany.com/91615759/openai-is-accused-of-taking-dangerous-cyber-risks-for-privat...
2•skinfaxi•14m ago•0 comments

Show HN: Token compression CLI to save Codex/Astra costs

5•yolandac•14m ago•0 comments

Show HN: Honorfit – a pushup app blocker for your phone

https://honorfitapp.com/
3•tnrich•15m ago•1 comments

What it means for a machine to understand what we want

https://medium.com/deepsense-ai/agents-are-already-super-powerful-but-can-they-understand-us-8a06...
1•nvmdbljstm•16m ago•0 comments

Trump decrees era of 'Super Intelligence' upon us

https://www.theregister.com/ai-and-ml/2026/09/30/trump-decrees-era-of-super-intelligence-upon-us/...
2•beardyw•16m ago•1 comments

RFK Jr outlines expansive vision for collecting US health data at Maha event

https://www.theguardian.com/us-news/2026/sep/29/rfk-jr-collect-health-data
1•text0404•16m ago•0 comments

Hijacking Copilot Cowork's AI Gateway to Bypass Sandboxing and Exfiltrate Files

https://www.promptarmor.com/resources/hijacking-copilot-coworks-ai-gateway-to-exfiltrate-files
2•jerryShaker•17m ago•0 comments

Bayesians Are Frequentists

https://statmodeling.stat.columbia.edu/2018/06/17/bayesians-are-frequentists/
2•theanonymousone•18m ago•0 comments

OpenZL v0.3.0: a major upgrade of native LZ engine and Compression Transformer

https://github.com/facebook/openzl/releases/tag/v0.3.0
1•ksec•19m ago•0 comments

Extraordinary Multi-Agent Delusions and the Madness of Crowds

https://freesystems.substack.com/p/extraordinary-multi-agent-delusions
1•paulpauper•21m ago•0 comments

Half of world’s population exposed to dangerous levels of ozone in 2026

https://www.scientificamerican.com/article/dangerous-ozone-threatened-half-of-worlds-population-t...
2•ck2•21m ago•0 comments

Weather risk is reflected in Florida home prices

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7509658
2•paulpauper•21m ago•0 comments

Yu Deng

https://time.com/collection/time100-next/2026/yu-deng/
1•tzury•21m ago•0 comments

Does Fertility Stabilize at Low Levels? Evidence from Period, Cohort

https://www.nber.org/papers/w35824
1•paulpauper•21m ago•0 comments

Apple's HomePad smart home hub launching on Oct 13 with iMac G4 design

https://9to5mac.com/2026/09/30/apples-homepad-smart-home-hub-launching-on-oct-13-bloomberg/
1•apparent•26m ago•2 comments