frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

https://cognition.com/blog/swe-2
58•seelos•1h ago

Comments

mydreamof•35m ago
Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
harmonic18374•28m ago
Probably, FrontierCode is made by Cognition itself. The model also seems worse in every way than DeepSeek v4.1 Flash, launched today.

Also the submitter's account is very new which makes me suspicious of self-promotion.

_doctor_love•34m ago
SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.
Tsarp•33m ago
"SWE-2 is post-trained from Kimi K3"
airstrafer•17m ago
Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.

Maybe still worth it if their "64% cheaper" figure holds.

Tsarp•12m ago
With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.
samyok•12m ago
SWE-2 is free for all subscribers on the CLI to try out for the next month :)
xlbuttplug2•5m ago
I presume post training is significantly easier than the distillation/training the top Chinese labs are doing.

I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of derived models.

scronkfinkle•32m ago
Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
samyok•15m ago
SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI (https://docs.devin.ai/cli)

:)

monkeydust•27m ago
As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere.
postalcoder•18m ago
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

TheJCDenton•17m ago
> SWE-2 is post-trained from Kimi K3

On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

nullbio•13m ago
Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?

I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

ltsSmitty•10m ago
Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at
llmslave•10m ago
At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).

I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.

Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job

eyeris•7m ago
Wonder if this was the model that drove factoring the rsa-260

The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.

Shopify moves back to Native from React Native

https://shopify.engineering/back-to-native
360•fnthawar2•2h ago•244 comments

Rust Is Tier-1 Language at Microsoft

https://rustfoundation.org/media/guest-post-rust-is-tier-1-language-at-microsoft/
262•mmastrac•3h ago•125 comments

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

https://cognition.com/blog/swe-2
64•seelos•1h ago•23 comments

Neki by PlanetScale

https://neki.dev/
35•handfuloflight•57m ago•4 comments

More questions about whether researchers can trust OpenAI with unpublished math

https://mathstodon.xyz/@andreasthom/117240535270608201
87•pred_•9h ago•302 comments

Hitachi launches CO2 heat pump water heaters with solar-friendly tariff controls

https://www.pv-magazine.com/2026/09/07/hitachi-launches-co2-heat-pump-water-heaters-with-solar-fr...
201•thelastgallon•1d ago•157 comments

NASA Color Trick Was Meant for Mars. Now It's Unveiling Rock Art on Earth

https://gizmodo.com/this-nasa-color-trick-was-meant-for-mars-now-its-unveiling-rock-art-on-earth-...
19•gumby•1h ago•1 comments

>10x More Efficient Pretraining

https://magic.dev/blog/pretraining#
45•ronfriedhaber•1d ago•12 comments

DeepSeek v4.1 Flash

https://twitter.com/deepseek_ai/status/2097930608790167907
753•Liwink•10h ago•402 comments

Neki

https://planetscale.com/blog/introducing-neki
36•simon_weber•1h ago•5 comments

Casablanca: How an unproduced play marched into movie history

https://www.thecollector.com/casablanca-unproduced-play-movie-history/
14•mdp2021•53m ago•3 comments

Software Drives People Insane

https://graybeard.ing/software-drives-people-insane/
11•rglover•31m ago•2 comments

What algorithm did Windows XP use to choose your initial user picture?

https://devblogs.microsoft.com/oldnewthing/20260909-00/?p=112683
259•soheilpro•7h ago•125 comments

List of references on Sony websites to players "owning" their digital games

https://consumerrights.wiki/w/Sony_PlayStation_digital_game_ownership_lawsuit
215•haunter•4h ago•71 comments

One resignation turned the embers of AI fear into a wildfire

https://www.interconnects.ai/p/one-resignation-turned-the-embers
13•pretext•35m ago•10 comments

Stockfish 19

https://stockfishchess.org/blog/2026/stockfish-19/
186•atiedebee•3d ago•120 comments

iPhone Duo

https://www.apple.com/iphone-duo/
1337•thecosmicfrog•22h ago•2334 comments

To write non-fiction, draw the trunk, then the rest of the tree

https://devz.cl/posts/how-to-write/
47•DanielVZ•2d ago•9 comments

Python sets and dictionaries can have quadratic-time performance

https://lemire.me/blog/2026/09/03/python-sets-and-dictionaries-can-have-quadratic-time-performance/
18•ibobev•2d ago•5 comments

Serverless DTLS

https://proxylity.com/docs/listeners/dtls.html
7•mlhpdx•55m ago•3 comments

Show HN: What if the speed of light was 5 km/h?

https://rivendell.dmitrybrant.com/relativity/
527•dmitrybrant•14h ago•224 comments

Native Python and TypeScript Drivers for ArcadeDB, from OpenAPI and Protobuf

https://arcadedb.com/blog/arcadedb-native-drivers-python-typescript/
4•lvca•30m ago•0 comments

The first drink-driving conviction may have happened in London

https://www.ianvisits.co.uk/articles/the-worlds-first-drink-driving-conviction-may-have-happened-...
12•beardyw•9h ago•15 comments

Show HN: Filament – Fast data movement engine in Go

https://github.com/galaxy-io/filament
14•ikswolzok•2d ago•1 comments

Silicon Valley Is Transforming the Military-Industrial Complex

https://costsofwar.watson.brown.edu/paper/how-big-tech-and-silicon-valley-are-transforming-milita...
4•paimapi•57m ago•1 comments

What do Visa and Mastercard do? An intro to card networks

https://tautology.town/2026/06/01/card-networks.html
632•evakhoury•1d ago•378 comments

Show HN: Art – draw one stroke, let symmetry complete it

https://mrdee.in/mandala/
68•cyb0rg0•5d ago•28 comments

Show HN: Syq – copy files between machines fast (better than rsync)

https://greaber.github.io/syq/
6•greaber•1h ago•2 comments

Who Dung It? (Turdle.fun)

https://turdle.fun/
12•nb_quant•2h ago•17 comments

Growing proof that autonomous cars save lives

https://spectrum.ieee.org/are-self-driving-cars-safe
427•bookofjoe•23h ago•750 comments