Mistral slightly proving me wrong (and I'm not mad).
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
And if not, why do they exist?
> that isn't owned
I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.
Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.
That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.
Wouldn't it be better to do something more like Cursor, and RL on an existing pretrained model if you're not innovating anyway?
Tais-toi et prends mon argent!
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
I'd live to have one like that but EU made.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I'm excited to try this out today.
Deepseek essentially releases instruction manuals in paper form.
DeepSeek is basically a research lab founded by a hedge fund guy, more than anything else.
I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.
This model looks reasonably cheap. Though not deepseek levels.
Going to test it with Hermes, wondering where it will land in term of capability.
Bon chance, Mistral!
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
Europe needs profitable AI companies, not money pits.
GLM-5.3: 753B, 40 Active
I was hoping for something that hinted at smaller models too, but I guess not.
Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.
Good enough to show competence, and instill confidence in the team/company. Later releases can be more efficient.
I think it's a great release with that framing.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
And if not, why do they exist?
Their reason to exist is to ensure the sovereignty of France.
Not everyone is wild about being downstream of either the Chinese or US governments, particularly when it comes to things like cybersecurity
And if not, why do they exist?
Its very xenophobic of you to say China has zero intention of helping humanity, and just wants to "wage economic warfare".
Last time I checked, it was ourselves (USA) waging economic warefare on 2/3rds of the world.
I dont get this cope people have where people have this idea that its impossible for a Chinese company (that make billions of dollars) to have done something by their own merit, but instead its always some Chinese Communist Party conspiracy where the main goal is to destroy America.
Lay off twitter for a bit.
We hear this about literally every industry the Chinese excel in - that it's only because the government subsidizes them that they succeed. For chip manufacturing, for batteries, for EVs, for solar, for AI. I don't see how the chinese government can afford to subsidize all of these industries and still have them contribute to the GDP.
So the most common way to publish manuals?
But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.
The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.
If they can do that, they'll have customers.
There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.
I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.
And yes, open weights are still behind, but are catching up.
"Now, here, you see, it takes all the running you can do, to keep in the same place. If you want to get somewhere else, you must run at least twice as fast as that!"
People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.
Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.
Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."
Yeah, not that smart.
As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.
I hear this literally every other week about whatever the newest FoTM model is.
Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.
If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.
The second moat is convenience, which all the big labs make it (comparatively) easy to glide into their models.
If they can carve out a niche of industrial and governmental partners who rely on them for sovereignty reasons, it may be enough.
delillos•39m ago
sofixa•38m ago
Important for sovereignty, multi-language support, and choice.
imjonse•36m ago
Dr4kn•25m ago
imjonse•19m ago
baby•35m ago
rwillmann•30m ago
bitnovus•37m ago
eigenspace•35m ago
It is of vital strateigic importance for Europe (and really the rest of the world too) that there are non-American, non-Chinese options for AI.
badatnames•30m ago
I realised after writing this, the EU as is uncomfortably often the case, may be the real forcing function for what happens with US policy irrespective of the media campaigns we're presently seeing. Here's hoping for a steady trickle of stale ChatGPT weights leaking from EU infra providers in the long term.
Tade0•16m ago
Specialists are expensive everywhere. China is functionally 80s Japan surrounded by several Brasils and all the frontier AI work is being done in that first part, where labour is expensive - if only due to competition for top talent.
eigenspace•13m ago
Yes, if the Chinese stop trading, you can still use the existing solar panels (unlike the gas which you literally set on fire), but it's nonetheless a major vulnerability to not have any local know-how in creating important infrastructure.
Giving up all AI know-how and expertise to China just because they currently share their models would be a generational mistake.
Europe already got burned hard by this sort of thing too recently, and is now very sensitive to strateigic depencencies, and is working hard to lift them where possible.
233mhz•33m ago
rahen•19m ago