frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Reright: Human approval for agent-written text

https://reright.it/
1•paulofilip3•36s ago•0 comments

Voice assistant that runs on your own Claude or ChatGPT plan

https://memoe.app/
1•yisraelgottz•1m ago•0 comments

Screens Aren't Destroying Young Minds. I Should Know

https://humanprogress.org/screens-arent-destroying-young-minds-i-should-know/
1•obscurette•1m ago•0 comments

GPT 6 Astra as an Embodied Policy

https://anonymous-report-421.github.io/public-website/?view=1
2•gmays•7m ago•0 comments

Show HN: Jotbus – a shared encrypted scratchpad for coding agents

https://jotbus.com/
3•shake-n-fries•7m ago•0 comments

Show HN: ChatGPT answers calls on a normal SIM. No Twilio, just a $6 ESP32

https://www.youtube.com/watch?v=zIVXu6g8qdY
2•nadam54321•8m ago•0 comments

Oracle Triggered the Implosion of the AI Bubble

https://medium.com/predict/what-everyone-had-been-waiting-for-and-fearing-oracle-just-triggered-t...
4•mpweiher•9m ago•0 comments

Show HN: Genjutsu AI – Replace objects in existing videos

https://genjutsuai.net/
2•CoderLim110•10m ago•0 comments

Agentic Software Development Hypothesis

https://brooker.co.za/blog/2026/05/20/hypothesis.html
3•hnisjafx40•10m ago•0 comments

What Is Codemode

https://lucumr.pocoo.org/2026/10/6/codemode/
3•Tomte•11m ago•0 comments

Southern Olive Oil

https://ambrook.com/offrange/supply-chain/southern-olive-oil
2•surprisetalk•12m ago•0 comments

If somebody tries to hot-patch an already-hot-patched function

https://devblogs.microsoft.com/oldnewthing/20261005-00/?p=112755/
1•ibobev•12m ago•0 comments

Ship the Code – Release the Feature

https://osscontainertools.org/blog/featureflags/
1•martizih•14m ago•0 comments

Doorman Fallacy

https://www.jaakkoj.com/concepts/doorman-fallacy
1•softwaredoug•15m ago•0 comments

How HN: Duck Out – webcam duck hunt you play with your hands

https://getduckout.com/
1•eekay•16m ago•0 comments

Show HN: Kahawai – An open source, modular media system

https://kahawai.net/
3•berm_•16m ago•0 comments

Show HN: Vev – Ask questions about a screenshot, get a probability per answer

https://github.com/Xiaooolong/vev
1•CountingSheeep•17m ago•0 comments

Why Semiconductor Bottlenecks Move Rather Than Disappear

https://michaelhillaert1.substack.com/p/why-semiconductor-bottlenecks-move
1•michaelhillaert•18m ago•0 comments

Show HN: TagBackup – Back up files to S3 with human-readable tags

https://tagbackup.com/
2•joncombe•19m ago•1 comments

Auteur Theory of Design [video]

https://www.youtube.com/watch?v=xk3UcgbbmxQ
2•tosh•20m ago•0 comments

Jiro's Dream (2012)

https://karrisaarinen.com/jiro/
1•tosh•21m ago•0 comments

I built this extension because students waste too much time on pointless exams

https://canvascrack.com/en
3•iamtheamerican•21m ago•0 comments

AI Certification Council (AICC) – AI education and AI technical skills

https://www.aicertificationcouncil.com/
1•aicc•21m ago•0 comments

Show HN: Parseable, an open observability datalake, handles 100M time-series/min

https://www.parseable.com
2•yashdotrv•22m ago•0 comments

Ask HN: What models and harnesses are you using that are not Claude or Codex?

1•dataviz1000•22m ago•2 comments

Show HN: Azurade, image and video generation with MCP, credits never expire

https://azurade.com/developers/
1•delneg•23m ago•0 comments

Extra Headroom in prod: Input -34%, Output -33% across 183 Claude Code users

https://extraheadroom.com/blog/claude-code-savings-real-usage
1•gghootch•24m ago•0 comments

Show HN: ADHDev – one task queue per repo for any coding CLI, any machine or OS

https://github.com/vilmire/adhdev
2•vilmire•25m ago•0 comments

Show HN: Codebase-guide: get the onboarding doc nobody ever had time to write

https://github.com/dimitritholen/codebase-guide
1•scriptdude•26m ago•0 comments

ChronoRoam - Simple trip planner, no maps, for Notes and Sheets people

https://chronoroam.app/#/demo
1•solnguyen93•26m ago•0 comments
Open in hackernews

Mistral Large 4

https://docs.mistral.ai/models/mistral-large-4-0
235•Philpax•37m ago

Comments

delillos•31m ago
wow, another large language model from another company. groundbreaking.
sofixa•30m ago
It's the only non-Chinese open weights frontier model, so while not groundbreaking per se, still quite important.

Important for sovereignty, multi-language support, and choice.

imjonse•28m ago
There are many non-Chinese open weights models, just not very good ones.
Dr4kn•17m ago
That's why he has written "frontier"
imjonse•11m ago
my bad, I had missed that word.
baby•27m ago
Inkling!
rwillmann•22m ago
You have a number of other European providers of open-weight LLMs, such as Bielik and Aleph Alpha. They are not trying to compete at the frontier, but they sometimes develop original architectures. I am also a big fan of PleIAs' research in this regard, have a look at their Baguettotron
bitnovus•29m ago
Mistral was silent on the frontier for quite some time so this is exciting to me.
eigenspace•27m ago
To Europeans at least, this is a big deal, especially because of cybersecurity capabilities.

It is of vital strateigic importance for Europe (and really the rest of the world too) that there are non-American, non-Chinese options for AI.

badatnames•22m ago
Non-American maybe, but non-Chinese is likely impractical. We might not want their APIs, but we can't compete on energy and labour for training, even if it means buying licenses to host the weights (which CADA encourages). From China's perspective, what's the case for baking weights they can't sell? Some Chinese models already don't even have the Taiwan politics stuff baked in, those filters are only in their APIs.

I realised after writing this, the EU as is uncomfortably often the case, may be the real forcing function for what happens with US policy irrespective of the media campaigns we're presently seeing. Here's hoping for a steady trickle of stale ChatGPT weights leaking from EU infra providers in the long term.

Tade0•8m ago
> we can't compete on energy and labour costs for training,

Specialists are expensive everywhere. China is functionally 80s Japan surrounded by several Brasils and all the frontier AI work is being done in that first part, where labour is expensive - if only due to competition for top talent.

eigenspace•6m ago
Using Chinese models is a bit like using Chinese solar/batteries instead of American gas.

Yes, if the Chinese stop trading, you can still use the existing solar panels (unlike the gas which you literally set on fire), but it's nonetheless a major vulnerability to not have any local know-how in creating important infrastructure.

Giving up all AI know-how and expertise to China just because they currently share their models would be a generational mistake.

Europe already got burned hard by this sort of thing too recently, and is now very sensitive to strateigic depencencies, and is working hard to lift them where possible.

233mhz•25m ago
It's hacker "news" not hacker "scientific breakthrough"
rahen•11m ago
Mistral is the only major non-Chinese contender in the AI race releasing open-weight models. I’d rather trust that French weights haven’t been backdoored than Chinese ones. Heck, I’d even trust Mistral more than Anthropic for sensitive work.
crimsoneer•30m ago
Woah, this seems like a big deal (assuming the benchmarks are as good as claimed)?

Mistral slightly proving me wrong (and I'm not mad).

cbg0•30m ago
Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)
Narciss•30m ago
Le Chaton Fat is here!
HelloUsername•26m ago
Le Chonk https://www.youtube.com/watch?v=hD51W2txi1Y
Narciss•24m ago
Yeah just saw that, I'm gonna keep converting it in my head.
Philpax•29m ago
https://venturebeat.com/technology/mistral-debuts-large-4-le...
imjonse•29m ago
The blog post https://mistral.ai/news/mistral-large-4/
staticman2•27m ago
Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.
keithnoizu•26m ago
touche
throwa356262•18m ago
It certainly has helped OpenAI and Anthropic get their KV cache costs under control.
Tade0•12m ago
It's no secret that everyone is dis-stealing from everyone else.
aeneas_ory•27m ago
Benchmarks are better than expected! And probably got there without distillation ;)
water-drummer•7m ago
Is there a reason to believe why they wouldn't distill locally running open weights Chinese models?
mcbuilder•27m ago
Looks like they are doing 50% off to stay price competitive with DS Flash V4.1
prodigycorp•24m ago
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.

Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.

Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.

I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.

erichocean•22m ago
Does Mistral ever advance the state of the art on any dimension?

And if not, why do they exist?

CJefferson•17m ago
Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?
embedding-shape•17m ago
Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.
eigenspace•15m ago
Has the French military advanced the state of the art on any dimension in the last hundred years?

And if not, why do they exist?

pizzafeelsright•13m ago
Second hand rifle distributor?
azan_•6m ago
Of course they have!
swiftcoder
tdubey•22m ago
Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?
irl_zebra•10m ago
Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!
scrubmunch•22m ago
wowza le models a heckin chonker
nsbk•21m ago
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.

Tais-toi et prends mon argent!

quadruple•8m ago
I believe it is already available in the API no?
4rtem•20m ago
Previous one is barely in top 50 on arena.ai
maxdo•20m ago
Not bad only two major releases behind top tier. Edit : checked its rather 3 generations behind . Oh well
Roark66•20m ago
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.

I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.

I'd live to have one like that but EU made.

saberience•20m ago
Looks like it's about a year behind still. i.e. its intelligence is behind models from roughly a year ago.

https://www.vals.ai/benchmarks/vals_index

eigenspace•19m ago
Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.

I certainly wouldnt have predicted that 10 years ago.

Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.

I'm excited to try this out today.

apexalpha•16m ago
I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.

Deepseek essentially releases instruction manuals in paper form.

wg0•13m ago
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
fsmedberg•9m ago
DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government. The reason they release the AI models is economic warfare against US, not because of charity or kindness. It's great for us consumers, but the goal is not to help humanity or open-source.
disgruntledphd2•7m ago
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.

DeepSeek is basically a research lab founded by a hedge fund guy, more than anything else.

apexalpha•17m ago
Excited to hear this!

I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.

This model looks reasonably cheap. Though not deepseek levels.

Going to test it with Hermes, wondering where it will land in term of capability.

Bon chance, Mistral!

jakozaur•16m ago
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).

Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.

zkmon•13m ago
Europe needs a lot of these. Quickly. Way to go, Mistral! Keep 'em coming.
spiderfarmer•10m ago
> Europe needs a lot of these.

Europe needs profitable AI companies, not money pits.

ofirg•13m ago
where does sit on the pareto distribution compered to Le Chaton Fat?
tosh•12m ago
sorting the charts like that gives off weird vibes

https://mistral.ai/news/mistral-large-4/

walrus01•9m ago
The terminalbench 4.0 score is encouraging as a sign of it not doing anything "stupid" when put in a proper harness.
Luker88•7m ago
Mistral Large 4: 1050B, 49 Active

GLM-5.3: 753B, 40 Active

I was hoping for something that hinted at smaller models too, but I guess not.

Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.

skc•5m ago
We're probably fast approaching the scenario where the cheapest models will win out.
•
14m ago
So that there exists an EU-native option in the near-frontier LLM space?

Not everyone is wild about being downstream of either the Chinese or US governments, particularly when it comes to things like cybersecurity

whstl•7m ago
Does erichocean ever advance the state of the art on any dimension?

And if not, why do they exist?

eckelhesten•7m ago
Source on the allegations = crackpipe
tacomagick•6m ago
So is American models. They are subsidised heavily but still can't provide cheaper access. Their fault is to assume all countries can afford them.
AIblemblio•16m ago
For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.

But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.

eigenspace•12m ago
I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.

The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.

If they can do that, they'll have customers.

spiderfarmer•11m ago
Less and less work requires a frontier model though.
senordevnyc•11m ago
Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...

There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.

I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.

calgoo•11m ago
Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.
senordevnyc•8m ago
We've been hearing the line about them only being a few months behind for a year now, during which time O/A have grown their revenue like 10x, haven't they?
wg0•11m ago
No it is not. Only maybe for the noobs or vibe coders.

People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.

Aldipower•5m ago
Despite Opus 5.5 got really bad the last days for me. Looks like they nerfed it again. This is extremely unreliable.
senordevnyc•16m ago
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?

Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.

If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?

swiftcoder•12m ago
> But in terms of actual revenue, is there really any chance of anyone catching the big labs?

I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.

If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.

senordevnyc•10m ago
Maybe. I can't freaking wait for the IPO filings so we can finally put all this to rest. (haha, like that'll actually put it to rest on HN, but at least we'll have better data)
WarmWash•11m ago
No, because compute, not model ability, is the moat.

The second moat is convenience, which all the big labs make it (comparatively) easy to glide into their models.

eigenspace•11m ago
The big lab revenue may not be catchable, but im not sure it needs to be.

If they can carve out a niche of industrial and governmental partners who rely on them for sovereignty reasons, it may be enough.

senordevnyc•7m ago
I completely agree, I think AI is a vast ecosystem will all kinds of profitable niches and sub-markets.