frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Xiaomi Mimo 2.6 live post-training dashboard

https://mimo.xiaomi.com/rl/
106•krackers•1h ago

Comments

wolttam•56m ago
Hah, it would be great to see more labs pick this up.
krm01•55m ago
This is pretty neat. What would be a good reason for the other Model providers to not do this?
kibae•45m ago
Speculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed.

Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.

jwpapi•4m ago
I think first of all it’s not an obvious idea, also the marketing surplus for other providers is not as big for openai/anthropic as for xiaomi and last but not least I’m pretty sure you can withdraw methodology from here.

I’m saying who has a million dollars for me, so I can make my own model?

speedgoose•52m ago
I didn't know 2 thirds of the training data would be source code.
jerrygenser•46m ago
that is the the "data used to improve the model" when signing up for the subscription plans
leothetechguy•46m ago
this is the rl run, not the pretraining run
ahmadyan•10m ago
even in pre-training, usually 30%-50% is code these days.
thehamkercat•48m ago
This is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies
medlazik•21m ago
Neoliberalism, that famously open and transparent economic ideology
rozab•47m ago
Why are they doing this? To try head off accusations about distillation?
bayindirh•40m ago
Sometimes you're confident about what you're doing and show how you work to the world.

Keeping the garage door open, or at least making the door translucent. It's always cool.

jampekka•24m ago
That China's official policy is now to prefer open models and open model development may be a part of it.
Aboutplants•18m ago
With that policy in place, labs might be incentivized to be creative in their openness. This being fun/free PR
liuliu•47m ago
When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
lucrbvi•44m ago
They are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.
jampekka•39m ago
Kinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data.

I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.

https://en.wikipedia.org/wiki/Training,_validation,_and_test...

liuliu•35m ago
Correct. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.
SwellJoe•37m ago
You gotta have something to aim at. And, presumably, the benchmark is not part of the training data, it is the test against which the model is tested at each stage; is behavior moving in the right direction?
ProfessorLayton•46m ago
2.6 Pro: >started 2026-09-15 10:32 UTC

For some reason I thought training took much, much longer than what the progress bar suggests.

This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.

GaggiX•41m ago
These are post-training reinforcement learning steps.
krackers•38m ago
Yes, updated the submission title to say "post-training" to hopefully prevent further confusion
joelwallis•45m ago
I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.

The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.

-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

james2doyle•41m ago
2.5 Pro or the regular 2.5?

I always found that those Mimo models to be really good at tool calling and following instructions

walrus01•24m ago
I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.
jwpapi•2m ago
May I ask why you ended up there instead of just using the heavy subsidized subscription. I’m actually curious.
levocardia•42m ago
You'd think they would make it less obvious that they are running their whole operation with Claude
SwellJoe•38m ago
It's not obvious to me. What's the tell?
impulser_•17m ago
The Chinese labs are just making fun of the US labs at this point.

Where is the cool shit from the US labs?

Training a 4B model to produce 81% faster query plans than Postgres

https://rohanbansal.com/qorl
226•polyphilz•2h ago•36 comments

Xiaomi Mimo 2.6 live post-training dashboard

https://mimo.xiaomi.com/rl/
108•krackers•1h ago•30 comments

Breaking the 1.58-bit Barrier for Ternary LLMs

https://arxiv.org/abs/2609.16338
28•matt_d•40m ago•0 comments

macOS 27 Golden Gate – Review

https://arstechnica.com/gadgets/2026/09/macos-27-golden-gate-the-ars-technica-review/
53•Brajeshwar•1h ago•39 comments

Small programming tricks

https://will-keleher.com/posts/small-programming-tricks-matter/
301•signa11•5h ago•154 comments

Reversing Factorio's RNG

https://gegell.github.io/posts/factorio-rng/
61•jheitmann•4d ago•6 comments

Accurate Models of AMD Matrix Cores

https://arxiv.org/abs/2609.14845
42•matt_d•2h ago•5 comments

Performance Improvements in .NET 11

https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-11/
63•soheilpro•1d ago•3 comments

Vectorized and performance-portable Quicksort (2022)

https://opensource.googleblog.com/2022/06/Vectorized%20and%20performance%20portable%20Quicksort.html
154•mococa•3h ago•24 comments

How good are frontier models at physics?

https://arxiv.org/abs/2609.13009
39•qt31415926•2h ago•13 comments

AWS says it can't restore some data from mideast facilities struck by Iran

https://www.wsj.com/world/middle-east/aws-says-it-cant-restore-some-data-from-mideast-facilities-...
59•berkeleyjunk•23h ago•20 comments

Anatomy of a Texture

https://agentlien.github.io/texture/
52•Agentlien•7h ago•10 comments

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

https://arxiv.org/abs/2609.14858
165•bananaflag•7h ago•48 comments

Mistral X Mozilla: Private, Multilingual AI Browsing

https://mistral.ai/news/mistral-x-mozilla/
495•vertigoruntime•13h ago•175 comments

Japan's book scene is moving from bookstores to libraries

https://untranslatedjp.substack.com/p/japans-book-scene-is-quietly-moving
40•herbertl•3d ago•4 comments

WalShadow: Sub-second Postgres replication to ClickHouse from physical WAL

https://clickhouse.com/blog/introducing-walshadow
18•spathak•5d ago•3 comments

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations

https://github.com/arnegiacomo/fugleramme
2005•arnemunthekaas•1d ago•234 comments

The Siberian Ice Maiden and the Scythian World

https://patrickwyman.substack.com/p/the-siberian-ice-maiden-and-the-scythian
40•NaOH•1d ago•1 comments

Anecdotally, programmers dislike "reduce"

https://evanhahn.com/posts/2026-09-13-programmers-dislike-reduce/
38•vinhnx•2d ago•74 comments

Tell the speakers that you liked their talks

https://ohhelloana.blog/tell-the-speakers/
252•whisper2020•1d ago•70 comments

Training Text-to-Image Models 3.6× Faster

https://www.linum.ai/field-notes/jit-ddt
12•schopra909•4h ago•1 comments

Show HN: Restarted – a 2026 remake of the classic 2015 startup generator

https://restarted.io/
11•zxcvbn4038•1d ago•1 comments

Show HN: AttaLambda: a language where types and data are made of untyped lambdas

https://attalambda.com
27•kserrec•2d ago•0 comments

A warning about 'model welfare'

https://mustafa-suleyman.ai/a-warning-about-model-welfare
161•andsoitis•7h ago•416 comments

Kyber (YC W23) Is Hiring a Forward Deployed Engineer

https://www.ycombinator.com/companies/kyber/jobs/eturrAR-forward-deployed-engineer
1•asontha•9h ago

The DeepMind Institute

https://institute.deepmind.com/
103•vertigoruntime•7h ago•33 comments

Claude Cowork and chat are now one Claude

https://claude.com/blog/cowork-is-now-claude
178•vertigoruntime•5h ago•185 comments

How big are factorials?

https://eli.thegreenplace.net/2026/how-big-are-factorials/
71•ibobev•1d ago•30 comments

Douglas Adams and the exterminated Doctor Who adventure

https://www.bbc.co.uk/news/articles/c8jdp38z4jgo
67•6LLvveMx2koXfwn•2d ago•63 comments

Randomized query complexity can beat certificate complexity

https://arxiv.org/abs/2609.15063
3•porridgeraisin•1d ago•0 comments