5 ms·
Xiaomi Mimo 2.6 live post-training dashboard
- Toslink 9d ago[dead]
- wolttam 10d agoHah, it would be great to see more labs pick this up.
- krm01 10d agoThis is pretty neat. What would be a good reason for the other Model providers to not do this?
- kibae 10d agoSpeculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed. Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.
- jwpapi 10d agoI think first of all it’s not an obvious idea, also the marketing surplus for other providers is not as big for openai/anthropic as for xiaomi and last but not least I’m pretty sure you can withdraw methodology from here. I’m saying who has a million dollars for me, so I can make my own model?
- Bolwin 9d agoI don't really remember a situation, which of those models supposedly beat the other? I still opus 4.6 though not for code
- nikcub 9d agothis is remarkable transparency in an otherwise hyper competitive and secretive industry
- speedgoose 10d agoI didn't know 2 thirds of the training data would be source code.
- jerrygenser 10d agothat is the the "data used to improve the model" when signing up for the subscription plans
- leothetechguy 10d agothis is the rl run, not the pretraining run
- ahmadyan 10d agoeven in pre-training, usually 30%-50% is code these days.
- leothetechguy 9d agoThat would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
- thehamkercat 10d agoThis is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies
- medlazik 10d agoNeoliberalism, that famously open and transparent economic ideology
- atemerev 9d agoAh, one Donald Trump, a famous neoliberal.
- stymaar 9d agoWere Sam Altman and Dario Amodei different men before Trump was in charge?
- brookst 9d agoTo some degree, sure. Remember “open” AI? I don’t think Trump changed them, but Trump is absolutely a symptom of larger social collapse in the US, and that collapse has affected Altman and Amodei. We’re not even pretending that truth matters or that the wealthy can ever suffer consequences, and those two seem quite liberated by that.
- rozab 10d agoWhy are they doing this? To try head off accusations about distillation?
- bayindirh 10d agoSometimes you're confident about what you're doing and show how you work to the world. Keeping the garage door open, or at least making the door translucent. It's always cool.
- jampekka 10d agoThat China's official policy is now to prefer open models and open model development may be a part of it.
- Aboutplants 10d agoWith that policy in place, labs might be incentivized to be creative in their openness. This being fun/free PR
- culi 10d agoBRICS just had a New Delhi meeting where Xi pushed a 5-point plan on AI cooperation that centered on open source models
- anemic 10d agoBottom of the page says "Open is what we value."
- hsbalanxvxjsmab 9d agolol that’s rich
- brookst 9d agoI don’t see how it would head off such accusations. This is post-training, and even it’s data could be pulled from other models or run against other models in realtime. Not saying that’s the case, just that the dashboard does not disprove.
- liuliu 10d agoWhen you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
- lucrbvi 10d agoThey are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.
- jampekka 10d agoKinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data. I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks. https://en.wikipedia.org/wiki/Training,_validation,_and_test_data_sets https://en.wikipedia.org/wiki/Training,_validation,_and_test...
- liuliu 10d agoCorrect. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.
- SwellJoe 10d agoYou gotta have something to aim at. And, presumably, the benchmark is not part of the training data, it is the test against which the model is tested at each stage; is behavior moving in the right direction?
- esafak 10d agoNot if you don't train against them.
- 10d ago
- ProfessorLayton 10d ago2.6 Pro: >started 2026-09-15 10:32 UTC For some reason I thought training took much, much longer than what the progress bar suggests. This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
- joelwallis 10d agoI been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it. -- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
- james2doyle 10d ago2.5 Pro or the regular 2.5? I always found that those Mimo models to be really good at tool calling and following instructions
- walrus01 10d agoI've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.
- deleted 9d ago[deleted]
- girvo 9d agoThe fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.
- walrus01 9d agoYeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it.
- levocardia 10d agoYou'd think they would make it less obvious that they are running their whole operation with Claude
- SwellJoe 10d agoIt's not obvious to me. What's the tell?
- ricardobeat 9d agoIf you're thinking of the UI style, definitely not Claude. It is incapable of writing a clear sentence like "what each step's samples are made of", would have used all-caps for everything, more padding and gradients.
- conception 9d agoI hope this is /s because it’s very easy to get Claude to write sensibly. That’s why AI slop writing is so annoying because it’s so easy to avoid with any amount of effort at all.
- ricardobeat 9d agoIn my experience Opus and Sonnet 5 subtly ignore most instructions related to writing style, and continue to sound the same half of the time. Do you have a successful skill/prompt to share?
- conception 9d agoMy AI Slop Tells Reference - https://pastebin.com/ZG4S485y https://pastebin.com/ZG4S485y is referenced by... My STE-Rewrite Skill - https://pastebin.com/8Z9GUqAX https://pastebin.com/8Z9GUqAX which uses... My readability skill - https://pastebin.com/sMkMgEUM https://pastebin.com/sMkMgEUM which uses... the analysis script - https://pastebin.com/hmc2zn1i https://pastebin.com/hmc2zn1i My use case is generally easy to read instructions for lay people of an international/ESL audience. Have it write it's whatever and then run that on it and it comes out... actually pretty good. Use it for emails, etc, when it doesn't need a personal touch and just needs to be clear. These skills won't get you a snazzy blog post, but I imagine could be augmented to produce something significantly better than the incomprehensible non-sense that it spews out by default.
- impulser_ 10d agoThe Chinese labs are just making fun of the US labs at this point. Where is the cool shit from the US labs?
- culi 10d agoWith other software, devs convince their managers of the importance of using open source stuff in their stack. With AI, it's usually managers choosing what models to use for the devs. The US labs don't need to give a damn how much devs like open source
- noir_lord 10d ago> The US labs don't need to give a damn how much devs like open source In the short term, true. In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.
- impulser_ 10d agoThis isn't about liking open source. This is about the labs just being cool and doing cool shit instead of the opposite which is Anthropic where all they talking about is killing everyone and taking everyone's job.
- dlisboa 9d agoThese labs are still (for the time being) made of people, who reflect their lives onto the work. The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.
- culi 9d ago> This is about the labs just being cool and doing cool shit The point still stands. Devs like "cool shit". Upper management doesn't care
- 9d ago
- esafak 10d agoThat's the kind of transparency we need! That DeepSWE benchmark puts it in frontier territory: https://artificialanalysis.ai/agents/coding-agents?coding-agents-performance-chart=deep-swe-v1.1 https://artificialanalysis.ai/agents/coding-agents?coding-ag...
- fzysingularity 10d agoVery cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.
- dr_dshiv 10d agoWell, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.
- dzonga 9d agothe open burial started when zAI served their latest model on all Chinese chips. now we r just noticing the grave getting dug deeper.
- skybrian 9d agoFor my own usage, Luna is cheap enough that I don't care if other models are cheaper. I'm interested if another model is in some way better and not too expensive.
- rapind 9d agoLuna is great but makes a lot of mistakes at high and lower in my experience (large rust codebase). I use Luna Max for asynchronous subagent reviews and am very happy with its work, but it’s slow af.
- ijidak 9d agoWhat plan are you on? Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.
- epolanski 9d ago
- passive 10d agoNeat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far. I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size. The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.
- ricardobeat 9d agoFor reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great. Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort). https://deepswe.datacurve.ai/blog/deepswe-v1-1 https://deepswe.datacurve.ai/blog/deepswe-v1-1
- Cookingboy 9d ago2.6-pro just reached 63.7% by step 10, it's on step 11 right now. Even flash reached 60.7% by step 12, and it's on step 16 now. This is so exciting lmao.
- arcanemachiner 9d agoDeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.
- brookst 9d agoIt is saturated, but that doesn’t mean worthless. Seeing 72% is low-signal, but 30% is still meaningful.
- markasoftware 9d agogemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure
- ehsankia 9d ago> and google just started letting all their engineers use claude That's misleading. 1. Having different models available is useful for A/B testing and helping improve Gemini itself. 2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
- buffalobuffalo 9d ago
- ernsheong 9d agoMino 2.5 has been my workhorse for coder and tester agents (the ones planner agents delegate tasks to)
- dude250711 9d agoDistillation in real-time? Very interesting!
- Cookingboy 9d agoThat "training cost" is just live revenue count for Anthropic/OpenAI API calls! /s
- jstummbillig 9d agoWow, spending money on training an almost-frontier-model is much more time intensive than I thought it was.
- dr_kiszonka 9d agoVery curious that everyone here (so far) seems to assume this dashboard presents real data.
- hsbalanxvxjsmab 9d agoHaha yeah pretty wild how easily you can see the data is fake by the repeating numbers (refresh the page the progress goes back in time constantly) + watch for restarts. They say they happen but 0 data correlates the log messages. Just a replay of old data or being fed by an llm so they convince people they are open
- Bolwin 9d agoThe intermediate tickers are fake but real data comes in and resets it. Its like a progress bar essentially. We don't call progress and bars fake
- hsbalanxvxjsmab 9d agoI do when their fake like this site is. Insane people blindly believe this stuff
- ttul 9d ago$5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.
- stymaar 9d agoWhich isn't that much when you compare to the kind of DC that US actors are using.
- brookst 9d agoDo we know what kind of DC US actors are using specifically for training, versus inference and delivery?
- stymaar 8d agoI remember Zuck bragging about using a 100MW DC for training, and Musk's Colosus was supposed to be a training data center (but they fucked up the design so they had to repurpose it to an inference one).
- lostmsu 9d agoThat's posttraining. Pretraining is the expensive part.
- heronbank 9d ago[dead]
- rao-v 9d agoI absolutely love that someone is doing this! Why isn’t IBM for Granite or Google for Gemini? If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world? I’m genuinely learning quite a bit just from the dashboard
- gtirloni 9d agoThey think they have the special sauce. Even if they do, what would they get in return for doing that?
- pppkin 9d ago[dead]
- thenews 9d agobeen using the 2.5 mimo for side projects, works amazing
- hsbalanxvxjsmab 9d agoThis is so very clearly fake? See the message stating the flash 2.6 flash run was restarted and 0 graphs correlate that restart
- Retro_Dev 9d agoA restart of the process does not necessarily mean reverting the model state. I don't know why you would even do that, because you'd lose all the progress you made.
- hsbalanxvxjsmab 9d agoIt said restarted step 15 5 mins ago and the progress showed they were working on step 16 for a day
- tcbbd 9d ago[dead]
- kkotak 9d agoWouldn't us observing this break down the model superposition and make it dumber? :)
- brcmthrowaway 9d agoFound the Dark Matter (2024) watcher
- ssn2000 9d agoTotal run cost is $1.2M until now, what resources are they using to train their model? Wish they shared more details on that and what the MFU metrics are.
- sinuhe69 9d ago[flagged]
- wg0 9d ago"Slow down this much openness in AI or we won't get our trillion dollars valuations!" Google had this GPT long go and a wise man within Google noted: "We don't have any maot neither does anyone else." The AI bubble burst is guaranteed and is only delayed by IPOs.
- b3lvedere 9d agoI wonder what we will do with the discarded data centers and its hardware..
- Ylpertnodi 9d agoCopper can be stole, but that's already in progress.
- user43928 9d agoNothing is guaranteed. Open models have not yet caught up with February's Mythos checkpoint. Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.
- wg0 9d ago"Stealing millennium problems" would be more complete if not accurate description. And that 99.99% of the market is not interested in solving millennium problems is the other fact.
- Alifatisk 9d agoCan we call this open AI?
- monneyboi 9d agoRefreshing, now let's make this a default feature. I imagine a "Upcoming models" list with links to these kind of dashboards.
- alescalaios 9d ago[dead]
- singularity2001 9d ago"Claude Distill Requests":'hidden'
- Retro_Dev 7d agoDoes it even matter if it distills other models? ethically no, because the information was originally stolen from us users... it is indirectly harming anthropic if it is happening, but who cares? It is better for everyone except anthropic that anthropic doesn't have a monopoly.