17 ms·
MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
- Oras 4mo ago1k TPS is great, but I’m more fascinated by the amount of AI generated comments in this thread!
- eli 4mo agoLike what?
- adam_arthur 4mo agoThere are many with subtle tells. Not nearly as obvious as the ones from 6 months ago, but seems to be more the use of hyperbolic phrasing in a particularly unnatural way. The assess/explain, then hyperbole at the end kind of structure. Top comment looks suspicious from this perspective, but it's kind of a losing battle to be able to differentiate them with sufficient accuracy anyway
- marknutter 4mo agoThis is very reminiscent of the "everyone's a Russian bot" era of social media, where everyone would just lob that accusation at people without any real proof.
- adam_arthur 4mo agoThere is no way to prove, but what is definitely true is that many people are attempting to use LLMs on forums and otherwise. So if you think none of these comments are written by LLMs, you're probably mistaken too. In the end we accept that we can't tell anymore and move on (barring some biometric protocol that can't be gamed via automation)
- trollbridge 4mo agoComments at 1,000 TPS is a terrifying future.
- deleted 4mo ago[deleted]
- 0xbadcafebee 4mo agoI prefer a thousand smart AI comments to a thousand dumb human comments
- wartywhoa23 4mo agoWell, you can just vibecode a complete AI echochamber version of HN!
- deleted 4mo ago[deleted]
- m00dy 4mo agoboom!
- atemerev 4mo agoI test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.
- 0cf8612b2e1e 4mo agoWhich ones fail?
- navigate8310 4mo agoDeepkseek
- atemerev 4mo agoI tested DeepSeek V4 Pro, Qwen 3.6 Max, Qwen 3.7, Kimi K2.6, MiniMax M2.7 - they all fail to answer. Curiously, MiniMax M3 answers correctly.
- nkmnz 4mo agoNo idea why you've been downvoted. This is excellent news.
- paulinho1 4mo agoBecause this never gets brought up about US models, which have just as much censorship as the Chinese ones.
- storus 4mo agoNo, US models have alignment. Only Chinese models have censorship.
- happyopossum 4mo agoPlease educate us - which accurate and provable events in history are censored by US based LLMs as part of a government enforced reeducation campaign?
- slopinthebag 4mo agoI hope this is the next frontier AI labs push. Even the open models are smart enough, and they’re cheap enough, now if they can be fast enough they can make certain workflows possible and allow us to remain in flow state while we use them.
- elar_verole 4mo agoYeah, this seems to be the easiest path for overall agents efficiency in the short term
- minraws 4mo agoAssuming they mean 8xA100 or similar, that's some rather insane performance, and at just 3x the cost, it still quite cheap-ish. With some optimisations this might be quite interesting. I think the margins are getting quite compressed with this one, since it isn't included in token plan and the actual costs increase are much higher than just 3x. But still fairly decent.
- throwa356262 4mo agoSuspect this will be included once out of beta but at a higher credit/token ratio. Remember, these guys are not VC backed. Anything they do must break even
- JayStavis 4mo ago> must break even Understand the spirit of this, but probably not true. I don't think Xiaomi, or any big tech company, needs to break even on their new model releases.
- varispeed 4mo agoChinese "companies" are not companies in the western sense, but more like government departments with capitalist styling to deceive the western audience. From that point of view, they have as much money as they need. That's why there is no "VC", because Chinese government assumes that role.
- throwaway67678 4mo agoHuge L for free market economies if true
- Qdulf 4mo agoMust be Blackwell for native fp4 support.
- deleted 4mo ago[deleted]
- maxloh 4mo agoThe generation speed in the demo video is crazy, to say the least, and completely beyond my impressions of LLMs. The Xiaomi team really brought something to the table.
- ilaksh 4mo agoI think these type of demo videos should allow people to get a sense of super intelligence. Because it's very hard to imagine something that is say three times as smart as you -- by definition you wouldn't be able to comprehend it's thoughts -- but this shows clearly what something that can think 100 times faster than you is like.
- deleted 4mo ago[deleted]
- npn 4mo agoHow? edit: now I read the article fully, seems like they utilize some very effective MTP algorithm. and somehow the quality is still decent enough. though, I doubt that the quality really only drip a bit like they claimed. maybe for the benchmarks, but for general uses the heavily quantized models very often so worse result.
- lostmsu 4mo agoThey say they are using https://github.com/tile-ai/TileRT https://github.com/tile-ai/TileRT - persistent CUDA kernel - tiled processing with overlapping read/writes - model designed with specific constraints in mind
- aitchnyu 4mo agoExcuse me, do aliens live among us? 17 commits, 99% Python and multiplying the speed of GLM, Deepseek V4, MiMO 2.5?
- zander_jiang 4mo agotilert is closed source, the repo is just a python wrapper that invokes the binary.
- 2001zhaozhao 4mo agoi wonder if it will be possible to hardcode a model with some kind of MTP-adjacent algorithm to use a smaller portion of it to generate most of the tokens but route to the real experts every once in a while to steer it towards good thinking directions. (Perhaps this is done only when it's generating its thinking block, and the training takes it into account) Could result in very high efficiency and still good intelligence without having to resort to fundamental adjustments like going to a diffusion LLM
- npn 4mo agoI doubt you can do that. MTP magic happens because for texts, we have a lot of low value fixed tokens that almost always get generated in the sequence (like punctuation, function words, language keywords etc). for most important ones (the entities, the content words, variables) you still need the full model. so there is alwasy a maximum limit for how well MTP can do.
- moffkalast 4mo ago42B active params, sliding window attention. There's your tradeoff.
- vlovich123 4mo agoSliding window for the draft model, not for the main. 42B for active params because it’s a sparse MoE which is a common technique for the larger models to not get bottlenecked by memory bandwidth.
- moffkalast 4mo agoSeems to be for both according to the spec [0], maybe it's wrong though. 128 sounds really tiny, I wonder if they mean some kind of blocks? [0] https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash#4-model-summary https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash#4...
- E-Reverance 4mo agoNo > It uses 384 routed experts (top-8) with hybrid attention (full-attention + sliding-window 128 at 6:1 ratio) over 70 layers (1 dense + 69 MoE) https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.5-Pro https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.5-Pro
- bearjaws 4mo agoGiven how "smart" some of the 26b dense models are now, I would not be surprised to see a strong 40b MoE.
- irthomasthomas 4mo agoI don't understand, given all they say, why this would not be made available to everyone at once? Why the limited release? They should have no trouble scaling it if it runs on a single rack.
- gekoxyz 4mo agoMaybe they don't have enough racks. The news indicate that China isn't in a really good situation with GPUs, so probably they want to keep most of them for other stuff. Also because since the price is so cheap they probably want to use the other GPUs for stuff that has higher margins.
- jdthedisciple 4mo agoBecause presumably then it won't be 1000 t/s for everyone anymore given hardware limitations?
- HarHarVeryFunny 4mo agoMaybe they only have a finite number of racks ;-)
- slaw 4mo agoChinese companies are blocked from buying modern ASML lithography machines. The most modern scanner China is still allowed to buy is NXT:1980i from 2015.
- boutell 4mo agoI wonder about this too. The other objections miss the point: if it's faster, and otherwise the same, and doesn't require different hardware, then why not just announce that the standard tier of MiMo-v.25-Pro is now ridiculously fast and raise the price? What does "limited high speed resources" mean if it runs on the same hardware as the rest of their pool? I think the answer is that there's a tradeoff here where additional throughput for a single person can be achieved only by tying up more resources than a normal request would, even when you take into account the fact that the normal request takes longer to finish. I'm not an expert, but some of the optimizations they describe, particularly the parallel prediction stuff, sound like they could take up extra resources.
- maxothex 4mo ago[flagged]
- kingstnap 4mo agoGiven that MiMo is as cheap as Deepseek ( previous discussion: https://news.ycombinator.com/item?id=48282814 https://news.ycombinator.com/item?id=48282814 ) multiplying that by 3x for ultra speed is still shockingly cheap.
- miroljub 4mo agoMiMo and DeepSeek are not cheap. Anthropic and OpenAI are expensive for what they provide.
- ignoramous 4mo agoThe Chinese "Neijuan" is real & well reported: https://www.reuters.com/business/autos-transportation/what-is-involution-chinas-race-to-the-bottom-competition-trend-2025-09-14/ https://www.reuters.com/business/autos-transportation/what-i... It is another thing the BigLabs accuse open weight models of benefiting from distillation & other techniques & essentially avoid higher training costs (which typically bleed into bills end users pay for inference). Ex A: https://www.anthropic.com/research/2028-ai-leadership https://www.anthropic.com/research/2028-ai-leadership Ex B: https://www.reuters.com/world/china/openai-accuses-deepseek-distilling-us-models-gain-advantage-bloomberg-news-2026-02-12/ https://www.reuters.com/world/china/openai-accuses-deepseek-...
- flexagoon 4mo agoTrue, but why would end users care about that? If anything, training on synthetic AI output is more ethical than on scraped human works (of course, not to say the Chinese labs aren't doing the latter)
- trollbridge 4mo agoWe buy cheap Chinese goods all the time. Absolutely nothing wrong with that. In this case, at least it’s threatening multimillion dollar salary jobs instead of entire towns of working class people in America or Mexico. And the Chinese labs actually release their weights. You could call it… open AI.
- serpix 4mo agoI may sound like a shill, but exponential growth and all. We are going to get near instant software from prompt, multiple ones and then choose the best one. Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.
- 9cb14c1ec0 4mo agoAnyone remember the old days when a new frontend framework came out every 3 months. That has pretty much stopped. No one cares anymore.
- mountainriver 4mo agoIt’s even discouraged now as LLMs wouldn’t have the documentation built in
- osti 4mo agoBut I think the eventual goal is that documentations won't even be needed. LLM should just itself understand the nuances of frameworks by analyzing their codebase.
- LASR 4mo agoOh you wait until LLMs come up with frameworks that allow multiple LLMs to collaborate effectively. Then you’ll have new frameworks every 3 days.
- asveikau 4mo ago> when a new frontend framework came out every 3 months. > No one cares anymore. I never cared about this. I think this captures something that I've been searching for the words for. (Maybe I should have gotten an LLM to write the words for me.) Some of the biggest AI boosters are the kind of dev that would have cared about the new frameworks of the last 3 months. They had a "the framework does all the thinking for me" attitude already, so it is easy for AI to slot into that.
- 4mo ago
- amunozo 4mo agoThese price and speed optimization from Chinese providers, combined with the raising prices from American ones will change the game sooner than later. Many companies are finding issues with the AI bills already.
- throwaway894345 4mo agoI wonder what are the economics driving these pricing decisions? Are the Chinese companies just subsidizing their models to a greater degree than the US, or is this an emergent property of energy policy between countries?
- Octoth0rpe 4mo agoThrowing out another factor: Chinese companies have been banned and/or limited from buying nvidia, and turned to local companies for their hardware. I haven't actually seen pricing/benchmarks comparing Chinese AI accelerators, but it wouldn't surprise me if that also worked out in their favor as well.
- lokar 4mo agoAnd, possibly, state subsidies at every level.
- Schlagbohrer 4mo agoI have to point out the massive state subsidies in the united states for the tech companies and datacenter builders.
- throwaway67678 4mo agoLower cost of labor, lots of under the hood optimizations (e.g. cache hits for DS), many of these companies have existing infra (fewer upfront costs for deployment), etc
- ecshafer 4mo ago
- scosman 4mo agoCerebras is trialing Kimi K2.6 at 3000t/s (invite only). I'm excited for when the fast hardware gets more mainstream for frontier models. Models designed for speed on Nvidia are nice addition that could bridge the gap.
- lostmsu 4mo agoCerebras currently does not provide any discounts for prefix caching making its use for agentic workloads sqr(n_turns) more expensive.
- michael-ax 4mo agonow that's what i call a software development breakthrough/platform! thanks for the heads up!
- adrian_b 4mo agoTFA mentions that until now special very expensive hardware like Cerebras was required for reaching this kind of speeds, and it emphasizes that what is novel in their results is that they have obtained over 1000 token/s for a model with over 1 T parameters by using just standard hardware, i.e. one server with 8 GPUs.
- btian 4mo agoSource? Their website says 1000t/s https://www.cerebras.ai/blog/which-is-faster-gemini-3-5-flash-or-kimi-k2-6-on-cerebras https://www.cerebras.ai/blog/which-is-faster-gemini-3-5-flas...
- scosman 4mo agoThis is likely correct, sorry for the bad info. Was working from memory.
- johndough 4mo agoCerebras got lucky that they IPOed last month instead of now.
- GaggiX 4mo agoIf MiMo v2.5 Pro can run at >1000tk/s on GPUs then I will soon expect the same from OpenAI/Anthropic/Google.
- 59nadir 4mo agoI wouldn't expect any of the american labs to be particularly great (or have much desire) to work on efficiency, they've been consistently proven to be uninterested (if not incapable) of actually improving on those types of things. The closest we've seen lately is that maybe GPT-5.5 (and Opus 4.{7,8}?) are more token-efficient, i.e. they solve things with less tokens...? It hasn't been coupled with any other kind of efficiency bump, though, and we're seeing higher costs anyway in most places where the american labs are involved. The only players that seem to be capable of a consistent pattern of doing more with less currency are the chinese labs.
- holoduke 4mo agoSpeed is indeed a next big thing what should happen with LLM frontier models. The possibilities with current models but 1000 times faster would be super useful. Earlier this week it took Claude at least full time a week with two max subscriptions to solve a complex issue where we wanted to mimic a occlusion mapping variant used in the game Crimson Desert. Pretty complex mathematical challenge. With a ultra fast LLM and a proper self verification process it would be awesome.
- astlouis44 4mo agoInteresting. For your occlusion mapping variant, what engine is the game you're making with made with that you're implementing this for? Do you have Claude hooked up to Unity or Unreal?
- MaxikCZ 4mo agoId also be interested in more details as sibling comment. I find that when I try to build stuff, its like building skyscraper from straw. What methods are moving you forward the most?
- __natty__ 4mo agoWith this at 1k tps and Kimi 2.6 1k tps by Cerebras, I believe we are entering the next stage of LLMs, where companies will also compete on throughput
- FastAnchor 4mo ago[dead]
- qsera 4mo agoTokens per seconds is the "Megapixels" of AI marketing!
- Octoth0rpe 4mo agoI mean, sure, in the sense that they're a real and meaningful number for most of the spectrum on offer, and only gets silly when the number gets too high? There's a pretty big usability difference between 10t/s and 100t/s, and I can imagine similarly for 100->1000. I don't know about > 1000, but let's not pretend that the number is meaningless.
- qsera 4mo agoIt is pretty meaningless for something that calls itself intelligent.
- orbital-decay 4mo agoDefinitely not, there's a ton of potential realtime use cases and high throughput/low TTFT is exactly what they need.
- qsera 4mo agoOf course, megapixels are also useful if you want to print large sizes.
- orbital-decay 4mo agoCompletely incomparable. Large printing is a narrow niche in art and technical photography, part of which is already covered by composites, and pixel size is a physical tradeoff for sensors. Cases for reasoning at realtime speeds are much, much more diverse, infinitely more diverse than anything we're currently using the big models for. Consider the fact that large models don't necessarily imply language. Speed is the major limiting factor for high-level automation. Coding is simply the immediate killer app that is useful right now, given the current state of AI - just like roleplaying and chatbots were previously.
- harel 4mo agoA few things in life I can't fully grasp why they are so sought after. One is that constant need to exhibit growth. As if being massive and staying as massive is not good enough, one has to always and continuously grow. The other is constant speed increases. We're already operating at 50x speed. My output is much wider and so much faster, I am sometimes my own bottleneck. And now as if that is not enough we want more speed. "I want a full software product from scratch in 12 seconds, Because 5 minute is too long and I got things to do..." Really?
- philipkglass 4mo agoI remember when I had to wait minutes to get a high resolution image over a dialup connection. When computer and communications hardware advanced enough that I could get 30 high resolution images every second, there were brand new uses. In the case of LLMs, I could imagine that much faster operations allow you to introduce them as parts of systems that need to react to the real world at high speed, like factory equipment. Showing that a model can do the usual LLM tasks at extremely high speed is just a demo proving that the approach works.
- harel 4mo agoThe example in the video was a generation of a dashboard app of some sort. I can do that with a "normal speed" Claude in a few minutes. The difference is a few minutes. This is compared to a few weeks in old school development time. I don't have a problem with taking it a little "slow" (as in - few minutes) and lending my thought to it rather than just going for fast generation and who knows what's inside. I get your use case, but this is a specialised one, and not the one 90% of people will think of - everyone want that fast app in 12 seconds... Or so it seems from me being downvoted on that comment.
- srdjanr 4mo agoI frequently tell agent to do something, wait ~10 min (which is just enough that I can't/don't want to start anything else), ask it to change something, wait a few minutes again, and so on. So I'm basically idle while waiting for agent, and it would be great if it was faster. It's like your compile times were ~10 min. Sure, it's not a huge deal, but it's sooo anoying
- eli 4mo agoNeat. The frontier models have gotten pretty impressive, but they're all a bit too slow for interactive, human-in-the-loop coding. It incentivizes vibecoding and running multiple agents in parallel. A fast agent feels more like a partner. For a while I was running Cerebras GLM 4.7 for a bunch of tasks. Not a very smart model, but it's fantastic to be have a live prototype of a site up and be able to type "make the fonts bigger. No not that big" and see it change in real time. And MiMo 2.5 is a lot more capable than GLM 4.7.
- ignoramous 4mo ago> And MiMo 2.5 is a lot more capable than GLM 4.7 MiMo 2.5 is not the same model as MiMo 2.5 Pro. GLM 5.1 is z.ai's lastest iteration & is one of the popular open weight coding models. If you've had the chance, how does GLM 5.1 (which is now more expensive than MiMo 2.5 Pro after its recent 70% price drop) compare?
- eli 4mo agoGLM 5.1 is very good. Definitely a contender for best open weight coding model. Nothing like 4.7. But quite a bit more expensive than MiMo 2.5 Pro. Like 5x to 10x more on my little tests, at least by the API rates.
- maxdo 4mo agoi tried glm 4.7 for agents that write code. simple scripts 200-1000 LOC. extremely bad . Had to abandon cerebras oferning, their smart models are only on enterprise plan.
- jona-f 4mo agoglm 4.7 is quite old by now. I don't even use 5.1 anymore, cause I found kimi k2.6, mimi 2.5 pro, deepseek v4 pro and qwen 3.7 all better than glm 5.1
- goyozi 4mo agoFast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and regret it later. On the other, AI that takes a few seconds to max few minutes to solve what used to take hours or days? That’s a game changer and I don’t even know where we fit in.
- ipkstef 4mo agoasking for curiosities sake. What kind of PR loop are you running that takes a few hours?
- ketzo 4mo agonot OP but usually for me this means long verification loop; waiting 10min on CI checks, that kind of thing, rather than actual 1hr wall clock of token generation
- devmor 4mo agoOr slow MCP servers that are waiting on HTTP calls from APIs, playwright/other UI instrumentation, etc.
- RussianCow 4mo agoBut those things won't be sped up by a faster LLM, so I feel like that's not what the OP is talking about.
- goyozi 4mo agoWell, I used an extreme example. OTOH, I’ve done quite a few of those „fix CI” or „migrate X” prompts recently and while there is a fixed component like running CI / builds, I’d say the LLM time is still around or above 50%, especially at the beginning of the project. Then there’s also regular tasks that now take minutes per message which completely get me out of the zone. I imagine iterating on those in near real time would be a big change.
- h14h 4mo agoThe gated "ultra-speed" phenomenon seen here and with the Cerebras Kimi K2.6 release, while understandable, is somewhat troubling IMO. Getting ~1000 TPS on near-frontier intelligence is a step change, and enables whole new use-cases for applications. Seeing limited compute resources beget selective access makes me worry for the future of competition.
- trilogic 4mo agoPfff time wasting. 1 password between 8-16 characters, and this and that... What??? 2 Captcha after captcha, come on 3 Service unavailable This service is not available in your region yet. Are you kidding me. Come back when you are ready for the users. I was hopping to try it, what a frustration.
- prplfsh 4mo agoThis will be really powerful for voice. Being able to reason makes LLM so much smarter but with voice your latency budget is so tight that you can't spare the time typically.
- jeffrallen 4mo agoThis is true for humans too. Lol
- pullshark91 4mo agoIt's interesting but not game-changing IMO. Speed here is not a bottleneck.
- gertlabs 4mo agoMiMo V2.5 Pro (regular speed) remains the strongest open weights agentic coding model we've tested -- it's been interesting to see how little attention it has received relative to some lower performing releases. And the "fast mode" pricing is very competitive here. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- unrvl22 4mo agowhy is deepseek v4 pro a lot lower than flash? where is mimo 2.5?
- deleted 4mo ago[deleted]
- gertlabs 4mo agoDeepSeek v4 Pro struggles with a custom harness, and all the models ranked above it don't, so it gets downweighted in the agentic coding benchmarks (although it ranks better than Flash in one-shot problem solving: https://gertlabs.com/rankings?ow=1&mode=oneshot_coding https://gertlabs.com/rankings?ow=1&mode=oneshot_coding). We ran plenty of samples. MiMo v2.5 is on there, as well as the pro version. We found a few anomalies in our evaluations, which makes sense -- if every new sub-release is better across the board in every area of the model card, that should raise alarms about benchmaxxing. But the main thing we found is that hype != performance, and I trust our benchmark methodology significantly more than the model cards the labs add to their press releases.
- digdugdirk 4mo agoCan you explain more about how it struggles? I haven't noticed any issues in my usage, so I'm just curious what is meant by this.
- gertlabs 4mo agoIt's likely overfit to common harnesses and iteration patterns, so it struggles with formatting tool calls and json in our testing which use our own harnesses (although there is a lot of overlap with tools that would be found in any coding harness like bash, apply_patch, etc.) We didn't love the results because it draws negative scrutiny to our benchmark, but the results are real and done at scale and I think DeepSeek V4 Pro's inability to do agentic work outside of environments it was trained on is an important thing to measure, especially when so many other models can generalize to new environments just fine. Google models also struggle with tools, but they have very strong initial answers, so there is more potential for them to bridge the gap with some better post-training.
- isusmelj 4mo agoNo note about the specific GPU they use. One might speculate. B200? H200? H100?
- PhunkyPhil 4mo agoObligatory taalas mention: https://taalas.com/ https://taalas.com/ Despite the performative UI components they have a shipped (demo) product: https://chatjimmy.ai/ https://chatjimmy.ai/ This is only 3.1 8B and a very small context window, but at 17k tokens per second it's likely enough to reliably call tools which would make a huge difference in agentic applications. Assuming they can bake in better models I'm just as bullish or even moreso on this, considering this opens up edge computing at the extremely low power requirement. High tok/s is the future IMO.
- desireco42 4mo agoI didn't use their pro speed but regular Mimo-v2.5, not even pro, it seems really fast. I have plenty of tokens and subscriptions but this is really impressive. I really don't need another one, but I am tempted simple because it works so fast, can't imagine how this fast service can be.
- dakiol 4mo agoSo, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot machine with the hope of it giving you the right answer with the right prompt. So, if any, I would say it's worse for us. Obviously, it's the completely opposite situation for corporations and executives: they are loving the AI situation so much!
- fullstop 4mo agoIt's making things less fun, for me at least.
- linsomniac 4mo agoOdd, I'm having the opposite experience. The thing I really love about working with computers is when I achieve something. That's the thing that makes me figuratively, and sometimes literally, throw my fists into the air and go "Yeaaah!" With the AI tooling, I'm getting those more like a couple times a week. Plus, I'm using AI to attack the things in my day that are "a drag", and getting them done too. The highs are more frequent and the lows are not so low.
- fullstop 4mo agoOh, sure, I can make things with it. But I have an extraordinarily hard time saying that I made something. It feels like it cheapens the whole thing. Maybe I'm just old, because I remember people saying the same thing about code completion in Visual Studio back in the late 90s. This is so much more than code completion, though.
- dd8601fn 4mo agoExactly how I feel. I didn’t make a damn thing. I essentially asked a chatbot to. Did I ask for better things with some important concepts pre-rolled? Yeah, of course. But that’s so, so much less interesting than having actually made a thing. I try to remind myself that the output of my projects have nothing to do with who I am, but the honest truth is they always mattered to me. Now that’s dead, and it’s never coming back. It ain’t exactly existential dread, but it is something I’ve lost.
- jbellis 4mo agoit is hard to understand what the actually meaningful innovations are here / what TileRT is bringing to the table. - dflash: new-ish but February is ancient by the standards of the pace of AI innovation lately, I guess applying it to a 1T model is new-ish in the sense that the dflash researchers don't have the hw budget to prove that out - persistent engine kernel: this is like CUDA 101 - warp specialization: I think this just means "keep different gpu resources all busy w/ pipelining" which is CUDA 201, some of it is even baked into pytorch now - MXFP4 QAT: not new - TileRT: hard to tell what this actually does, there's a PyPi wheel with support for DS 3.2 and GLM 5 but binary only
- zander_jiang 4mo agotilert is a highly optimized megakernel, its a single kernel that does the entire decode pass, this enables overlapping weight loading with computation, eliminates cuda launch overhead (CUDA graph does not, contrary to what most people think), allows for more fine-grained pipelining. There're lots of blogs/papers on it. Its currently the best approach to maximize memory bandwidth. But megakernels are incredibly hard to optimize, and only work for small batch sizes (low throughput, hence high price), thats why we don't see them much in production.
- GodelNumbering 4mo agoBelow is the part I found most interesting > "However, naively applying FP4 across the entire model causes degradation in complex reasoning, logic, and code generation. Given the MoE (Mixture of Experts) architecture of Xiaomi MiMo-V2.5-Pro — where Experts constitute the vast majority of parameters and exhibit the highest tolerance to quantization — we selectively quantize only the MoE Experts to FP4 while preserving original precision for all other modules. Through FP4 QAT (Quantization-Aware Training), we dramatically reduce model size and maximize hardware bandwidth utilization while keeping the model's overall capability essentially on par with the original, as shown below"
- buildbot 4mo agoThe 120B and 20B GPT-OSS models by OpenAI did this last year for what it’s worth; the MoEs where MXFP4
- 0xbadcafebee 4mo agoThis is the value prop of Groq and Cerebras. They don't have the best models, but they have the fastest inference, and Groq has both the lowest cost and fastest speed.
- pants2 4mo agoWith a tps and a token price you can calculate approx. price per hour of running the model! $2.61/M tokens * 1,000 tok/s = $9.40/hr That would be pretty cheap for an 8-GPU node which would typically run around $45/hr or more. Guess this depends on how many parallel streams it can handle.
- aplomb1026 4mo ago[flagged]
- jingpostmedia 4mo ago[flagged]
- wartywhoa23 4mo agoAn exercise for the near future: Albert has a chalet in swiss alps and an uncles' fortune, burning tokens at 11 kHz. Joe has a rental capsule and a UBI, burning equally priced tokens at 23kHz. Who's the first to solve the problem of maniacs in power?
- aburayhanalif 4mo agoit is good i think
- _pdp_ 4mo agoDo you know what will be cool? It will be cool to measure models based on their RAW performance and measure them in terms of ROI - not some benchmark but something meaningful like we used this model to solve X. That will be a massive mind shift and might justify the token expenditure.
- HDBaseT 4mo agoAren't benchmarks exactly that? We used the AI to solve given problem with x% adherence/quality/correctness?
- siddbudd 4mo agoto try the demo you need to sign up. why? to sign up you need a password 8-16 chars. Why limit at 16? geez, I hate Chinese IT companies with a passion. update: AFTER signing up, and only then, am I told: 'This service is not available in your region yet.'
- HerShin5 4mo ago[dead]
- overgard 4mo agoPretty cool, although I can't help but think this would be a very easy to way rack up a GARGANTUAN bill. That company that blew 500 million on Claude in a month might have competition soon..
- sheeshkebab 4mo agoOpus regularly bitches and wines to me how long something will take and that I should think before asking it to do it. But then it does it anyway in 15 minutes.
- temikus 4mo agoI’ve personally found MiMo models a hit and miss. I have some personal agentic projects and I found them to hallucinate hard at least 10% of the time. And do so in pretty sinister ways - making up people, names, places, etc. I switched back to Kimi for now.
- RachelF 4mo agoI wonder how fast it performs on just a CPU? If the model performs say 10x on a GPU cluster, would it also perform faster on a CPU? This could bring proper desktop AI to the average laptop user, which could be a game changer for running local models.
- mrwaffle 4mo agoWhat a ripoff you have to make an account then 'apply' to try this demo.
- digitaltrees 4mo agoAm I the only one that doesn’t care about speed? I want it to not do stupid stuff and to be cheaper.
- Npovview 4mo agoGenerally thinking tokens are the ones which are verbose. So the speed helps with reducing time for thinking tokens generations and you get your actual output code very fast.
- 59nadir 4mo agoI prefer faster, dumber models because I provide the intelligence myself and I use them only for things that can be verified pretty easily; they do research (with sources) for me, do certain types of code analysis and code search, boilerplate generation, etc., so a fast model is really key. I don't have any desire (or think it's a good use of LLMs) to one-shot features because even SotA models are incredibly bad at this. I'm optimizing for what they actually seem to be able to do reliably and pretty well, and I want those things to be done fast so I can get on with things.
- digitaltrees 4mo agoFair point and good counter argument. Too bad jimmychat.ai doesn’t have api access anymore.
- kopirgan 4mo agoWill this list for trillion dollar valuation as well?
- Frannky 4mo agoI tried this model it was pretty bad at coding. Maybe it was me. 1k tokens/sec pretty cool tho. Deepseek V4 pro is better. I wonder tweak pi + deepseek pro v4+ 1k tokens/sec if would actually be better than Claude code
- LoganDark 4mo agoI was just playing with Cerebras a few days ago because it's the fastest inference provider by far. Unfortunately, the only model anywhere near economical to run that fast is gpt-120b-oss which sucks at Pi's tool calling. So I've been hoping for something faster ever since, especially since my local hardware has a paltry 128GB of unified memory. Hopefully this pans out and fast models (that are also not ridiculously dumb) become the norm. It's amazing what you can unlock with even a single order of magnitude's speed improvement.
- bryabaek 4mo agoi tried to test it and after logging in, i get "You don't have access to this event trial" and can't even log out until i clear my cookies. despite having good model, why such a bad website?
- girvo 4mo agoSame. I also found out that my old Xiaomi account is apparently considered "mainland china" and I can't put any phone number except a chinese one on it lol. I'm not trusting these people with anything that's for sure, useless. I'm australian and have never been to china in my life!
- yanhangyhy 4mo agohave anyone give it a try? even in china, it's not popular...but xiaomi is really good at make price go down on everything...
- PhilippGille 4mo agoThe interesting bits on how they achieved it: > On the model side, we applied FP4 quantization > introduced DFlash, an efficient speculative decoding method based on block-level masked parallel prediction > On the system side, TileRT perfectly adapts to the dynamic characteristics of these algorithms > 1000+ tokens/s output [...] using just a single standard 8-GPU commodity node
- ljlolel 4mo agoCan try it now in seconds on https://trustedrouter.com/ https://trustedrouter.com/
- zero0529 4mo agoCool, what is the price pr. Million token. I am using a 300 t/s model for a project I am doing and speed is crucial over precision, so this seems like an upgrade. However if it is 10$ pr. M tokens then it is not worth an upgrade.
- GaggiX 4mo ago$0.435/$0.87 for the standard speed, this one should be 3 times that.
- Yatsui 4mo ago[flagged]
- adithyaharish 4mo ago[flagged]
- megous 4mo agoThis just means you can blow through monthly budget in 1h instead of in 4h on the cheapest plan. :)
- trollbridge 4mo agoIf you didn't apply already, you should - they turned around my application in a day. This thing is seriously fast and was good enough to switch it in for the other model I was using. I tried it for both planning, executing, and subagent tasks and it performed adequately in all 3. So, this is another one to add to the list next to DeepSeek-V4-Pro and Qwen-3.7-Max...
- linzhangrun 4mo agoMaybe it is very suitable for some scenarios like autonomous driving, where reaction speed matters a lot, if they can find a way to put the hardware into a car at an acceptable cost. Maybe this is not impossible, because the hardware cost of current high-end assisted driving is already quite high? At present, intelligent driving still feels, in general, like a beginner driver who drives mainly by reaction. FSD is a little better. But it still lacks the kind of “spirit” human drivers have. How to say it: when a human driver sees the car in front shaking left and right, he can guess that the driver may not be fully conscious, and then keep away from it. Current assisted driving systems are still quite weak in this kind of understanding of the world. The most important thing in driving is prediction. But driving itself does not need very deep or very complicated reasoning. Recently I tried using Mimo for development, and I believe the understanding ability it can provide is absolutely more than enough for driving scenarios. Sadly, the Pro version does not have multimodal ability. And this US version seems to be trying to solve the biggest problem of using LLMs in control systems: latency. Xiaomi’s car is good, but its assisted driving level is near the bottom in the same class. Compared with new EV makers, its route is quite “traditional”, just like comparing lap times with Porsche at the Nürburgring. Xiaomi’s large model team may change this.