4 ms·
Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
by gpugreg 13d ago
Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
- beastman82 13d agocan confirm. I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
- tomega2134 13d agoIs a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
- mathisfun123 13d agosame reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn't immune from vanity).
- throwaway27448 13d agoIf you seriously think apple products are nothing but a status item, you're deluding yourself and probably have been for decades.
- deleted 13d ago[deleted]
- _hugerobots_ 13d agoThis 1000%. Data centres don't equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
- bigyabai 13d agoIt's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
- _hugerobots_ 13d agoSo if it isn't a comparative ability, now it's a power cost issue? This reads like goal post moving.
- mathisfun123 13d agoIf you seriously think apple cares about anything other than cell phones, you're deluding yourself and probably have been for decades.
- tom_ 13d agoThey've been selling phones for less than 20 years at this point? Though I suppose 1.9 is not equal to 1, so it gets the plural.
- throwaway27448 13d ago...did you mean profit? I don't think they're manufacturing iphones just on the hope they delight you. This is also true of Google et al. I don't get these weird parasocial emotional attachments/beefs people have with brands. Talk to a therapist.
- mathisfun123 13d agobrother my point is they don't care about their product offerings outside of their phones. this post/thread is about one of their product offerings which is not a phone which is inferior to their competitors'. simple.
- selectodude 13d agoMy M1 Pro MBP is 6 years old and continues to be the best computer I own, so if that’s Apple not trying, god help everybody else once they do.
- mathisfun123 13d agoJesus Christ reading comprehension has completely fallen off a cliff - we're talking specifically about GPU performance here. Do you understand the words that are coming out of my mouth?
- deleted 13d ago[deleted]
- nacs 13d agoPeople don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.
- ProllyInfamous 13d agoExactly; when I first got my RTX 5070 Ti (16gb, to game with!!!, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models <8gb, primarily llama3.1. My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster." This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
- selectodude 13d agoOnce I get some kind of settlement after getting beaten up by a cop my first purchase will be some RTX Pro 6000s.
- throwaway27448 13d agoA) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.
- bigyabai 13d ago> for me at least a GPU is completely useless for anything but being a token generator. No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.
- throwaway27448 13d ago> No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland. Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games so a GPU would probably be better off in some kid's computer.
- Eisenstein 13d agoA 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
- beastman82 13d agoOff the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
- deleted 13d ago[deleted]
- medvezhenok 13d agoI think you’re missing that MTP can predict more than 1 token in advance.
- _hugerobots_ 13d agoHave a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work. The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.
- cyanydeez 13d agoQwen3.8-Flash-Next is pretty damn worth the extra ram you need.
- liuliu 13d agoBoth are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.
- ActorNightly 13d ago[flagged]
- bee_rider 13d agoI guess it is possible, but Apple has had very vocal fans for decades. I suspect, rather than astroturfing, it is just people who are in their ecosystem.
- abletonlive 13d agoTok/sec is 0 on a 3090 for most of the models that the mac can run
- ActorNightly 13d agoRunning very large models on Mac is unusable at 10 tok/sec. You get more average inference over the day using free Google Gemini. And for the price of a Mac that can run a large model, you can get 2 3090s humming along running a small model so fast that it can simulate a lot of the behavior in large models just through sheer number of context it generates. For example, editing code means that by the time your large model on your Mac is finished writing a file, the smaller models have generated the code, written the code to file, ran it, and debugged any issues. So given that, which one of these is true about you? 1. You are paid by Apple to push marketing on HN 2. You are a hardcore Apple fanboy and just think that owning a Mac studio is a flex
- searealist 13d ago... or with llama.cpp with MTP.
- nuccy 13d agoHonest question (as a person outside of AI industry): I understand the title "M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents" requires ... but if those LLM benchmarks are so settings-sensitive, why not to dedicate at least half of the article to some general benchmarks? Maybe ultra-large matrix/tensor/N-body calculations, some memory bandwidth measurements, which all contribute to LLM performance, but which allow to compare more-or-less apples to apples for those different devices (M3 Ultra, M5 Ultra, RTX 5090, et al.)
- GeekyBear 13d agoThe issue is that the moment you want to run the more capable models that won't fit in a single 5090's memory, performance falls off a cliff.