6 ms·
Performance per dollar is getting faster and cheaper
- minraws 3mo agoCan you folks add performance per watt as a metric to these comparisons, I honestly want to understand where AMD fits in the stack in terms of actual performance to dollars. I have had talks with companies wanting to build data centers outside of US and find it hard to source anything Nvidia in sufficient capacity and scale. If AMD is competitive performance per watt and roughly reliable in terms of software support which is what most folks outside of US prioritize above all else, since outside of China and US electricity tends to at a relative premium. Maybe if they make smaller data centers viable at the right price, AMD could be part of the stack outside of US where ever Nvidia is more limited in supply. Though I have genuinely no idea what sourcing an AMD GPU looks like. I have never seen a company use AMD outside of wafer and a couple others mostly in US. Genuinely intriguing or maybe not really (could be this stuff is common knowledge) and I am just stuck in my Nvidia bubble here.
- craftkiller 3mo ago> I have never seen a company use AMD Meta is using AMD: https://www.amd.com/en/newsroom/press-releases/2026-2-24-amd-and-meta-announce-expanded-strategic-partnersh.html https://www.amd.com/en/newsroom/press-releases/2026-2-24-amd... And OpenAI: https://www.amd.com/en/newsroom/press-releases/2025-10-6-amd-and-openai-announce-strategic-partnership-to-d.html https://www.amd.com/en/newsroom/press-releases/2025-10-6-amd...
- Schiendelman 3mo agoIt's not clear when this will be - AMD has slipped these dates likely to 2027.
- minraws 3mo agoOpenAI maybe, but a few friends in Meta said they don't so dunno man. Seems sus atm. But it's meta they can get a GW up of AMD in a year
- Twirrim 3mo ago> I have never seen a company use AMD outside of wafer and a couple others mostly in US. There's a few using them, and even more starting to experiment with them. AMD has long been a source of disappointment around this side of things, so I'm hesitant to feel optimistic we'll finally get some competition. The market really needs viable competition to Nvidia, especially performance/watt.
- deleted 3mo ago[deleted]
- technoabsurdist 3mo agoAMD MI355X uses 1,400W per GPU and NVIDIA B200 uses 1,200W. So AMD uses about 16% more power.
- vlovich123 3mo agoNot how you measure performance per watt but generally it’s 20-60% worse at tok/s/watt not 16. It does have ~50% more memory (~100gb) which complicates the comparison.
- kingstnap 3mo agoA DGX B200 costs like ~$0.5 M and uses around 14 kW. If you plan to run it straight for 8 years 100% max usage thats around 1 GWhr. A gigawatt hour is a lot of energy but its not that much compared to the price of the actual machine. In Germany for example with its expensive energy thats about €100k worth, which spread over 8 years is pretty minor compared to the up front half mill. The real issue with high power consumption is not really the cost of energy but the limited powersupply you can get for a datacenter. A more efficient setup is highly desirable because it means you can fit more in the limited power hookup.
- dannyw 3mo agoIt’s more than power supply. Cooling and ventilation becomes a MUCH bigger deal at rack scale, and that costs electricity too.
- thereisnospork 3mo agoCooling demand is only fractional with respect to the load: cooling 1MW of heat will only cost a few 10's to low 100's of kW, depending on the specifics. 10-20% overhead on cooling is probably a close enough estimate for napkin math.
- psychoslave 3mo agoAnd datacenters have impact on everything around them. If at the end of the day to result is a few more yachts and jets and, a lot more of miserable humans starving in ruined ecosystems, maybe that’s not the best go-to direction.
- butvacuum 3mo agoYou say they have a large impact, but having lived somewhere with some of the largest data centers- they very much don't. At least not more so then any other structure that paves over greenery. love to debate actual discission points. pull up "datacenter dfw" on google maps for mine.
- latchkey 3mo ago> I have never seen a company use AMD outside of wafer and a couple others mostly in US. Just because you haven't seen it doesn't mean it doesn't exist. We've serviced over 700 customers on our MI300x.
- deleted 3mo ago[deleted]
- 7thpower 3mo agoTypically any company that can’t get Nvidia to fill their orders will have at least some AMD.
- embedding-shape 3mo agoWhat type of company are you talking about here? Granted, nowadays I mostly interact with ML-adjacent companies but almost none would go "Hmm, hard to get nvidia hardware today, lets dump all expertise and knowledge of CUDA et al we have and start using AMD hardware until we can get nvidia", everyone would just wait or rent in the meantime.
- wongarsu 3mo agoInference workloads are usually a lot less picky about the exact hardware than model training. At least in the cases I know of the models are trained on Nvidia hardware, then exported and run on a mix of Nvidia and AMD
- minraws 3mo agoAt scale for inference it's almost non-existent for a data center company to go for AMD because they couldn't get or afford Nvidia atm. They instead start the build out and plug in stuff they can, then take a loan or ask Nvidia to help fund it. (I am not joking) I believe the case is if you can prove to Nvidia you can install and provide more Nvidia capacity they help out because more Capacity going online today is in the best interest of Nvidia. Spot prices of Nvidia GPUs going up is not good news for Nvidia btw. The people renting Nvidia has the least amount of friction in moving off Nvidia, especially with AI tools you could build and get up to speed with AMD stack much sooner... So if Nvidia is truly not an option and you entire company is not a bet on Nvidia then you will move off but only as a renter not as a buyer unless they truly can't fund Nvidia I suppose. But again I repeat if you build a datacenter and provide good enough base Nvidia will help fund you to a mostly complete data center. People might not like it but that's the reason Nvidia is so unreasonably dominant even now when otherwise given the scale of investments it might have been cheaper to look for alternatives. This is why Nvidia doesn't like the China stack.
- embedding-shape 3mo ago> I have never seen a company use AMD outside of wafer and a couple others mostly in US. Worth remembering AMD basically "owns" (not literally) the hardware-side of things in video games consoles for good many years now, with no end in sight.
- ekianjo 3mo agoBecause they have x86 CPU licenses.
- embedding-shape 3mo agoEvery single video game console of the last generation (and probably further back) are using AMD Radeon for graphics too FWIW. I think the Switch might be the only outlier recently using nvidia graphics.
- wongarsu 3mo agoConsoles used to all be custom architectures. If Intel was the only one doing x86 and AMD had offered the same price, performance and features as they do now, but in another architecture, my bet is that in that universe AMD would still have gotten the contract. Using x86 is a big deal to simplify things, but so is AMD's APU with unified memory between CPU and GPU (similar to what Apple now does with their silicon)
- duped 3mo agoAMD invented x86_64
- minraws 3mo agoI was talking in the data center gpu context, EPYCs are pretty common in data centers these days. I have a huge EPYC based data center like 200-300+km from my house on the outskirts of the city a few dozen miles from a IT industry tech park(place with lots of IT company offices).
- jingpostmedia 3mo ago[flagged]
- jingpostmedia 3mo ago[flagged]
- oDot 3mo agoDo these providers have 80+% gross margins or is something eating into them? Maybe utilization?
- technoabsurdist 3mo agohi i work at wafer. no the margins are lower averaging at about ~40%. utilization is one of the highest order bits in determining margins here, yes.
- keynha 3mo ago[dead]
- yieldcrv 3mo agoAgentic coding drivers for different architectures is a massive unlock for the world So much compute is under utilized waiting for a savant or company to prioritize an architecture, and now all the other engineers can tackle this at any time if they get inspired on the right prompts
- yogthos 3mo agoPersonally, I can't wait till something like this starts getting to consumer level. https://www.anuragk.com/blog/posts/Taalas.html https://www.anuragk.com/blog/posts/Taalas.html
- yieldcrv 3mo agoThat’s pretty fascinating, Apple has some innocuous LLMs and transformers baked into its devices and leveraging their neural chipset So I could see something like this where the neural chipset has an LLM that cant be so easily updated baked into it, until you get a new device
- yogthos 3mo agoExactly, it'd be the same as regular chip designed evolving. You get a specific model version baked into the chip, if it does what you need then it's fine. If you need more capability in the future, you just buy a new chip. I also think the dynamic would be really different if model inference can run at ridiculous speeds. You could make a genetic algorithm loop around it, so it can generate a population of proposals at each step, then have those tested and whittled down iteratively. If inference happens at thousands of tokens per second, then from user perspective it would still be really fast, and even a small model could solve complex problems.
- technoabsurdist 3mo agothis is exactly our thesis at wafer :) thank you for the support
- 3mo ago
- AussieWog93 3mo agoThe 2600 tok/s is an "aggregate", not the actual throughput.
- technoabsurdist 3mo agoyes it is 213 tok/s single stream (so per user)
- 3836293648 3mo agoSo per subagent*.
- alienbaby 3mo ago*per stream, I guess is more accurate than either?
- unrvl22 3mo agothat 213 wasn't achieved when saturated though. was probably more like 30 tps per stream when doing 2.6k tps.
- Schiendelman 3mo agoI'm not surprised to see competition with Blackwell. Rubin is 5x faster than Blackwell at inference - Blackwell is the last generation Nvidia didn't optimize specifically for inference. If I'm missing something, please let me know!
- nullc 3mo agohow do you get 5x faster at inference when inference is memory bandwidth limited? getting 5x the memory bandwidth of a h100 seems physically difficult.
- Schiendelman 3mo agoRubin has 22TB/s of memory bandwidth vs Blackwell's 8TB/s. NVLink 6 doubles interconnect speed. Plus they're moving to 3nm from ~4nm. (Previously this comment said Rubin did native NVFP4, but Blackwell does too! Rubin just also trains with native NVFP4, which Blackwell does not.)
- boredatoms 3mo agoMoving to lower bits is not a slam dunk, the model itself might degrade too much
- Schiendelman 3mo agoOf course, but for most workflows it's fine.
- zackangelo 3mo agoBlackwell supports nvfp4 natively.
- Schiendelman 3mo agoYou're right - Rubin is better at NVFP4 training, not inference, thank you for catching me!
- p1esk 3mo agoThere’s noticeable accuracy degradation when they switched from fp8 to mxfp4
- throwdbaaway 3mo agoAnd somehow they claimed that it is "lossless".
- greyb 3mo agoWafer discontinued their own "Wafer Pass" flagship coding plan within weeks of launch and had to issue prorated refunds. Now they're bragging about squeezing costs down even further via quantization, even though their implementation is clearly lacking. [1] https://www.ycombinator.com/launches/Q9i-wafer-pass-flat-rate-access-to-the-fastest-open-source-llms https://www.ycombinator.com/launches/Q9i-wafer-pass-flat-rat...
- alienbaby 3mo agoI'm interested if anyone knows how much legwork the assumed 60% cache hit, plus running a quantised model is doing? Esp. compared to what the headline half implies is a full fat GLM5.2
- killingtime74 3mo agoNo word on what this actually means as a consumer. What's the price. Is it lower than NVIDIA serving?
- mixtureoftakes 3mo agoThey seem to be serving it at 3x the price while also struggling with maintaining uptime on openrouter; while the vercel router advertizes even bigger speeds but has no clear uptime stats I guess you really do have to try it at least for some time to actually know
- nxtfari 3mo agoI think we should make it illegal to not specify the quantization in the headline for these types of posts.
- ahmadyan 3mo agoIts MXFP4
- ozgrakkurt 3mo agoA nice filter is checking for the `.ai` in the end. It is very likely slop if you see that. Slop meaning low-effort/clickbait/shallow/useless/scam etc.
- 48484949 3mo ago[flagged]
- IshKebab 3mo agoAnd to use the heading "Why this matters".
- zahlman 3mo agoI don't know what you mean by "quantization". I guess you're asking to explain how they measure "performance", but that kind of thing often won't neatly fit in a title. I do think the phrasing is weird, though. The performance may be improving, but it isn't the thing getting "faster" (e.g. responses to queries might get "faster"). And the dollars aren't getting cheaper; the performance is. "Performance per dollar" is a rate; it is not getting faster or cheaper. It should just say "Performance per dollar is increasing".
- hmry 3mo agoThey're asking what precision the model parameters were quantized/rounded to, look up "LLM quantization" to learn more. The old headline was something like "We served GLM5.2 on AMD MI355X at 2626 tok/s/node". They think it's bad to advertise performance numbers like that without specifying how much you quantized the model (Since you can 'cheat' more performance by quantizing more, and not mentioning the reduced output quality)
- beffjezos 3mo agoThis is very interesting and yet not at the same time. This looks to be optimized for single-stream LLM traffic which is not viable to serve in a production setting. It's only interesting to hobbyists that want to run the model locally. It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark hacking that doesn't stand up to real-world application.
- technoabsurdist 3mo agohi yes it’s not optimized for single stream it’s optimized for total node throughput
- beffjezos 3mo agoOh, that's much better then. A good metric to share is the tokens per second per user for the node rather than the total throughput of the node. It disambiguates what's being optimized for much better than your blog post currently does.
- technoabsurdist 3mo agosounds good feedback taken, thanks beffjezos
- foobar10000 3mo agoWell, for a lot of agentic stuff nowadays, having 250k-500K context is where things live - and the benchmarks don't really show that unfortunately - but they could :)
- wmf 3mo agoYou got it backwards; it's ~200 on single stream so the 2,600 is achieved with ~13 streams.
- beffjezos 3mo ago
- hassaanr 3mo agoWhile cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.
- google234123 3mo agoFirst thing I noticed as well
- tw1984 3mo agofrom memory, it is like 96-98% of the accuracy.
- lgessler 3mo agoAccuracy isn't a meaningful metric here without reference to a specific task.
- flawn 3mo agoAdditionally, I'd imagine quantization to have more side-effects than just slightly lower performance (on whatever task). You are basically removing information, and that information could be by chance what the model needs to fulfill it exactly the way you'd want to do - although it's still fully capable. I am not sure if this is really different from "lower performance" but open to hear your opinions.
- EduardoBautista 3mo agoAnd that 2%-4% makes all the difference.
- fpaf 3mo agoYes, it's like saying "we took off a big chunk of his brain but look! He can still breathe autonomously, swallow food and walk almost straight, which is like 95% of what he did before!"
- villgax 3mo agoThey fail to mention non speculative numbers & whether baseline was nvfp4 as well. So much for erosion against an older gen
- zuzululu 3mo agoyeah but we are still far far away from being able to run the frontier model equivalents locally without significant quantization even having something like opus 4.8 locally would completely change the landscape
- calin2k 3mo agothen why is token per dollar getting more expensive?
- AtlasBarfed 3mo agoBecause they are dumping/subsidizing it token processing to try and get companies to fire as many people as possible. So they'll be dependent upon the companies when they have to Jack the rates
- FeepingCreature 3mo agoBecause lots of people are willing to pay more dollar for smarter token.
- ilaksh 3mo agoThere are a limited number of these available in comparison to demand. I think people figured out that LLMs and VLMs can do real work that can replace a lot of humans. And for plenty of jobs, it's good enough to reduce already outsourced staff by 75-90% at a fraction of the cost.
- shevy-java 3mo agoBut RAM prices skyrocketed! The AI companies owe use money. As does e. g. NVIDIA for becoming a cartel.
- bitwize 3mo ago(in a high-pitched, pathetic regency-era British orphan voice) Please sir, may I have some compute as well?
- hahahaa 3mo agoWhat is a knee, in performance talk?
- johanvts 3mo agoThat sounds literally impossible.
- dtgriscom 3mo agoAgreed. The writer is pretty loose with their comparisons: * What does it mean for "performance per dollar" to get faster? Higher, maybe; rise faster than it has in the past, maybe, but just "faster"? Nope. * The article cites some equipment as being "2x cheaper". I think they mean "half the cost", but if so they should say it.
- gowthamsaiyadav 3mo agoworld is not limited by Nvidia, AMD can be used
- adammarples 3mo agoSlight criticism of the headline there, you can't get cheaper per dollar.
- deleted 3mo ago[deleted]
- tim333 3mo agoNot a new phenomena - performance per dollar has been fairly steadily exponentialling since 1900 or so 1900 - 2010 https://www.thekurzweillibrary.com/exponential-growth-of-computing https://www.thekurzweillibrary.com/exponential-growth-of-com... 1939 - 2023 https://medium.com/@timventura/kurzweils-law-for-the-ai-age-b00e82c84d5c https://medium.com/@timventura/kurzweils-law-for-the-ai-age-...
- pullrun 3mo ago[flagged]
- jessinra98 3mo ago[flagged]
- conorcleary 3mo ago*especially as many currencies weaken
- BurningFrog 3mo agoSo... the headline is about performance per dollar per dollar?
- ilaksh 3mo agoCan you actually rent an MI355X per hour anywhere right now?
- ilaksh 3mo agoThe compute-in-memory and neuromorphic paradigms are likely to push this much, much farther over the next decade as more radical improvements make it out of the lab. Sooner or later it will involve new materials and new nano devices and providing multiple orders of magnitude better efficiency. And just scaling up existing things like MRAM.
- mchusma 3mo agoI was hoping they would be discussing some path to improving things faster and cheaper. But in this post it looks like they offer quantized version for the same price as full version, and a fast version at much higher cost.
- gcanyon 3mo agoIsn't this pretty much a given? Performance per dollar has to be a ratcheting function because how would something more expensive replace something less expensive?
- paulreaney 3mo ago[dead]
- sometimelurker 3mo agoI like the metric of tok/joule a lot. it really brings to mind a lot of really nice ideas about energy and work and ideas and thought and efficiency
- servola 3mo ago[flagged]
- investmuse 3mo ago[flagged]