15 ms·
AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
- qeternity 2y agoWhy the hell are we doing 128 input token benchmarks in 2024. This is not representative of most workloads, and prefill perf is incredibly important.
- ta12653421 2y agoFor understanding: What would be a suitable input length in your oppinion? And why isnt this a good one: Are real-life queries shorter? Or longer? If i count one word as a token, then in my case most of the queries are less than 128 words.
- spacecadet 2y agoIn most cases thats not enough
- Gasp0de 2y agoIncluding the initialization prompt and your history if you have one? I use ChatGPT for a very simple task, to map chat messages to one of 5 supported function calls, and the function definitions alone already take up 200 tokens I think
- stefs 2y agoIt's not just the current prompt, but the whole conversation, if possible. Or, if you want the AI to summarise an article, the article has to fit in. If I understood that correctly, context length is something like session storage or short term memory. If it's too small the AI starts to forget what it's talking about.
- qeternity 2y agoI think today 512 tokens is a minimum. It's not just the query (if you're running a chatbot, which many of us are not). It's the entire context window. It's not uncommon to have a system prompt that is > 512 tokens alone. I would like to see benchmarks for 512, 1024, 4096 and 8192 token inputs.
- rfoo 2y agoIMO the relevant benchmark for now is a mixed stream of requests with 50 (20%), 500 (50%), 2000 (10%) and 50k (20%) input tokens, ignore EOS and decode until you get around 300 output tokens.
- SuchAnonMuchWow 2y agoI'm really interested, do you have a source for those percentages ? I tried to look for some service provider to publish this kind of metrics, but haven't found any.
- rfoo 2y agoSorry, I can't. My employer doesn't publish this kind of metrics, either. What I posted was definitely just some very rough number off my brain.
- zxexz 2y agoIs this an ad for a new, closed-source, GPGPU backend?
- BoredPositron 2y agoPretty much and the test suit is optimized to get the results they wanted.
- nottorp 2y agoPretty sure a useful benchmark for this kind of thing would calculate performance per watt (or per watt and dollar). That info is conspicuously absent from the article.
- acchow 2y agoThe electricity consumption in the cloud is not really important. The H100 rents for about $4.5/hr consuming 0.7kWh in that hour which will likely cost them less than 7 cents.
- nottorp 2y ago> The electricity consumption in the cloud is not really important. That just says you don't run a cloud for profit :)
- latchkey 2y agoHere is the open source backend... https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmarking_brilliance_single_amd_mi300x_vllm/ https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmar...
- m_a_g 2y ago"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications." I suggest taking the report with a grain of salt.
- epolanski 2y agoWell, there's the beauty of specifying exactly how you ran your benchmark, it is easy to reproduce and disprove or confirm (assuming you got the hardware).
- scotty79 2y agoAs easy as getting yourself 8 H100 and 8 MI300X. Fun weekend project for anybody.
- idiliv 2y agoYou can rent them online for ~ 4-5 $ per hour per GPU. Not cheap, but definitely feasible as a weekend project.
- _zoltan_ 2y agowhere can I rent a H100 for 4-5 dollars an hour? AWS doesn't let you use p5 instances (not getting a quota as a private person), lambda cloud is sold out.
- lhl 2y agoIt looks like Runpod currently (checked right now) has "Low" availability of 8x MI300 SXM (8x$4.89/h), H100 NVL (8x$4.39/h), and H100 (8x$4.69/h) nodes for anyone w/ some time to kill that wants to give the shootout a try.
- darrick_horton 2y ago
- rjzzleep 2y ago> Hardware: TensorWave node equipped with 8 MI300X accelerators, 2 AMD EPYC CPU Processors (192 cores), and 2.3 TB of DDR5 RAM. > MI300X Accelerator: 192GB VRAM, 5.3 TB/s, ~1300 TFLOPS for FP16 > Hardware: Baremetal node with 8 H100 SXM5 accelerators with NVLink, 160 CPU cores, and 1.2 TB of DDR5 RAM. > H100 SXM5 Accelerator: 80GB VRAM, 3.35 TB/s, ~986 TFLOPS for FP16 I really wonder about the pricing. In theory the MI300X is supposed to be cheaper, but whether is that is really the case in practice remains to be seen.
- huevosabio 2y agoRunPod [0] is pricing MI300X at $4.89/hr vs $3.89-4.69/hr for H100s. So, probably around the same price? The tests look promising, though! [0] https://runpod.io/ https://runpod.io/
- latchkey 2y agoWe are starting at $4.50/hr [0]. The catch is that we won't have availability until mid August. The weird thing on Runpod is the virtual CPUs, you can't run MI300x in virtual machines yet. It is a missing feature that AMD is working on. [0] https://hotaisle.xyz/pricing/ https://hotaisle.xyz/pricing/
- sigmoid10 2y agoIt doesn't matter. AMD has offered better compute per dollar for a while now, but noone switched because CUDA is the real reason why all serious ML people use Nvidia. Until AMD picks up the slack on their software side, Nvidia will continue to dominate.
- michaelnny 2y agoI'm wondering if the tensor parallel settings have any impact on the performance. My naive guess is yes but not sure. According to the article: """ AMD Configuration: Tensor parallelism set to 1 (tp=1), since we can fit the entire model Mixtral 8x7B in a single MI300X’s 192GB of VRAM. NVIDIA Configuration: Tensor parallelism set to 2 (tp=2), which is required to fit Mixtral 8x7B in two H100’s 80GB VRAM. """
- renonce 2y agoI personally find such comparisons unfair. A good comparison should optimize for each device configuration, which means use a model within the VRAM limit and quantize to 8 bits where it boosts performance etc and avoid shortcomings of both devices unless necessary.
- DarkmSparks 2y agohopper (H100) is the predecessor to the current blackwell architecture. This is a new AMD vs last generation nvidia benchmark.
- triblemaster 2y agoBlackwell won't be here till next year.
- acchow 2y agoNvidia expects to ship 420k Blackwell chips this year.
- DarkmSparks 2y agoGB200 based on blackwell launched in March of this year. https://www.theregister.com/2024/03/21/nvidia_dgx_gb200_nvk72/ https://www.theregister.com/2024/03/21/nvidia_dgx_gb200_nvk7... MI300X launched 3 months earlier at the end of December. H100 launched March 2023,
- muxr 2y agoYou can paper launch as much as you want but Blackwell isn't shipping until Q4. Also Blackwell's lead will be short lived, because mi350x is coming out next year, and it will have a node and architecture advantage. So AMD will be ahead again.
- iAkashPaul 2y agoINT8/FP8 benchmarks would've been great, both cards could have loaded them with around 60GB VRAM instead of TP=2 on H100.
- sva_ 2y agoI try to be optimistic about this. Competition is absolutely needed in this space - $NVDA market cap is insane right now, about $0.6 trillion more than the entire Frankfurt Stock Exchange.
- Rinzler89 2y agoIt's more how little the Frankfurt stock Exchange is worth. And European devs keep wondering why our wages are lower than in the US for the same work. That's why.
- raverbashing 2y agoYes But there's a long list of German companies not on the DAX (though Germany DAX really deserves to be worth less than NVidia)
- earthnail 2y agoThe DAX is made of the 40 most valuable German companies. That’s how it is defined. So the companies not in it, again by definition, matter less.
- qsi 2y agoThe CDAX index has about 360 companies and appears to have a market cap of around 2 trillion EUR vs 1.7 trillion for the DAX 40.
- pell 2y ago> The DAX is made of the 40 most valuable German companies. Not to be too nitpicky here but these are only the publicly traded companies. You have a number of pretty large German companies that are still entirely private such as Aldi, Schwarz Group, Boehringer or Bosch.
- bigfudge 2y agoAlso many medium sized companies which are productive and competitive but not public, so never grow to huge sizes. I view this as a feature not a bug… smaller companies have a more direct connection with their workforce and tend to behave better with them.
- instagraham 2y agoGiven that a lot of projects are written or optimised for CUDA, would it require an industry shift if AMD were to become a competitive source of GPUs for AI training?
- irusensei 2y agoEvery hardware vendor is working to provide something with their own technology. I don't know if it's possible but a lot of very resourceful companies are doing their best to break the CUDA dominance. I really hope it works and hopefully a non proprietary standard emerges.
- yobbo 2y agoThe model code is comparatively tiny compared to pytorch or CUDA itself. Translating models from CUDA/C could be laborious but not a barrier. Making AMD work effortlessly with pytorch et al should make the switch transparent.
- mistymountains 2y agoThese kinds of comments make me think few people have actually tried. My experience has been 1 work day of getting things set up to work the same as before for training and testing (PyTorch).
- imtringued 2y agoYou have to consider that the average person who tried to do machine learning on AMD GPUs got burned in the past decade and has no reason to change their opinion. Also in the past it was much harder to get access to cutting edge GPUs from AMD. The fact that AMD drops GPU support for ROCm quickly also earns them scorn. I don't think it is an unfair assessment. They earned their reputation.
- muxr 2y agoROCm has improved a lot. And you can rent mi300x in the cloud now. So if you have a program that runs on Nvidia GPUs, it takes no time to test it on a cloud mi300x. If it works you can use it and save some money in the process.
- lccerina 2y agoWithout proper statistical metrics (why use average when 95% percentile is widely used?) and performance/watt this is a useless comparison.
- whereismyacc 2y agoaverage says more about throughput, right? 95% would be nice too
- DrNosferatu 2y agoAnd performance/price -> that's the bottom line.
- jvlake 2y agoCool story. How supported is OpenCL compared to CUDA again?
- chillee 2y agoI'm skeptical of these benchmarks for a number of reasons. 1. They're only comparing against VLLM, which isn't SOTA for latency-focused inference. For example, their vllm benchmark on 2 GPUs sees 102 tokens/s for BS=1, gpt-fast gets around 190 tok/s. https://github.com/pytorch-labs/gpt-fast https://github.com/pytorch-labs/gpt-fast 2. As others have pointed out, they're comparing H100 running with TP=2 vs. 2 AMD GPUs running independently. Specifically, > To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2. This is uhh.... very misleading, for a number of reasons. For one, at BS=1, what does running with 2 GPUs even mean? Do they mean that they're getting the results for one AMD GPUs at BS=1 and then... doubling that? Isn't that just... running at BS=2? 3. It's very strange to me that their throughput nearly doubles going from BS=1 to BS=2. MoE models have an interesting property that low amounts of batching doesn't actually significantly improve their throughput, and so on their Nvidia vllm benchmark they just go from 102 => 105 tokens/s throughput when going from BS=1 to BS=2. But on AMD GPUs they go from 142 to 280? That doesn't make any sense to me.
- jvlake 2y agoAt this point in history were still at ROCm vs CUDA... Schmicko hardware is only as good as the software you can write for it.
- DrNosferatu 2y agoThe comparison is between setups with different amounts of GPU RAM and there's no quantification of final performance/price.
- Gasp0de 2y agoSo? If you get twice the RAM at a comparable price and that leads to twice the performance, what's wrong with comparing that?
- DrNosferatu 2y agoNothing wrong - just for transparency. Also, the price difference is not quantified. Additionally, CUDA is a known and tangible software stack - can I try out this "MK1 FLywheel" on my local (AMD) hardware?
- muxr 2y agoNone of those things matter, if all you're looking to do is run your existing workloads on mi300x in the cloud. You get more bang per buck by going AMD.
- robblbobbl 2y ago1.Investing (wasting) the billions. 2. Receive downvotes on ycombinator lol
- huntertwo 2y agoAMD has better seemingly better hardware - but not the production capacity to compete with Nvidia yet. Will be interesting to see margins compress when real competition catches up. Everybody thinks it’s CUDA that makes Nvidia the dominant player. It’s not - almost 40% of their revenue this year comes from mega corporations that use their own custom stack to interact with GPUs. It’s only a matter of time before competition catches up and gives us cheaper GPUs.
- Refusing23 2y ago> but not the production capacity to compete with Nvidia yet. thats just a question of negotiating with tsmc or their few competitors (also didn't tsmc start production of some factories in the US and/or EU?) I mean, nvidia use tsmc, so does amd.
- huntertwo 2y agoYes it is - but Nvidia has larger contracts _right now_. Nvidia has been investing more money in producing more GPUs for longer, so it’s only natural that they have an advantage now. But now that there’s a larger incentive to produce GPUs, their moat will eventually fall. TSMC runs at 100% capacity for top tier processes - their bottleneck is more foundries. These take time to build. So the question becomes - how long can Nvidia remain dominant? It could be quarters or it could be years before any real competitor convinces large customers to switch over. Microsoft and Google are producing their own AI hardware too - nobody wants to depend solely on Nvidia, but they’re currently forced to if they want to keep up.
- fmajid 2y agoIsn't their moat primarily software (CUDA) rather than supply-chain strength?
- pastaguy1 2y agoCan you explain the cuda-less stack a little more or provide a source?
- amelius 2y agoAre these fabbed at the same process node? (Otherwise it's apples and oranges)
- qeternity 2y agoIt's not apples and oranges. These are the top of the line offerings from the respective companies today.
- mark_l_watson 2y agoA good start for AMD. I am also enthusiastic about another non-NVidea inference option: Groq (which I sometimes use). NVidia relies on TMSC for manufacturing. Samsung is building competing manufacturing infrastructure which is also a good thing, so Taiwan is not a single point of failure.
- nextworddev 2y agoWe need more competition in the training space, not inference. For consumer grade inference, there's already many options available.
- KaoruAoiShiho 2y agoPretty bad benchmarks to the point of being deliberately misleading. They benchmarked vllm which is less than half the speed of the inference leader lmdeploy: https://bentoml.com/blog/benchmarking-llm-inference-backends https://bentoml.com/blog/benchmarking-llm-inference-backends They also used Flywheel for AMD while not bothering to turn on Flywheel for Nvidia, which is crazy since Flywheel improves Nvidia performance by 70%. https://mk1.ai/blog/flywheel-launch https://mk1.ai/blog/flywheel-launch In this context the 33% performance lead by AMD looks terrible, and straight up looks slower.
- tgtweak 2y agoThe market (and selling price) is reflecting the perceived value of nvidia's solution vs AMDs - comprehensively including tooling, software, TCO and managability. Also curious how many companies are dropping that much money on those kind of accelerators just to run 8x 7B param models in parallel... You're also talking about being able to train a 14B model on a single accelerator. I'd be curious to see how "full-accelerator train and inferrence" workloads would look ie: Training a 14B param model then inferrence throughput on a 4x14B workload. AMD (and almost every other inferrence claim maker so far... intel and apple specifically) have consistently cherry picked the benchmarks to claim a win over, and ignored the remainder which all show nvidia in the lead - and they've used mid-gen comparison models as many commenters here pointed out in this article.
- fvv 2y agomi300x win in some inference workloads, h100 win in training and some others inference workloads ( fp8 inference with tensorRT-llm , rocm is young but is growing fast ) in a single system ( 8x accelerators ) LLMs, mi300x has very competitive inference TCO vs h100 . also : AMD Instinct MI300X Offers The Best Price To Performance on GPT-4 According To Microsoft, Red Team On-Track For 100x Perf/Watt By 2027 https://wccftech.com/amd-instinct-mi300x-best-price-performance-gpt-4-microsoft-on-track-100x-perf-watt-2027/ https://wccftech.com/amd-instinct-mi300x-best-price-performa...
- zhyder 2y agoShouldn't the right benchmark be performance per watt? It's easy enough to add more chips to do LLM training or inference in parallel. Maybe the benchmark should be performance per $... though I suspect power consumption will eclipse the cost of purchasing the chips from NVDA or AMD (and costs of chips will vary over time and with discounts). EDIT: was wrong on eclipsing; still am looking for a more durable benchmark (performance per billion transistors?) given it's suspected NVDA's chips are over-priced due to demand outstripping supply for now, and AMD's are under- to get a foothold in this market.
- mmoskal 2y agoNot quite. Assume 1kW power consumption (with cooling). At $0.08/kWh (avarage US industrial rate) this is $700 per year. Adjust for more cooling etc and for say 5 years of usage but you still won't be anywhere near the $25k MSRP for H100.
- deleted 2y ago[deleted]
- mistymountains 2y agoI’m a AI Scientist and train a lot of models. Personally I think AMD is undervalued relative to Nvidia. No, chips aren’t as fast as Nvidia’s latest and yes, there are some hoops to get things working. But for most workloads in most industries (ignoring for the moment that AI is likely a poor use of capital), it will be much more cost effective and achieve about the same results.
- latchkey 2y agoWe just got higher performance out of open source. No need for MK1. https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmarking_brilliance_single_amd_mi300x_vllm/ https://www.reddit.com/r/AMD_MI300/comments/1dgimxt/benchmar...