11 ms·
"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance
by m_a_g 2y ago
"TensorWave is a cloud provider specializing in AI workloads. Their platform leverages AMD’s Instinct™ MI300X accelerators, designed to deliver high performance for generative AI workloads and HPC applications."
I suggest taking the report with a grain of salt.
- epolanski 2y agoWell, there's the beauty of specifying exactly how you ran your benchmark, it is easy to reproduce and disprove or confirm (assuming you got the hardware).
- scotty79 2y agoAs easy as getting yourself 8 H100 and 8 MI300X. Fun weekend project for anybody.
- idiliv 2y agoYou can rent them online for ~ 4-5 $ per hour per GPU. Not cheap, but definitely feasible as a weekend project.
- _zoltan_ 2y agowhere can I rent a H100 for 4-5 dollars an hour? AWS doesn't let you use p5 instances (not getting a quota as a private person), lambda cloud is sold out.
- lhl 2y agoIt looks like Runpod currently (checked right now) has "Low" availability of 8x MI300 SXM (8x$4.89/h), H100 NVL (8x$4.39/h), and H100 (8x$4.69/h) nodes for anyone w/ some time to kill that wants to give the shootout a try.
- darrick_horton 2y agoWe'd be happy to provide access to MI300X at TensorWave so you can validate our results! Just shoot us an email or fill out the form on our website
- Jlagreen 2y agoIf you're able to advertise available GPU compute in some public forums then it's enough to tell us about the demand of MI300X in cloud ...
- lhl 2y agoYou're joking/trolling right? There are literally 10's of thousands of H100s available on gpulist right now, does that mean there's no cloud demand for Nvidia gpus? (I notice from your comment history that you seem to be some sort of bizarre NVDA stan account, but come on, be serious)
- deleted 2y ago[deleted]
- impulser_ 2y agoIf they used Nvidia's chip would this somehow make the blog post better?
- aurareturn 2y agoFor one, they didn't use TensorRT in the test. Also, stuff like this is hard to take the results seriously: * To make an accurate comparison between the systems with different settings of tensor parallelism, we extrapolate throughput for the MI300X by 2. * All inference frameworks are configured to use FP16 compute paths. Enabling FP8 compute is left for future work. They did everything they can to make sure AMD is faster.
- braiamp 2y agoI see it as they did everything they can to compare the specific code path. If your workload scales with FP16 but not with tensor cores, then this is the correct way to test. What do you need for LLM inference?
- HPsquared 2y agoCouldn't they find a real workload that does this?
- ebalit 2y agovLLM inference of Mixtral in fp16 is a real workload. I guess the details are there because of the different inference engine used. You need the most similar compute tasks to be ran but the compute kernels can't be the same as in the end they need to be ran by a different hardware.
- davidguetta 2y agoSo they just multipiled their results per 2 ^^ ?
- ebalit 2y agoYou need 2 H100 to have enough VRAM for the model whereas you need only 1 MI300X. Doubling the total throughput (for all completions) of 1 MI300X to simulate the numbers for a duplicated system is reasonable. They should probably show separately the throughput per completion as the tensor parallelism is often used for that purpose in addition to the doubling the VRAM.
- nabla9 2y agoThe salt is in the plain sight. The do the standard AMD comparison: 8x AMD MI300X (192GB, 750W) GPU 8x H100 SXM5 (80GB, 700W) GPU The fair comparison would be against 8x H100 NVL (188GB, <800W) GPU Price tells a story. If AMD performance would be in par with Nvidia they would not sell their cards for 1/4 price.
- fleischhauf 2y agoAMDs deep learning libraries are very bad the last time I checked, nobody uses amd in that space for that reason. Nvidia has a quazi monopoly, that's the main reason for the price difference IMHO.
- nabla9 2y ago> that's the main reason for the price difference IMHO. Explain why the performance difference does not matter? AMD does only 33% better with a chip that has 2X transistors and 2X memory.
- More-nitors 2y agothis... nearly 95% of deeplearning github repos are "tested using cuda gpu - others, not so sure" the only way to out-run nvidia is to have 3~10x better bang-for-buck. Or AMD can just provide a "DIY unlimited gpu RAM upgrade" kit -- a lot of people are buying macstudio 128gb ram because of its "bigger ram-for-buck" than nvidia gpus
- sroussey 2y agoI heard apple m4 ultra using 256gb HBM for studio and pro, but I don’t buy it. The 256GB maybe. But a HBM memory control that would go unused on laptops doesn’t pass the smell test.
- tracker1 2y agoI think their best option might be more/better prosumer options at the higher end of consumer pricing. Getting more hobbyists into play just on the value proposition.