13 ms·
Whisper: Nvidia RTX 4090 vs. M1 Pro with MLX
- whywhywhywhy 3y agoFind these findings questionable unless Whisper is very poorly optimized the way it was run on a 4090. I have a 3090 and an M1 Max 32GB and and although I haven't tried Whisper the inference difference on Llama and Stable Diffusion between the two is staggering, especially with Stable Diffusion where SDXL is about 0:09 seconds 3090 and 1:10 minute on M1 Max.
- agloe_dreams 3y agoIt all is really messy, I would assume that almost any model is poorly optimized to run on Apple Silicon as well.
- kamranjon 3y agoThere has been a ton of optimization around whisper with regards to apple silicon, whisper.cpp is a good example that takes advantage of this - also this article is specifically referencing the new apple MLX framework which I’m guessing your tests with llama and stable diffusion weren’t utilizing.
- sbrother 3y agoI assume people are working on bringing a MLX backend to llama.cpp... Any idea what the state of that project is?
- tgtweak 3y agohttps://github.com/ml-explore/mlx-examples https://github.com/ml-explore/mlx-examples Several people working on mlx-enabled backends to popular ML workloads but it seems inference workloads are the most accelerated vs generative/training.
- oceanplexian 3y agoM1 Max has 400GB/s of memory bandwidth and a 4090 has 1TB/s of memory bandwidth, M1 Max has 32 GPU cores and a 4090 has 16,000. The difference is more about how well the software is optimized for the hardware platform than any performance difference between the two, which are frankly not comparable in any way.
- codedokode 3y agoI think that 4090 has 16000 ALUs, not "cores" (let's call a component capable to execute instructions independently from others, a "core"). And M1 Max probably has more than 1 ALU in every core, otherwise it resembles an ancient GPU.
- rsynnott 3y agoYeah; 'core' is a pretty meaningless term when it comes to GPUs, or at least it's meaningless outside the context of a particular architecture. We may just be thankful that this particular bit of marketing never caught on for CPUs.
- segfaultbuserr 3y ago> M1 Max has 32 GPU cores and a 4090 has 16,000. Apple M1 Max has 32 GPU cores, each core contains 16 Execution Units, each EU has 8 ALUs (also called shaders), so overall there are 4096 shaders. Nvidia RTX 4090 contains 12 Graphics Processing Clusters, each GPC has 12 Streaming Multi-Processors, and each SM has 128 ALUs, overall there are 18432 shaders. A single shader is somewhat similar to a single lane of a vector ALU in a CPU. One can say that a single-core CPU with AVX-512 has 8 shaders, because it can process 8 FP64s at the same time. Calling them "cores" (as in "CUDA core") is extremely misleading, so "shader" became the common name for a GPU's ALU due to that. If Nvidia is in charge of marketing a 4-core x86-64 CPU, they would call it a CPU with 32 "AVX cores" because each core has 8-way SIMD.
- jrk 3y agoActually each of those x86 CPUs probably has at least two AVX FMA units, and can issue 16xFP32 FMAs per cycle – it’s at least “64 AVX cores”! :)
- ps 3y agoI have 4090 and M1 Max 64GB. 4090 is far superior on Llama 2.
- astrodust 3y agoOn models < 24GB presumably. "Faster" depends on the model size.
- brucethemoose2 3y agoIn this case, the 4090 is far more memory efficient thanks to ExLlamav2. 70B in particular is indeed a significant compromise on the 4090, but not as much as you'd think. 34B and down though, I think Nvidia is unquestionably king.
- michaelt 3y agoDoesn't running 70B in 24GB need 2 bit quantisation? I'm no expert, but to me that sounds like a recipe for bad performance. Does a 70B model in 2-bit really outperform a smaller-but-less-quantised model?
- brucethemoose2 3y ago2.65bpw, on a totally empty 3090 (and I mean totally empty). I woukd say 34B is the performance sweetspot, yeah. There was a long period where allow we had in the 33B range was llamav1, but now we have Yi and Codellamav2 (among others).
- jb1991 3y agoBut are you using the newly released Apple MLX optimizations?
- ps 3y agoIt's been approximately 2 months since I have tested it, so probably not.
- woadwarrior01 3y agoYou're taking benchmark numbers from a latent diffusion model's (SDXL) inference and extrapolating them to encoder-decoder transformer model's (Whisper) inference. These two model architectures have little in common (except perhaps the fact that Stable Diffusion models use a pre-trained text encoder from clip, which again is very different from an encoder-decoder transformer).
- brucethemoose2 3y agoThe point still stands though. Popular models tend to to have massively hand optimized Nvidia implementations. Whisper is no exception: https://github.com/Vaibhavs10/insanely-fast-whisper https://github.com/Vaibhavs10/insanely-fast-whisper SDXL is actually an interesting exception for Nvidia because most users still tend to run it in PyTorch eager mode. There are super optimized Nvidia implementations, like stable-fast but their use is less common. Apple, on the other hand, took the odd step of hand writing a Metal implementation themselves, at least for SD 1.5.
- tgtweak 3y agoModest 30x speedup
- kkielhofner 3y agoThis will determine who has a shot at actually being Nvidia competitive. What I like to say is (generally speaking) other implementations like AMD (ROCm), Intel, Apple, etc are more-or-less at the “get it to work” stage. Due to their early lead and absolute market dominance Nvidia has been at the “wring every last penny of performance out of this” stage for years. Efforts like this are a good step but they still have a very long way to go to compete with multiple layers (throughout the stack) of insanely optimized Nvidia/CUDA implementations. Bonus points nearly anything with Nvidia is a docker command that just works on any chip they’ve made in the last half decade from laptop to datacenter. This can be seen (dramatically) with ROCm. I recently took the significant effort (again) to get an LLM to run on an AMD GPU. The AMD GPU is “cheaper” in initial cost but when the dollar equivalent (to within 10-30%) Nvidia GPU is 5-10x faster (or whatever) you’re not saving anything. You’re already at a loss unless your time is free just to get it to work (random patches, version hacks, etc) and then the performance just isn’t even close so the “value prop” of AMD currently doesn’t make any sense whatsoever. The advantage for Apple is you likely spent whatever for the machine anyway, and when you have it just sitting in front of you for a variety of tasks the value prop increases significantly.
- KingOfCoders 3y ago"I haven't tried Whisper" I haven't tried the hardware/software/framework/... of the article, but I have an opinion on this exact topic.
- xxs 3y agoThe topic is benchmarking some hardware and specific implementation of some tool. The provided context is n earlier version of hardware where known implementations perform drastically differently, an order of magnitude differently. That leaves the question why that specific tool exhibits the behavior described in the article.
- deleted 3y ago[deleted]
- tgtweak 3y agoReading through some (admittedly very early) MLX docs and it seems that convolutions (as used heavily in GANs and particularly stable diffusion) are not really seeing meaningful uplifts on MLX at all, and in some cases are slower than on the cpu. Not sure if this is a hardware limitation or just unoptimized MLX libraries but I find it hard to believe they would have just ignored this very prominent use case. It's more likely that convolutions use high precision and much larger tile sets that require some expensive context switching when the entire transform can't fit in the gpu.
- liuliu 3y agoBoth of your SDXL and M1 Max number should be faster (of course, it depends on how many steps). But the point stands, for SDXL, 3090 should be 5x to 6x faster than M1 Max and should be 2x to 2.5x faster than M2 Ultra.
- stefan_ 3y agoHaving used whisper a ton, there are versions of it that have one or two magnitudes of better performance at the same quality while using less memory for reasons I don't fully understand. So I'd be very careful about your intuition on whisper performance unless it's literally the same software and same model (and then the comparison isn't very meaningful still, seeing how we want to optimize it for different platforms).
- mv4 3y agoThank you for sharing this data. I've just been debating between M2 Mac Studio Max and a 64GB i9 10900x with RTX 3090 for personal ML use. Glad I chose the 3090! Would love to learn more about your setup.
- Flux159 3y agoHow does this compare to insanely-fast-whisper though? https://github.com/Vaibhavs10/insanely-fast-whisper https://github.com/Vaibhavs10/insanely-fast-whisper I think that not using optimizations allows this to be a 1:1 comparison, but if the optimizations are not ported to MLX, then it would still be better to use a 4090. Having looked at MLX recently, I think it's definitely going to get traction on Macs - and iOS when Swift bindings are released https://github.com/ml-explore/mlx/issues/15 https://github.com/ml-explore/mlx/issues/15 (although there might be some C++20 compilation issue blocking right now).
- brucethemoose2 3y agoThis is the thing about Nvidia. Even if some hardware beats them in a benchmark, if its a popular model, there will be some massively hand optimized CUDA implementation that blows anything else out of the water. There are some rare exceptions (like GPT-Fast on AMD thanks to PyTorch's hard work on torch.compile, and only in a narrow use case), but I can't think of a single one for Apple Silicon.
- MBCook 3y agoI wouldn’t be surprised a $2k top of the line GPU is a match/better than the built in accelerator on a Mac. Even if the Mac was slightly faster you could just stick multiple GPUs in a PC. To me the news here is how well the Mac runs without needing that additional hardware/large power draw on this benchmark.
- NorwegianDude 3y agoThe power draw is not impressive here. Sure, it's low, but if you account for performance/W then the GPU is much more efficient.
- rfoo 3y ago> but I can't think of a single one for Apple Silicon. The post here is exactly one for Apple Silicon. It compared a naive implementation in PyTorch which may not even keep 4090 busy (for smaller/not-that-compute-intensive models having the entire computation driven by Python is... limiting, which is partly why torch.compile gives amazing improvements) to a purposedly-optimized one (optimized for both CPU/GPU efficiency) for Apple Silicon one.
- bcatanzaro 3y agoWhat precision is this running in? If 32-bit, it’s not using the tensor cores in the 4090.
- DeathArrow 3y agoOk, OpenAI will ditch Nvidia and buy macs instead. :)
- baldeagle 3y agoOnly if Sam Altman is appointed to the Apple board. ;)
- tiffanyh 3y agoKey to this article is understanding it’s leveraging the newly released Apple MLX, and their code is using these Apple specific optimizations. https://news.ycombinator.com/item?id=38539153 https://news.ycombinator.com/item?id=38539153
- modeless 3y agoAlso, this is not comparing against an optimized Nvidia implementation. There are faster implementations of Whisper. Edit: OK I took the bait. I downloaded the 10 minute file he used and ran it on my 4090 with insanely-fast-whisper, which took two commands to install. Using whisper-large-v3 the file is transcribed in less than eight seconds. Fifteen seconds if you include the model loading time before transcription starts (obviously this extra time does not depend on the length of the audio file). That makes the 4090 somewhere between 6 and 12 times faster than Apple's best. It's also much cheaper than M2 Ultra if you already have a gaming PC to put it in, and still cheaper even if you buy a whole prebuilt PC with it. This should not be surprising to people, but I see a lot of wishful thinking here from people who own high end Macs and want to believe they are good at everything. Yes, Apple's M-series chips are very impressive and the large RAM is great, but they are not competitive with Nvidia at the high end for ML.
- isodev 3y agoIt also wasn’t optimised for Apple Silicon. Given how the different platforms performed in this test, the conclusions seem pretty solid.
- Lalabadie 3y agoThere will be a lot of debate about which is the absolute best choice for X task, but what I love about this is the level of performance at such a low power consumption.
- SlavikCA 3y agoIt's easy to run Whisper on my Mac M1. But it's not using MLX out of the box. I spend an hour or two, trying to run figure out what I need to install / configure to enable it to use MLX. Was getting cryptic Python errors, Torch errors... Gave up on it. I rented VM with GPU, and started Whisper on it within few minutes.
- xd1936 3y agoI've really enjoyed this macOS Whisper GUI[1]. It doesn't use MLX, but does use Metal. 1. https://goodsnooze.gumroad.com/l/macwhisper https://goodsnooze.gumroad.com/l/macwhisper
- tambourine_man 3y agoIs was released last week. Give it a month or two
- JCharante 3y agoHmm, I've been using this product for whisper https://betterdictation.com/ https://betterdictation.com/
- jonnyreiss 3y agoI was able to get it running on MLX on my M2 Max machine within a couple minutes using their example: https://github.com/ml-explore/mlx-examples/tree/main/whisper https://github.com/ml-explore/mlx-examples/tree/main/whisper
- etchalon 3y agoThe shocking thing about these M series comparisons is never "the M series is fast as the GIANT NVIDIA THING!" it's always "Man, the M series is 70% as fast with like 1/4 the power."
- kllrnohj 3y agoIt's not really that shocking. Power consumption is non-linear with respect to frequency and you see this all the time in high end CPU & GPU parts. Look at something like the eco modes on the Ryzen 7xxx for a great example. The 7950X stock pulls something like 260w on an all core workload at 5.1ghz. Yet enable the 105w eco mode and that power consumption plummets to 160w at 4.8ghz. That means the last 300mhz of performance, which is borderline inconsequential (~6% performance loss), costs 100W. The 65W option then cuts that in half almost again down to 88w (at 4ghz now), for a "mere" 20% reduction in performance. Or phrased differently, for 1/3rd the power the 7950X will give you 75% of the performance of a 7950X. Matching performance while using less power is impressive. Using less power while also being slower not so much, though.
- nottorp 3y ago> Or phrased differently, for 1/3rd the power the 7950X will give you 75% of the performance of a 7950X. So where is the - presumably much cheaper - 65 W 7950X?
- dotnet00 3y agoWhy would it be much cheaper? The chips are intentionally clocked higher than the most efficient point because the point of the CPU is raw speed, not power consumption, especially since 7950x is a desktop chip. His point is that M-CPUs being somewhat competitive but much more efficient is not as stunning since you're comparing a CPU tuned to be at the most efficient speed to a CPU tuned to be at the highest speed. Similarly a 4090's power consumption drops dramatically if you underclock or undervolt even slightly, but what's the point? You're almost definitely buying a 4090 for its raw speed.
- tgtweak 3y agoDoes this translate to other models or was whisper cherry picked due to it's serial nature and integer math? looking at https://github.com/ml-explore/mlx-examples/tree/main/stable_diffusion https://github.com/ml-explore/mlx-examples/tree/main/stable_... seems to hint that this is the case: >At the time of writing this comparison convolutions are still some of the least optimized operations in MLX. I think the main thing at play is the fact you can have 64+G of very fast ram directly coupled to the cpu/gpu and the benefits of that from a latency/co-accessibility point of view. These numbers are certainly impressive when you look at the power packages of these systems. Worth considering/noting that the cost of m3 max system with the minimum ram config is ~2x the price of a 4090...
- densh 3y agoApple’s silicon memory is fast only in comparison to consumer CPUs that stagnated for ages with having only 2 memory channels which was fine in 4 core era but wakes no sense at all with modern core counts. Memory scaling on GPUs is much better, even on the consumer front.
- mightytravels 3y agoUse this Whisper derivative repo instead - one hour of audio gets transcribed within a minute or less on most GPUs - https://github.com/Vaibhavs10/insanely-fast-whisper https://github.com/Vaibhavs10/insanely-fast-whisper
- thrdbndndn 3y agoCould someone elaborate how this is accomplished and if there is any quality disparity compared to original? Repos like https://github.com/SYSTRAN/faster-whisper https://github.com/SYSTRAN/faster-whisper makes immediate sense on why it's faster than the original implementation, and lots of others do so by lowering quantization precision etc (and worse results). but this one, it's not very clear how. Especially considering it's even much faster.
- lern_too_spel 3y agoThe Acknowledgments section on the page that GP shared says it's using BetterTransformer. https://huggingface.co/docs/optimum/bettertransformer/overview https://huggingface.co/docs/optimum/bettertransformer/overvi...
- mightytravels 3y agoFrom what I can see it is parallel batch processing - default for that repo is 24. You can reduce batches and if you use 1 it's as fast or slow as Whisper. Quality is the exact same (same large model used).
- claytonjy 3y agoAnecdotally I've found ctranslate2 to be even faster than insanely-fast-whisper. On an L4, using ctranslate2 with a batch size as low as 4 beats all their benchmarks except the A100 with flash attention 2. It's a shame faster-whisper never landed batch mode, as I think that's preventing folks from trying ctranslate2 more easily.
- theschwa 3y agoI feel like this is particularly interesting in light of their Vision Pro. Being able to run models in a power efficient manner may not mean much to everyone on a laptop, but it's a huge benefit for an already power hungry headset.
- darknoon 3y agoWould be more interesting if Pytorch with MPS backend was also included.
- LiamMcCalloway 3y agoI'll take this opportunity to ask for help: What's a good open source transcription and diarization app or work flow? I looked at https://github.com/thomasmol/cog-whisper-diarization https://github.com/thomasmol/cog-whisper-diarization and https://about.transcribee.net/ https://about.transcribee.net/ (from the people behind Audapolis) but neither work that well -- crashes, etc. Thank you!
- dvfjsdhgfv 3y agoI developed my own solutions, pretty rudimentary - it divides the MP3s into chunks that Whisper is able to handle and then sends them one by one to the API to transcribe. Works as expected so far, it's just a couple of lines of Python code.
- mosselman 3y agoI would like to know the same. It shouldn’t be so hard since many apps have this. But what is the most reliable way right now?
- iAkashPaul 3y agoThere's a better parallel/batching that works on the 30s chunks resulting in 40X. From HF at https://github.com/Vaibhavs10/insanely-fast-whisper https://github.com/Vaibhavs10/insanely-fast-whisper This is again not native PyTorch so there's still room to have better RTFX numbers.
- bee_rider 3y agoHmm… this is a dumb question, but the cookie pop up appears to be in German on this site. Does anyone know which button to press to say “maximally anti-tracking?”
- layer8 3y agoIf only we had a way to machine-translate text or to block such popups.
- bee_rider 3y agoI think it is better not to block these sorts of pop-ups, they are part of the agreement to use the site after all. Anyway the middle button is “refuse all” according to my phone, not sure how accurate the translation is or if they’ll shuffle the buttons for other people. It is poor design to have what appear to be “accept” and “refuse” both in green.
- layer8 3y agoAccording to GDPR, the site is not allowed to track you as long as you haven’t given your consent. There is no agreement until you actually agreed to something.
- throwaw33333434 3y agoMETA: is M3 Pro good enough to run Cyberpunk 2077 smoothly? Does Max really makes a difference?
- ed_balls 3y ago14 inch M3 Max may overheat.
- jauntywundrkind 3y agoI wonder how AMD's XDNA accelerator will fair. They just shipped 1.0 of the Ryzen AI Software and SDK. Alleges ONNX, PyTorch, and Tensorflow support. https://www.anandtech.com/show/21178/amd-widens-availability-of-ryzen-ai-software-for-developers-xdna-2-coming-with-strix-point-in-2024 https://www.anandtech.com/show/21178/amd-widens-availability... Interestingly, the upcoming XDNA2 supposedly is going to boost generative performance a lot? "3x". I'd kind of assumed these sort of devices would mainly be helping with inference. (I don't really know what characterizes the different workloads, just a naive grasp.)
- deleted 3y ago[deleted]
- lars512 3y agoIs there a great speech generation model that runs on MacOS, to close the loop? Something more natural than the built in MacOS voices?
- treprinum 3y agoYou can try VALL-E; it takes around 5s to generate a sentence on a 3090 though.
- deleted 3y ago[deleted]
- runjake 3y agoAnyone have overall benchmarks or qualified speculation on how an optimized implementation for a 4070 compares against the M series -- especially the M3 Max? I'm trying to decide between the two. I figure the M3 Max would crush the 4070?
- sim7c00 3y agolooking at the comments perhaps the article could be more eptly titled. the author does stress these benchmarks, maybe better called test runs, are not of any scientific accuracy or worth, but simply to demonstrate what is being tested. i think its interesting though that apple and 4090s are even compared in any way since the devices are so vastly different. id expect the 4090 to be more powerful, but apple optimized code runs really quick on apple silicon despite this seemingly obvious fact, and that i think is interesting. you dont need a 4090 to do things if you use the right libraries. is that what i can take from it?
- ex3ndr 3y agoSo running on M2 Ultra would beat 4090 by 30%? (since it has 2x of gpu cores)
- atty 3y agoI think this is using the OpenAI Whisper repo? If they want a real comparison, they should be comparing MLX to faster-whisper or insanely-fast-whisper on the 4090. Faster whisper runs sequentially, insanely fast whisper batches the audio in 30 second intervals. We use whisper in production and this is our findings: We use faster whisper because we find the quality is better when you include the previous segment text. Just for comparison, we find that faster whisper is generally 4-5x faster than OpenAI/whisper, and insanely-fast-whisper can be another 3-4x faster than faster whisper.
- moffkalast 3y agoIs insanely-fast-whisper fast enough to actually run on the CPU and still trascribe in realtime? I see that none of these are running quantized models, it's still fp16. Seems like there's more speed left to be found. Edit: I see it doesn't yet support CPU inference, should be interesting once it's added.
- atty 3y agoInsanely fast whisper is mainly taking advantage of a GPU’s parallelization capabilities by increasing the batch size from 1 to N. I doubt it would meaningfully improve CPU performance unless you’re finding that running whisper sequentially is leaving a lot of your CPU cores idle/underutilized. It may be more complicated if you have a matrix co-processor available, I’m really not sure.
- youssefabdelm 3y agoDoes insanely-fast-whisper use beam size of 5 or 1? And what is the speed comparison when set to 5? Ideally it also exposes that parameter to the user. Speed comparisons seem moot when quality is sacrificed for me, I'm working with very poor audio quality so transcription quality matters.
- atty 3y agoOur comparisons were a little while ago so I apologize I can’t remember if we used BS 1 or 5 - whichever we picked, we were consistent across models. Insanely fast whisper (god I hate the name) is really a CLI around Transformers’ whisper pipeline, so you can just use that and use any of the settings Transformers exposes, which includes beam size. We also deal with very poor audio, which is one of the reasons we went with faster whisper. However, we have identified failure modes in faster whisper that are only present because of the conditioning on the previous segment, so everything is really a trade off.
- accidbuddy 3y agoAbout whisper, anyone knows a project (github) about using the model in real-time? I'm studying a new language, and it appears to be a good chance to use and learning pronunciation vs. word.
- samx81 3y agoThis one uses faster-whisper as the backend, I've tried with small model and the performance is good. https://github.com/collabora/WhisperLive https://github.com/collabora/WhisperLive The is another one that uses huggingface's implementation, but I haven't tried it since my spec doesn't support flash-att2 https://github.com/luweigen/whisper_streaming https://github.com/luweigen/whisper_streaming
- accidbuddy 3y agoThanks. I'll try.
- brcmthrowaway 3y agoShocked that Apple hasn't released a high end compute chip competitive with NVIDIA
- atlas_hugged 3y agoTL;DR If you compare whisper on a mac with Mac optimized build Vs on a pc with a few NON-optimized NVIDIA build The results are close! If nvidia optimized is compared, it’s not even remotely close. Pfft I’ll be picking up a Mac but I’m well aware it’s not close to Nvidia at all. It’s just the best portable setup I can find that I can run completely offline. Do people really need to make these disingenuous comparisons to validate their purchase? If a mac fits your overall use case better, get a Mac. If a pc with nvidia is the better choice, get it. Why all these articles of “look my choice wasn’t that dumb”??
- 2lkj22kjoi 3y ago4090 -> 82 TFLOPS M3 MAX GPU -> 10 TFLOPS It is 8 times slower than 4090. But yeah, you can claim that a bike has a faster acceleration than Ferrari, because it could reach the speed of 1km per hour faster...