10 ms·
M1 and AMD GPU support. I'm personally more interested in the latter as I haven't yet upgraded my MacBook Pro and I expect that my Vega 20 to be faster than M1
by codelord 5y ago
M1 and AMD GPU support. I'm personally more interested in the latter as I haven't yet upgraded my MacBook Pro and I expect that my Vega 20 to be faster than M1 at ML training.
The raw compute power of M1's GPU seems to be 2.6 TFLOPS (single precision) vs 3.2 TFLOPS for Vega 20. This can give you an estimate of how fast it would be for training.
Just for reference Nvidia's flagship desktop GPU(3090)'s FP32 performance is 35.5 TFLOPS.
- jra101 5y agoVega 20 should be ~13.8 TFLOPs (single precision): https://www.anandtech.com/show/13923/the-amd-radeon-vii-review https://www.anandtech.com/show/13923/the-amd-radeon-vii-revi...
- codelord 5y agoSeems like AMD has been using Vega 20 to refer to two different things. I was talking about the mobile GPU on MacBook Pros which is based on a 14nm chip. The full name is Radeon Pro Vega 20: https://www.amd.com/en/graphics/radeon-pro-vega-20-pro-vega-16 https://www.amd.com/en/graphics/radeon-pro-vega-20-pro-vega-... https://www.techpowerup.com/gpu-specs/radeon-pro-vega-20.c3263 https://www.techpowerup.com/gpu-specs/radeon-pro-vega-20.c32... Vega 20 seems to also refer to a discrete GPU. This has been later rebranded to Radeon VII (maybe because of this confusion). The number you are quoting is for the discrete GPU.
- jra101 5y agoHuh, I had no idea they used Vega 20 both as a codename and a product name. Confusing.
- floatboth 5y agoSame with 10 LOL. VEGA10 codename is for the original Vega Frontier Edition/RX Vega 56/RX Vega 64. But now there's also "Vega 10" used as a description for the 10 Vega compute units on Ryzen APUs.
- 6gvONxR4sf7o 5y ago> Just for reference Nvidia's flagship GPU(3090)'s FP32 performance is 35.5 TFLOPS. In context of ML, nvidia’s flagship is the A100, which has 312 TFLOPS. You can also compare with a TPU device which has 180 TFLOPS (v2) or 420 TFLOPS (v3). You can use at least the TPU v2 reliably on colab for free.
- fartcannon 5y agoThe free part is limited to a set number of hours per month.
- codelord 5y agoIt's not really fair to compare a discrete GPU to a mobile GPU, I only provided this as a comparison for someone who maybe has one of these at home. And btw, you are talking about TF32 performance not FP32. TF32 actually uses 16 bits. A100's FP32 performance is actually lower than 3090, it's 19.5 TFLOPS: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia-a100-datasheet.pdf https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent... FP16 performance is also relevant as a lot of people now train in FP16. The default for pytorch/TF is still FP32.
- 6gvONxR4sf7o 5y agoI’ve been mostly using TPUs lately and non-TF32 GPUs at work, so I don’t have any practical experience with TF32, but the sales pitch seems pretty good. Do you have any personal experience on whether it’s as much of a drop in replacement for fp32 as they suggest?
- codelord 5y agoI haven't used TF32 personally, but I think the sales pitch is not too far off. Most of the time I use mixed-precision training which should be similar to FP16/TF32 in terms of performance. It does tremendously speed up training.
- rubatuga 5y agoTF32 is 19 bits! Default for pytorch and TF is TF32!
- orik 5y agoI wanted this for a long time, but eventually gave up on my MacPro5,1 that I had a Radeon VII in (before that, Vega Frontier, before that R9-280X). Would like to see some benchmarks with that sort of hardware or with a 6900XT, driver support which came to MacOS only recently. Now I just have an M1 Mac Mini & a PC with Nvidia.
- ksec 5y agoSo Apple would need 16x its GPU Core, or 128 GPU Core to reach Nvidia 3090 Desktop Performance. Or roughly 480mm2 Die Size, 192W TDP excluding memory controller and interconnect. Doesn't look too bad for Nvidia, especially when you consider 3090 is still on Samsung 8nm, which is equivalent to TSMC 10nm, compared to 5nm on Apple M1.
- Joeri 5y agoThe rumors say they’re going to have a high end of 128 gpu cores by using 4 32 gpu core chiplets.
- dharma1 5y agowouldn't that require pretty hefty active cooling, which doesn't fit so well with Apple devices? It would be great if they did it the eGPU route with TB4, in a slick package
- Tagbert 5y agoThis is more for the Mac Pro line. Not so much the laptops. We’ve only see the Apple equivalent of the i3 with integrated graphics. It’s going to get interesting over the next several months as Apple unveils their middle and upper performance solutions.
- dahfizz 5y agoDoes that mean they will actually produce a laptop with adequate cooling? Might be worth looking at.
- Ly5846N929Rm6bb 5y agoI think the 128 core GPU is rumored for desktop Mac Pro.
- ababaiem 5y agoNo. According to rumours: the high-end 64-128 GPUs will be exclusively for the Mac Pro and the MacBook Pro will have 16-32 cores GPUs instead. It think it's fair to assume that we won't see much changes with the cooling for their laptops.
- truth_ 5y agoPyTorch currently provides RoCM support. Has anyone tried that?
- yza 5y agoI tried it on a Radeon VII. It is little hard to get running as it does not work with latest kernel, but otherwise it kind of works with some quirks. One of the quirks is the first epoch of your training is very slow to start compared to Nvidia cards. Also, the speed of training is slow.
- andy_ppp 5y agoInteresting that AMD is supported, could this mean Apple Silicon 16inch MBP with an AMD GPU?
- defaultname 5y agoA large number of Apple devices in the field, and even still for sale by Apple, have AMD GPUs. It's most likely legacy support driving its inclusion -- if Apple abstracted the pluggable device, appropriately branching it for their existing AMD and new chips seems a given. It's extremely doubtful any new Apple Silicon device will come out with AMD GPUs.