7 ms·
If AMD developed CUDA transpilers / DL backends / CUDNN equivalents people would switch over in a heartbeat. It's fully their fault for not investing in the sof
by TTPrograms 8y ago
If AMD developed CUDA transpilers / DL backends / CUDNN equivalents people would switch over in a heartbeat. It's fully their fault for not investing in the software.
If they did that and had a card that got 2x performance / $ or more I would switch in a heartbeat.
- JoshuaJB 8y agoAMD has a CUDA Transpiler (https://github.com/ROCm-Developer-Tools/HIP https://github.com/ROCm-Developer-Tools/HIP) and a multitude of other open source libraries through GPUOpen ( https://gpuopen.com/professional-compute/ https://gpuopen.com/professional-compute/). The quality and open source nature of their tools has resulted in much of my research group (real-time vision) increasingly and voluntarily moving to work on AMD platforms (we were previously almost exclusively using CUDA).
- deleted 8y ago[deleted]
- ratsimihah 8y agoBut the GPU needs to support ROCm, which is quite recent right? e.g. The MBP 2015's Radeon m370x doesn't seem to support it.
- gcb0 8y agoless recent than a 1080 or 2080ti
- ratsimihah 8y ago2080 Ti is brand new :/
- auslander 8y agoHere is it: https://github.com/ROCm-Developer-Tools/HIP https://github.com/ROCm-Developer-Tools/HIP Also, AMD does not limit FP performance on consumer cards.
- twtw 8y ago> Also, AMD does not limit FP performance on consumer cards. I don't understand this meme. The consumer cards are different chips with slow fp64 hardware. In what sense is that "limiting" performance relative to the enterprise cards?
- auslander 8y agoOn AMD, if I'm not not mistaken, FP16 is twice of FP32, as it should be "For their consumer cards, NVIDIA has severely limited FP16 CUDA performance. GTX 1080’s FP16 instruction rate is 1/128th its FP32 instruction rate, or after you factor in vec2 packing, the resulting theoretical performance (in FLOPs) is 1/64th the FP32 rate, or about 138 GFLOPs." https://www.anandtech.com/show/10325/the-nvidia-geforce-gtx-1080-and-1070-founders-edition-review/5 https://www.anandtech.com/show/10325/the-nvidia-geforce-gtx-... "FP16 performance is 1/64th and FP64 is 1/32th of FP32 performance." https://blog.inten.to/hardware-for-deep-learning-part-3-gpu-8906c1644664 https://blog.inten.to/hardware-for-deep-learning-part-3-gpu-...