3 ms·
You seem to be saying AMD GPUs can run PyTorch but can't run GEMM? Can you explain? I thought PyTorch used GEMM extensively. I also don't understand the commen
by fancyfredbot 1y ago
You seem to be saying AMD GPUs can run PyTorch but can't run GEMM? Can you explain? I thought PyTorch used GEMM extensively.
I also don't understand the comment on "whatever ROCm can't do doesn't seem to be important because no-one has articulated why they need it in my line of sight". Isn't the problem with ROCm the lack of support? It's only officially supported on a tiny proportion of AMD's product line?
- benreesman 1y agoParent seems to be saying that stability issues on consumer RDNA cards are the issue as opposed to ROCm support in PyTorch.
- imtringued 1y agoMy pet theory is that the scheduler can't handle desktop graphics + ML workloads simultaneously, which leads to a deadlock in the firmware.
- compsciphd 1y agofor those who are using it for ML loads, what's the point of using it for desktop graphics (at least at the same time). i.e. I'd argue that unless one is getting server chips (which negates the desktop graphics comment), it seems the vast majority of modern CPUs come with iGPUs that are sufficient for running a desktop environment. Unless one is planning to game on the same machine (but again, probably also not at the same time), if the above is the problem, why not use the iGPU for the desktop and use the dGPU for your ML workloads.
- DiabloD3 1y agoYour pet theory needs work. 1) I have no issues throwing ROCm (which still uses its own path in the driver) at my Radeon at the same time I'm hammering it with "normal" path APIs (Legacy D3D, D3D12, OpenGL, Vulkan, etc), they are scheduled normally and compete for GPU resources normally. 2) ROCm memory allocation is weird in the driver. I have gotten my GPU to hardlock the entire system by allocating about 2x my VRAM, because I suspect its misusing/overusing mprotect().
- skirmish 1y agoThat matches my experience, local LLMs or diffusion models would lock up the GPU when VRAM allocated was close to the maximum available (as monitored by nvtop). After decreasing the batch size or reducing the number of layers offloaded to the GPU, the same workload would run stable for hours.
- deleted 1y ago[deleted]