3 ms·
I have been using ROCm for 2y+. The investment in this infrastructure was a big mistake. The biggest burner was the need to do a clean install on each new ROCm
by zmachinaz 5y ago
I have been using ROCm for 2y+. The investment in this infrastructure was a big mistake. The biggest burner was the need to do a clean install on each new ROCm release. Clean here means manually finding and deleting all traces from the previous ROCm version, and recompilation of apps like pytorch. Good upgrades took hours, bad ones days ... . Finally I settled to freeze the system and not touch it anymore until retirement of the cards, hopefully soon.
- supernovae 5y agothis is a problem with nvidia too.. i just made all my infra easy to reprovision and start clean and workloads ran as containers..
- jacquesm 5y agoInteresting, can you describe this in a bit more detail? It runs completely counter to my experience, so far NVidia for me has just been a long string of 'boring' in that it just works. Even applications written for older cards and older versions of CUDA have continued to work just fine.
- supernovae 5y agoIt was so bad, we just moved to immutable GPU infrastructure regardless of physical or virtual. When a new release of all the nvidia stuff comes out, we re-image the machine and install it. Cuda on linux with ml/gpu workloads is still kind of a hotmess and i'd say we're far from finding a winner like some suggest here. It's gotten better... but still far easier to treat it like a mess and start fresh with any install
- rocmasdf 5y agoCould you describe what is involved in installing a ROCm release from scratch? (I've mostly stayed in CUDA-land, but I'm curious and intrigued about the idea of ROCm, and AMD GPUs, though your experience suggests there is much room for improvement...)