12 ms·
Now all we need is better support for AMD gpus, both CDNA and RDNA types
by androiddrew 9mo ago
Now all we need is better support for AMD gpus, both CDNA and RDNA types
- mappu 9mo agoZLUDA implements CUDA on top of AMD ROCm - they are explicitly targetting vLLM as their PyTorch compatibility test: https://vosen.github.io/ZLUDA/blog/zluda-update-q4-2025/#pytorch-support-underway https://vosen.github.io/ZLUDA/blog/zluda-update-q4-2025/#pyt... (PyTorch does also support ROCm generally, it shows up as a CUDA device.)
- sofixa 9mo agoYou can run vLLM with AMD GPUs supported by ROCm: https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference/deploy-your-model.html https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/infer... However from experience with an AMD Strix Halo, a couple of caveats: it's drastically slower than Ollama (tested over a few weeks, always using the official AMD vLLM nightly releases), and not all GPUs were supported for all models (but that has been fixed).
- bildung 9mo agovLLM ususally only plays out its strength when serving multiple users in parallel, in contrast to llama.cpp (Ollama is a wrapper around llama.cpp). If you want more performance, you could try running llama.cpp directly or use the prebuilt lemonade nightlies.
- sofixa 9mo agoBut vLLM was half the t/s of Ollama, so something was obviously not ok.