3 ms·
As you said: everything works on llama.cpp Why it does not work on vllm? Of course you can say that it is AMD fault but there was an issue of abysmal performanc
by npodbielski 2mo ago
As you said: everything works on llama.cpp
Why it does not work on vllm? Of course you can say that it is AMD fault but there was an issue of abysmal performance of models on Strix Halo, that is open for half a year (https://github.com/vllm-project/vllm/issues/34579#issuecomment-5129108179 https://github.com/vllm-project/vllm/issues/34579#issuecomme...) and nothing is happening there. They do not care about those use cases. Seems like they are going with bit players that will run vllm inside datacenters racks. Hobbyists does not matter.