3 ms·
Have you tried running it on a single 5090? Dual 5090 require https://github.com/aikitoria/open-gpu-kernel-modules https://github.com/aikitoria/open-gpu-kernel-
by xfalcox 2mo ago
Have you tried running it on a single 5090? Dual 5090 require https://github.com/aikitoria/open-gpu-kernel-modules https://github.com/aikitoria/open-gpu-kernel-modules for higher perf. Are you using TP?
- trouve_search 2mo agoYes, I mentioned the setup, but on vllm you can only use TP with speculative decoding or pipeline parallelism without, so there's tradeoff to both. I gave general numbers of what I'm getting above, the performance ratios seemed similar regardless of setup (eg. getting a AWQ-in4 quant on a single GPU vs PP without speculative decoding vs TP with speculative decoding). Overall single GPU is fastest, and TP+speculative decoding is still faster than PP, but for fp8 models you need dual GPUs whether you want it or not.