3 ms·
If anyone else is running this on an RTX 5090, https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer as inference engine gets me ~138 tokens/seco
by kimsey0 2mo ago
If anyone else is running this on an RTX 5090, https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer as inference engine gets me ~138 tokens/second, roughly double what I get with a naive llama.cpp setup.
- wincy 2mo agoAmazingly enough, Hacker News decided to show this to me as the top comment, I'm also running an rtx 5090 and trying it out now, thanks for the tip! Edit: Absolutely blazing fast! Getting 163 tokens/sec on WSL and it generated a pretty sweet Pelican. https://gist.github.com/hansale/ed9e73fe35165a58ea2af6b1632afdb5 https://gist.github.com/hansale/ed9e73fe35165a58ea2af6b1632a...
- sgt 2mo agoAmazing. I'll give it a shot on my 5090. I already tried using vLLM but it ran out of GPU memory. I guess it's likely Llama.cpp will work.
- deleted 2mo ago[deleted]
- searealist 2mo agoJust enable MTP on llama.cpp and you will get the same decode speeds.
- pulse7 2mo agoIt there anything similar for RTX 3090 and RTX 4090?
- mmlkrx 2mo agoI'm not sure about a 4090 but there is a fork for 3090s: https://github.com/Don-Chad/ninfer-3090 https://github.com/Don-Chad/ninfer-3090
- mirekrusin 2mo ago[dead]
- ixaxaar 2mo agoI'm running it using a 4090 on using llama.cpp with Q5_K_S and its running at ~33 t/s
- glinkot 2mo agoYep, same, testing it now and it flies!