3 ms·
Amazingly enough, Hacker News decided to show this to me as the top comment, I'm also running an rtx 5090 and trying it out now, thanks for the tip! Edit: Abso
by wincy 2mo ago
Amazingly enough, Hacker News decided to show this to me as the top comment, I'm also running an rtx 5090 and trying it out now, thanks for the tip!
Edit: Absolutely blazing fast! Getting 163 tokens/sec on WSL and it generated a pretty sweet Pelican.
https://gist.github.com/hansale/ed9e73fe35165a58ea2af6b1632afdb5 https://gist.github.com/hansale/ed9e73fe35165a58ea2af6b1632a...
- sgt 2mo agoAmazing. I'll give it a shot on my 5090. I already tried using vLLM but it ran out of GPU memory. I guess it's likely Llama.cpp will work.