3 ms·
What is your setup?
by manx 2y ago
What is your setup?
- aappleby 2y agoMinisforum BD790i with 96 gigs of LPDDR5 RAM running Ollama. I run Qwen 110b on the CPU and get around a token per second, which is enough for me to dash off a question and get an answer a few minutes later. Smaller models like CodeLlama that can run 100% on the GPU are 10x faster, but I've found that any model under 70b params makes stupid errors when asked more complex questions.