4 ms·
Funnily enough, RAM I find I can use no matter how much I have, especially now that you have LLMs. I can run a quantised Mixtral 8x7B on 64GB RAM. But the 7950
by rakejake 3y ago
Funnily enough, RAM I find I can use no matter how much I have, especially now that you have LLMs. I can run a quantised Mixtral 8x7B on 64GB RAM.
But the 7950X is grossly underutilised, I don't know what to do with it.
- lolinder 3y agoShouldn't the quantized LLMs be using the CPU to good effect? I'd imagine your Mixtral is substantially faster than it would be on a weaker CPU.
- rakejake 3y agoI have a 4080 with 16GB of VRAM. I experimented with llama.cpp by offloading layers onto GPU and doing the remaining on CPU. I found that it gives me max tokens/sec if I set it to 8 CPU cores as opposed to the 16 available on the 7950X. I guess beyond that, the bookkeeping between the cores might be taking up more time than it is worth.
- mischief6 3y agothe answer i found is run gentoo and compile yocto projects.