2 ms·
I have a 4080 with 16GB of VRAM. I experimented with llama.cpp by offloading layers onto GPU and doing the remaining on CPU. I found that it gives me max tokens
by rakejake 3y ago
I have a 4080 with 16GB of VRAM. I experimented with llama.cpp by offloading layers onto GPU and doing the remaining on CPU. I found that it gives me max tokens/sec if I set it to 8 CPU cores as opposed to the 16 available on the 7950X. I guess beyond that, the bookkeeping between the cores might be taking up more time than it is worth.