4 ms·
Seconding this. I’ve found that llama-cpp doesn’t completely evict all of the GPU memory with its idle sleep parameter configured.
by gcr 2mo ago
Seconding this. I’ve found that llama-cpp doesn’t completely evict all of the GPU memory with its idle sleep parameter configured.