4 ms·
Does it have to be exclusively vram to be useful or is shared ram also useful here. I have a system that reports 8gb dedicated to GPU with another 15 that is sh
by ripvanwinkle 3y ago
Does it have to be exclusively vram to be useful or is shared ram also useful here. I have a system that reports 8gb dedicated to GPU with another 15 that is shared and thus a GPU memory of 23 GB
- brucethemoose2 3y agoDriver shared memory is effectively useless. The way it accesses CPU memory is (for now) extremely slow. First some background: llama is divided into prompt ingestion code, and "layers" for actually generating the tokens. There are different offloading schemes, but what llama.cpp specifically does is map the layers to different devices. For example, ~7GB of the model could reside on the 8GB GPU, and the other ~9GB would live on the CPU. During runtime, the prompt is ingested by the GPU all at once (which doesn't take much VRAM), and then for each word, the layers are run sequentially. So the first ~half of a word would run on your GPU, and the last half would run on your CPU, and the alternation repeats till all the words are generated. The beauty of llama.cpp is that its cpu token generation is (compared to other llama runtimes) extremely fast. So offloading even half or two thirds of a model to a decent CPU is not so bad.
- ripvanwinkle 3y agoThank you for the explanation