2 ms·
This is probably due to this: https://github.com/ggml-org/llama.cpp/issues/12637 https://github.com/ggml-org/llama.cpp/issues/12637. This GitHub issue is about
by tbocek 1y ago
This is probably due to this: https://github.com/ggml-org/llama.cpp/issues/12637 https://github.com/ggml-org/llama.cpp/issues/12637. This GitHub issue is about interleaved sliding window attention (iSWA) not available in llama.cpp for Gemma 3. This could reduce the memory requirements a lot. They mentioned for a certain scenario, going from 62GB to 10GB.
- nolist_policy 1y agoOllama supports iSWA.
- parched99 1y agoResolving that issue, would help reduce (not eliminate) the size of the context. The model will still only just barely fit in 16 GB, which is what the parent comment asked. Best to have two or more low-end, 16GB GPUs for a total of 32GB VRAM to run most of the better local models.