2 ms·
Then you might be missing SWA. Gemma models are extremely memory hungry without
by jodleif 2mo ago
Then you might be missing SWA. Gemma models are extremely memory hungry without
- CMay 2mo agoSo long as they have flash attention enabled, Llama.cpp enables Sliding Window Attention by default for Gemma 4 models. Even if they're using Ollama or LM Studio I would expect those to mostly be doing the right things.
- DiabloD3 2mo agoI would not expect Ollama to be doing the right thing fwiw.