4 ms·
Two things: 1. Because the support in llama.cpp is horizontal integrated within ggml ecosystem, we can optimize it to run even faster than ollama. For example
by ngxson 1y ago
Two things:
1. Because the support in llama.cpp is horizontal integrated within ggml ecosystem, we can optimize it to run even faster than ollama.
For example, pixtral/mistral small 3.1 model has some 2D-RoPE trick that use less memory than ollama's implementation. Same for flash attention (which will be added very soon), it will allow vision encoder to run faster while using less memory.
2. llama.cpp simply support more models than ollama. For example, ollama does not support either pixtral or smolvlm
- danielhanchen 1y agoBy the way - fantastic work again on llama.cpp vision support - keep it up!!
- ngxson 1y agoThanks Daniel! Kudos for your great work on quantization, I use the Mistral Small IQ2_M from unsloth during development and it works very well!!
- danielhanchen 1y ago:)) I did have to update the chat template for Mistral - I did see your PR in llama.cpp for it - confusingly the tokenizer_config.json file doesn't have a chat_template, and it's rather in chat_template.jinja - I had to move the chat template into tokenizer_config.json, but I guess now with your fix its fine :)
- ngxson 1y agoOhhh nice to know! I was pretty sure that someone already tried to fix the chat template haha, but because we also allow users to freely create their quants via the GGUF-my-repo space, I have to fix the quants produces from that source
- danielhanchen 1y agoGlad it all works now!
- roger_ 1y agoWon’t the changes eventually be added to ollama? I thought it was based on llama.cpp
- diggan 1y agoAs far as I understand (not affiliated, just a user who peeked at the code), Ollama started out using llama.cpp as a runner for everything. But eventually they wrote their own runner in Golang, which is where they add support for new models. So most models you run via Ollama uses llama.cpp, but new stuff their own Golang runner.
- nolist_policy 1y agoOn the other hand ollama supports iSWA for Gemma 3 while llama.cpp doesn't. iSWA reduces kv cache size to 1/6.
- vlovich123 1y agoWhat’s iSWA? Can’t find any reference online
- nolist_policy 1y agointerleaved sliding window attention
- imtringued 1y agoGemma 3 has some layers with a context size of 1024 tokens and others having full length. You need to read the Gemma technical report.