6 ms·
Bringing K/V context quantisation to Ollama
- smcleod 2y agoShout out to everyone from Ollama and the wider community that helped with the reviews, feedback and assistance along the way. It's great to contribute to such a fantastic project.
- octocop 2y agoshout out to llama.cpp
- satvikpendem 2y agoWhat's the best way to use Ollama with a GUI, just OpenWebUI? Any options as well for mobile platforms like Android (or, I don't even know if we can run LLMs on the phone in the first place).
- sadeshmukh 2y agoA lot of the UIs, including OpenWebUI have the feature to expose over LAN with users - that's what I did to use my GPU while still being on my phone. Not entirely sure about native UIs though. Also, I normally use Groq's (with a q) API since it's really cheap with no upfront billing info required - it's a whole order of magnitude cheaper iirc than OpenAI/Claude. They literally have a /openai endpoint if you need compatibility. You can look in the direction of Google's Gemma if you need a lightweight open weights LLM - there was something there that I forgot.
- smcleod 2y agoI personally use a mix of Open WebUI, Big AGI, BoltAI, AnythingLLM on the desktop. The mobile space is where things are really lacking at the moment, really I just end up browsing to Open WebUI but that's not ideal. I'd love a iOS native client that's well integrated into Siri, Shortcuts, Sharing etc...
- qudat 2y agoFor hosting a web gui for ollama I use https://tuns.sh https://tuns.sh It really convenient because it's just an SSH tunnel and then you get automatic TLS and it protects your home IP. With that you can access it from your mobile phone, just gotta require a password to access it.
- satvikpendem 2y agoI'm running OpenWebUI via Docker via OrbStack, it also automatically provides TLS and works pretty well.
- minwidth0px 2y agoI wrote my own UI[0] that connects over WebRTC using Smoke.[1] [0] https://github.com/minwidth0px/gpt-playground https://github.com/minwidth0px/gpt-playground and https://github.com/minwidth0px/Webrtc-NAT-Traversal-Proxy-Server https://github.com/minwidth0px/Webrtc-NAT-Traversal-Proxy-Se... [1] https://github.com/sinclairzx81/smoke https://github.com/sinclairzx81/smoke
- deleted 2y ago[deleted]
- huijzer 2y agoI have Open WebUI on a Hetzner instance connected to Deep Infra. Works on mobile by turning the web page into an app. I find the web framework that WebUI uses quite bloated/slow, but apart from that it does work reliably. Price at Deep Infra is typically about $0.04 per month even when actively asking lots of questions during programming.
- paradite 2y agoI built a custom GUI for coding tasks specifically, with built-in code context management and workspaces: https://prompt.16x.engineer/ https://prompt.16x.engineer/ Should work well if you have 64G vRAM to run SOTA models locally.
- deleted 2y ago[deleted]
- throwaway314155 2y ago> Should work well if you have 64G vRAM to run SOTA models locally. Does anyone have this? edit: Ah, it's a Mac app.
- paradite 2y agoYeah Mac eats Windows on running LLMs. My app does support Windows though, you can connect to OpenAI, Claude, OpenRouter, Azure and other 3rd party providers. Just running SOTA LLMs locally can be challenging.
- throwaway314155 2y agoI'm pretty satisfied with my linux nvidia gpu setup. I may not have as much memory on my card, but the speed is almost certainly competitive if not outright faster. Further there are lots of techniques that mitigate this issue like offloading/streaming layers in as needed, quantizing, etc. It also handles actual training/finetuning better.
- gzer0 2y agoM4 Max with 128 GB RAM here. ;) Love it. A very expensive early Christmas present.
- accrual 2y agoGreat looking GUI, I find simple black/white/boxy/monospace UIs very effective.
- rkwz 2y agoIf you’re using a Mac, I’ve built a lightweight native app - https://github.com/sheshbabu/Chital https://github.com/sheshbabu/Chital
- magicalhippo 2y agoAa a Windows user, who just wanted something bare bones for playing, I found this[1] small project useful. It does support multi-modal models which is nice. [1]: https://github.com/jakobhoeg/nextjs-ollama-llm-ui https://github.com/jakobhoeg/nextjs-ollama-llm-ui
- seb314 2y agoFor running llms _locally_ on Android, there's "pocketpal" (~7tok/s on a pixel 7 pro for some quant of llama 3.2 3B). (Not sure if it uses ollama though)
- vunderba 2y agoAs far as open source goes, I'd probably recommend LibreChat. It has connections for ollama, openai, anthropic, etc. It let's you setup auth so you can theoretically use it from anywhere (phone, etc.). Fair warning, it's relatively heavyweight in so far as it has to spin up a number of docker instances but works very well. https://github.com/danny-avila/LibreChat https://github.com/danny-avila/LibreChat
- zerop 2y agoMany are there, apart from what others mentioned I am exploring Anything LLM - https://anythingllm.com/ https://anythingllm.com/. Liked the workspace concept in it. We can club documents in workspaces and RAG scope is managed.
- deleted 2y ago[deleted]
- wokwokwok 2y agoNice. That said... I mean... > The journey to integrate K/V context cache quantisation into Ollama took around 5 months. ?? They incorrectly tagged #7926 which is a 2 line change, instead of #6279 where it was implemented, which made me dig a bit deeper and reading the actual change it seems: The commit (1) is: > params := C.llama_context_default_params() > ... > params.type_k = kvCacheTypeFromStr(strings.ToLower(kvCacheType)) <--- adds this > params.type_v = kvCacheTypeFromStr(strings.ToLower(kvCacheType)) <--- adds this Which has been part of llama.cpp since Dev 7, 2023 (2). So... mmmm... while this is great, somehow I'm left feeling kind of vaguely put-off by the comms around what is really 'we finally support some config flag from llama.cpp that's been there for really quite a long time'. > It took 5 months, but we got there in the end. ... I guess... yay? The challenges don't seem like they were technical, but I guess, good job getting it across the line in the end? [1] - https://github.com/ollama/ollama/commit/1bdab9fdb19f8a8c73ed85291f9acea5bc1c7075#diff-7c8fcee9a6ef35252c34bdc9910b1e605c5d480ea80d9f2fe1c67dc069e9888cR144 https://github.com/ollama/ollama/commit/1bdab9fdb19f8a8c73ed... [2] - since https://github.com/ggerganov/llama.cpp/commit/bcc0eb4591bec5ec02fad3f2bdcb1b265052ea56#diff-201cbc8fd17750764ed4a0862232e81503550c201995e16dc2e2766754eaa57aR907 https://github.com/ggerganov/llama.cpp/commit/bcc0eb4591bec5...
- BowBun 2y agoAuthor describes why it took as long as it did in the post, so I don't think they're trying to be disingenous. Getting minor changes merged upstream in large projects is difficult for newer concepts since you need adoption and support. Full release seems to contain more code[1], and author references the llama.cpp pre-work and that author as well This person is also not a core contributor, so this reads as a hobbyist and fan of AI dev that is writing about their work. Nothing to be ashamed of IMO. [1] - https://github.com/ollama/ollama/compare/v0.4.7...v0.4.8-rc0 https://github.com/ollama/ollama/compare/v0.4.7...v0.4.8-rc0
- smcleod 2y ago> this reads as a hobbyist and fan of AI dev that is writing about their work Bingo, that's me! I suspect the OP didn't actually read the post. 1. As you pointed out, it's about getting the feature working, enabled and contributed into Ollama, not in llama.cpp 2. Digging through git commits isn't useful when you work hard to squash commits before merging a PR, there were a _lot_ over the last 5 months. 3. While I'm not a go dev (and the introduction of cgo part way through that threw me a bit) there certainly were technicalities along the way, I suspect they not only didn't both to read the post, they also didn't bother to read the PR. Also, just to clarify - I didn't even share this here, it's just my personal blog of things I try to remember I did when I look back at them years later.
- lastdong 2y agoGreat project! Do you think there might be some advantages to bringing this over to LLaMA-BitNet?
- smcleod 2y agoToday I ran some perplexity benchmarks comparing F16 and Q8_0 for the K/V, I used Qwen 2.5 Coder 7b as I've heard people say things to the effect of Qwen being more sensitive to quantisation than some other models. Well, it turns out there's barely any increase in perplexity at all - an increase of just 0.0043. Added to the post: https://smcleod.net/2024/12/bringing-k/v-context-quantisation-to-ollama/#perplexity-measurements https://smcleod.net/2024/12/bringing-k/v-context-quantisatio...