4 ms·
Qwen3.5 9b seems to be fairly competent at OCR and text formatting cleanup running in llama.cpp on CPU, albeit slow. However, I have compiled it umpteen ways an
by Curiositry 7mo ago
Qwen3.5 9b seems to be fairly competent at OCR and text formatting cleanup running in llama.cpp on CPU, albeit slow. However, I have compiled it umpteen ways and still haven't gotten GPU offloading working properly (which I had with Ollama), on an old 1650 Ti with 4GB VRAM (it tries to allocate too much memory).
- WhyNotHugo 7mo agoIf you’re building from source, the vulkan backend is the easiest to build and use for GPU offloading.
- Curiositry 7mo agoYes, that's what I tried first. Same issue with trying to allocate more memory than was available.
- acters 7mo agoI have a 1660ti and the cachyos + aur/llama.cpp-cuda package is working fine for me. With about 5.3 GB of usable memory, I find that the 35B model is by far the most capable one that performs just as fast as the 4B model that fits entirely on my GPU. I did try the 9B model and was surprisingly capable. However 35B still better in some of my own anecdotal test cases. Very happy with the improvement. However, I notice that qwen 3.5 is about half the speed of qwen 3
- lioeters 7mo ago> GPU offloading working I had this issue which in my case was solved by installing a newer driver. YMMV. sudo apt install nvidia-driver-570
- dunb 7mo agoAre you running with all the --fit options and it’s not working correctly? You could try looking at how many layers are being attempted to offload and manually adjust from there. Walk down --n-gpu-layers with a bash script until it loads.
- AllegedAlec 7mo agoI found that the drivers I had were no longer compatible with the newer kernels. After upgrading to newer drivers it was able to offload again.