4 ms·
llama.cpp works pretty well for me on the Framework 13 laptop, but the current era of "move fast, break things, rarely fix" (sorry, that's how it feels), bites
by imrehg 2mo ago
llama.cpp works pretty well for me on the Framework 13 laptop, but the current era of "move fast, break things, rarely fix" (sorry, that's how it feels), bites here quite a bit.
Two examples:
- https://github.com/ggml-org/llama.cpp/pull/25863 https://github.com/ggml-org/llama.cpp/pull/25863 Someone's few lines change broke the native (ROCm) support for the AMD GPU inside Framework (and other integrated systems), and any rollback or proper fix is pending for almost a month. Fortunately there's workaround (switching to Vulkan rather than ROCm devices), but both the way the bug was introduced and the way it is not fixed just doesn't give much confidencen
- LM Studio is using llama.cpp internally for GGUF, they ship their own build with their closed source system as "runtimes". Their ROCm runtime does not enable the the AMD GPU inside the Framework, even thought the llama.cpp version would support it. So their runtime keeps telling me that there's no supported AMD GPU -- again, the solution is to use the GPU with the Vulkan devices. Not fixed since Jan at least https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1394 https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1...
I guess overall it's the worst runtime I've seen so far, except for all the other runtimes out there... I'm a fan, though in some cases I don't have enough knowledge, or I don't have access to fix things, and that feels like a bummer...
- d3Xt3r 2mo agoSo are there any alternatives which do actually work well with ROCm OOTB?
- imrehg 2mo agoI just switched to Vulkan, and be done with it. :) As much as I can tell, the ROCm version of llama.cpp would be a bit faster on prompt processing, but about the same on the token generation as Vulkan. Real life benchmarks don't seem to give any "ROCm or nothing" sort of vibes. And the difference between the performance of different models are way bigger than the difference between the llama.cpp versions (and versus different runtimes like the llama.cpp/GGUF and the MLX runtimes on Mac for the same models)... I've tinkered enough with the serving, that I'd rather do something with them with, say 10% slower speed, than spending hours on seting things up again... YMMV
- ljosifov 2mo agoHipfire works on my 7900xtx best of all. Biggest surprise - Qwen3.6-27B dense does not grind to a halt with context depths all the way up to 250K! Measured at 1K, 8K, 32K, 64K, 128K, 192K, 250K - Hipfire speed holds close to 40 tok/s. Finished 5 days run of gpu 100% inferencing, Hipfire did not crash even once afaics. I assumed speed dropping like a stone on Qwen models was a feature/bug of the model, and "nothing can be done about it". Hipfire showed me wrong - pleasantly surprised there.
- d3Xt3r 2mo agoThanks, sounds promising. Written in Rust, no PyTorch etc, sounds like my kinda tool. Surprised I haven't seen more folks mention it here. Also if I may ask, what does the rest of your stack look like (agent, harness etc)?
- ljosifov 2mo agoAtm OMP (oh-my-pi) with opencode-go subscription using deepseek-v4-flash-0731 (and mimo-v2.5 as advisor). Used pi before (still good, but too minimalist for all general use; works on Android under Termux btw) and OpenCode (which is ok too). I've used Codex the longest, but the $20 sub doesn't last that long. Before that Claude (and Claude on the web - about Nov'25), but nowadays with the new models not possible to do anything on the cheap sub. Got Zai (GLM-5.2) legacy sub to supplement too for variety. :-) Run local jobs on the 7900xtx (24gb vram - so dense models) batch bulk jobs. Non-dense models like nemotron-cascade-2-30b-a3b MoE run at 100 tok/s, looking forward to try the new lightning-3.5-30b-a3b. And got old MBP M2 Max 96gb uram (unified ram) for bigger sparse MoE models. Testing DS4F there under ds4 server/harness (too slow, 12-15 tok/s even with REAP25 reduced model), and (got agents) porting Ling-3.0-flash atm from llama.cpp -> ds4. Hoping only A5B (5B active, versus 13B of DS4F) will result in better tok/s. TBS. Back in Feb-Apr'26 looked local will be the only way to run llm jobs (plus 300w gpu job works out for office heating when cold), otherwise the ex-quota API payg costs would bankrupt me. But now with the oss weights models and providers like opencode - actually it's not only faster but also cheaper for me to run llm jobs on the api. Engines for local - llama.cpp for all, hipfire (mentioned) for the 7900xtx, and mlx on the apple ASI (that never quite work out for me). And ds4 trying forever when off-line :-) - e.g. on a flight - either ds4-agent, or ds4-server + agent pi. And I use Hermes as a general agent, non-coding agent, for everything that's not strictly coding.
- dbgobrrr 2mo agoLemonade-server works pretty well (most of the time). It wraps llama.cpp and other runtimes - it downloads the official binaries as far as I could see, and you can set alternative versions if needed. Works nicely with Strix Halo for a while now. https://lemonade-server.ai https://lemonade-server.ai
- etdznots 2mo agoThe first one multiple contributors highlighted the PR as urgent andits had lots of review but it appears to be waiting for another review and/or someone that owns the affected hardware to test that the PR fixes the issue, it wpuld be easy for you to test and report whether or not it does, and the second thing is not related to llama.cpp at all Yes ideally there would be testing every hardware + software combo but this costs engineering time and $$$ money, and you are running on master branch, no master branch of any software is stable, inherently, if you run into issues, just stick to the old hash where stuff worked, why are you insistent on both being at the bleeding edge and experience 0 breakage!
- imrehg 2mo agoI did report my test results on the first one. :) The second I didn't say it's any of llama.cpp's "fault", but it is _related_ to llama.cpp since it's being shipped in another system, aye? Can't stick to the old hash either, because older version have different bugs. E.g. on older versions the same Qwen3.6 model reliably fails to call specific tools due to template issues, while just having the newer llama.cpp version has that fixed. So different versions - different bugs, rather than no bugs. Why the beating you are trying to gimme, mate? :)
- etdznots 2mo agoIn some cases it may be that llama.cpp causes a bug or lack of functionality downstream but lack of ROCM support appears to be entirely LM Studio’s doing and unrelated to llama.cpp Sorry if I was too harsh, it’s just that my perception watching the repo has been that the llama.cpp devs are by far the most cautious and slow moving of all the inference implementations, so I found your perspective a bit surprising, I do think that the desire for stable software that never break, and software that supports the latest models and devices/device API’s are conflicting, nothing will do both, and I think that llama.cpp devs do a good job of balancing between shipping features and not breaking users.
- bjackman 2mo agoI think ROCm is just a total second class citizen in the space TBH. It's a shame coz it's not even really what we want, we would obviously all be better served if we could use Vulkan or something. But I guess it's inevitable that a generic framework lags behind here. If I was AMD I'd hire a whole ecosystem team to sit next to the ROCm people and just support big users like llama.cpp to work better on their HW, e.g. giving OSS maintainers access to their board farms. Maybe they have already done that, in which case I guess I should say I'd double the size of that team.
- roenxi 2mo agoYou probably can just use Vulkan. Why don't you think Vulkan is an alternative? Can it not handle BLAS and memory access? What is a graphics card supposed to be doing that isn't efficient in Vulkan?
- binary132 2mo agoSometimes it’s just that a particular backend is poorly optimized or has a regression on a particular platform as compared to the “mainstream” backends. For example, whisper.cpp’s Vulkan backend performs 2-3x worse than CPU on my Snapdragon X2 laptop when using the ggerganov v3 turbo model. It’s probably a simple fix, but it does need to be fixed.
- basedrum 2mo agoI have a framework 13, but I couldn't imagine running a local llm on it, how do you do it? Do you have a eGPU?
- mildred593 2mo agoIf you have AMD, you can benefit from the unified memory.
- lhl 2mo agoFor updated/validated updates, Donato Capitella maintains independent Strix Halo "toolboxes": https://strix-halo-toolboxes.com/ https://strix-halo-toolboxes.com/ A team from AMD maintains Lemonade, another all-in-one setp with convenient installers for setting everything up: https://lemonade-server.ai/ https://lemonade-server.ai/ These are probably better than running against llama.cpp ROCm directly as there are frequent/constant regressions on the main branch, especially for gfx1151 (Strix Halo), but RDNA in general. There are number of AMD-focused llama.cpp forks (nathanw1014, charlie12345, ciru-ai, justinappler, etc) - as well as a few alternatives like hipfire or my hipEngine. While ROCm has gotten a lot better, one of the things I've found after writing an inference engine that has completely custom tuned/fused C++/HIP kernels, is that while it's been pretty straightforward to match/beat llama.cpp ROCm performance, that Vulkan RADV has been a lot harder since RDNA3 support for ROCm has a few issues that make it underperform ACO on some common operations on both gfx1100 and gfx1151 (see: https://github.com/ROCm/ROCm/issues/6409 https://github.com/ROCm/ROCm/issues/6409 ) In general, for anyone just looking to run LLM models on an AMD card, I'd just recommend going with llama.cpp Vulkan and skipping ROCm completely.