3 ms·
Not sure why it's on the front page now, but I highly recommend using llama.cpp for running AI model locally vs using other inference framework, unless you have
by karimf 2mo ago
Not sure why it's on the front page now, but I highly recommend using llama.cpp for running AI model locally vs using other inference framework, unless you have a very specific requirement.
ggerganov and the team have done a stellar job maintaining the quality while still being fast to implement new models/improvements.
- walrus01 2mo agoAt this point the options are llama-server or vLLM if you're serious about running things at your desk in the under 256GB RAM size class (70B, 120B size models). In addition to, of course, 27B to 35B size things. With of course a ton of compile time build customization options for whatever specific hardware platform you want to run either llama or vllm on.
- embedding-shape 2mo ago> At this point the options are llama-server or vLLM Which last time I checked, both use different formats of the weights, the former GGUF while the latter .safetensors. I mostly end up using vLLM these days and I'm a bit more performance sensitive than what I used to be. Just a shame it's a hassle to share the weights between them with conversion and what not, either batched or on-startup.
- itake 2mo agodoes your comment depend on the OS? I thought MLX has better performance on MacOS than llama.cpp
- quantumleaper 2mo agoThe gap was MUCH larger in the past, but in my tests, oMLX and llama.cpp are now very similar (within 10%) in both prompt processing and generation speed. GGUF ecosystem provides a better selection of quants, in my experience Unsloth ones are excellent.
- MrScruff 2mo agoI thought the main advantage of oMLX is it's less likely to invalidate the KV cache when working with coding agents, which is key when working on a Mac because of the slower prompt processing.
- quantumleaper 2mo agollama-server also supports saving the kv cache to SSD. I had no issues with cache invalidation using pi.
- gcr 2mo agoTIL! When was this functionality added? It wasn’t in llamacpp when I looked in June
- markasoftware 2mo agopossibly hitting front page because this website is fairly new? For me, it's certainly the first time I've seen a one-liner curl|bash installer for llama.cpp, which was basically the only reason to use ollama.
- aniceperson 2mo agoI think it is due to the new website? it now looks like every other vibe coded site,the only upside is that is looks more saleable for people unfamiliar with it, e.g., explaining OSPO,IT the stack you are using. they should also add a pricing page for eenterprise where they promise 99.9% uptime for local models*.
- chamomeal 2mo agoWow it’s aggressively vibe coded. Nothing inherently wrong with that, but it looks a bit amateurish which is funny. I’m still waiting on 98.css to become the standard for vibe coded sites. You don’t have to read docs anyway if you’re just using LLMs! All you have to do is say “use 98.css” and you have a 10/10 site https://jdan.github.io/98.css/ https://jdan.github.io/98.css/
- dustypotato 2mo agoWow, gonna use that. Thanks
- trouve_search 2mo agoUsing 98.css would still leave you with the AI slop text wording. The core problem is that some people don't even seem to notice / care.
- LoganDark 2mo agoVanilla llama.cpp leaves a lot of performance on the table. I'm reaching 120 t/s with a custom inference engine for a model that llama.cpp can barely run at 70 t/s. Theoretical maximum on this hardware is around 147 t/s according to measured memory bandwidth.
- mirekrusin 2mo agoJust run /goal to optimise it and you should be good in less than an hour. Also best to use models that support speculative decoding.
- LoganDark 2mo agoOptimize llama.cpp? Hmm. WRT speculative decode, basically zero finetunes keep it. I'm testing with some ridiculous abliterated amalgamation so spec decode has been gone for most of its ancestry. Fable recommended n-gram speculation so I'm working on that now.
- mirekrusin 2mo agoOptimize startup params for llama-server for your hardware (not llama.cpp itself), on my 2x 4090 I got ~20% speedup after maybe 40 mins. ps. ngram didn't work for me very well, but dedicated speculative model works very well ps. 2. in my case I'm just maintaining Makefile that does everything from update/upgrade (git pull/recompile) to starting server with different models, stuff like: # over baseline at temp 0.6 (95 vs 45 tok/s), ~4x over naive layer-split baseline. Qwen3.6-27B-MTP-UD-Q8_K_XL: ./llama.cpp/llama-server \ -hf unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q8_K_XL \ --no-mmproj \ --parallel 1 \ --kv-unified \ --flash-attn on \ --fit off \ --split-mode tensor \ -ngl 99 \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --host 0.0.0.0 \ --tools all \ --jinja \ --ctx-size 262144 \ --spec-type draft-mtp \ --spec-draft-n-max 6 \ --temp 0.6 \ --top-p 0.95 \ --top-k 20 \ --min-p 0.0 \ --repeat-penalty 1.0 \ --presence-penalty 1.1 \ --threads 8 \ --reasoning-budget 2048 \ --reasoning on \ --chat-template-kwargs '{"preserve_thinking": true}' \ --reasoning-budget-message "reasoning budget consumed, time to answer now" ... Qwen: Qwen3.6 Qwen3.6: Qwen3.6-27B Qwen3.6-35B-A3B: Qwen3.6-35B-A3B-MTP Qwen3.6-35B-A3B-MTP: Qwen3.6-35B-A3B-MTP-UD-Q8_K_XL Qwen3.6-27B: Qwen3.6-27B-MTP Qwen3.6-27B-MTP: Qwen3.6-27B-MTP-UD-Q8_K_XL
- synergy20 2mo agotrue, switched from ollama to llama.cpp these days and it's good. wonder if this is also the best option for edge ai deployment(currently use it on desktop)