6 ms·
llama.cpp
- billybobbildo 2mo ago[dead]
- tosh 2mo agoI was a bit suspicious of the url but it is also listed on llama.cpp github https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp
- deleted 2mo ago[deleted]
- dlcarrier 2mo agoI tried to run in on my Arc A770, but all of the binary releases I could find were compiled without OpenVINO support enabled. I tried compiling it myself, but after two days of the compiler running it failed.
- madushan1000 2mo agoTwo days sounds like a lot, both llama.cpp and openvino only takes a few minutes to compile on any decent modern cpu.
- walrus01 2mo agothe full set of llama.cpp binaries builds in under 5 minutes with an unmodified build workflow straight from their github page on a literally ten year old dual xeon.
- gnull 2mo agoIt also even easier to get working and integrate into your system in a sustainable manner with NixOS. Do it by hand or throw an LLM at it, it will get you a declarative patch for your NixOS config that brings llama-cpp into your config that you can review and add under version control (no random `make install` build artifacts contaminating your system, no wondering "what was it that I ran? what are all these files? how do I do the same with a newer version?" a couple months later). There's also likely some build cache where Nixoids have already build what you want. I had a great experience with llama-cpp with Nvidia backend on NixOS. (Sorry for being that guy.)
- numpad0 2mo agoI don't know how it works, but does Vulkan not work on Arc?
- cptskippy 2mo agoDoes the A770 use the Xe driver? If so then it might work with the scripts that I've been using to build llama.cpp with SYCL support for the Arc Pro B70. https://github.com/cptskippy/battlemage-llm-gateway https://github.com/cptskippy/battlemage-llm-gateway It's designed so that you can re-run the scripts to pull the latest updates. When Muse Glimmer was released the other day I just ran the 02 script to build the latest version of llama.cpp with support for it.
- dlcarrier 2mo agoThanks for the link. I'm running an Arch-based system, by the way, so I won't be able to directly use the scripts, which are expecting Debian's package manager, but I'll look at what they're doing and see if I can get something working.
- OpenArcBob 2mo ago[dead]
- deleted 2mo ago[deleted]
- walrus01 2mo agoAnything that suggests curl into bash just plain sketches me out. (edit: I know, this isn't totally rational, it just seems weird to me. We download and trust a lot of software and run code from a bunch of package repositories as a regular activity...). Git clone llama.cpp and build it, it's not hard. https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md https://github.com/ggml-org/llama.cpp/blob/master/docs/build... literally just a few steps for the basics: git clone https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp cmake -B build cmake --build build --config Release
- ur-whale 2mo ago> Anything that suggests curl into bash just plain sketches me out. Yeah, 100% and it's becoming more and more of a thing, see rust install for example. OTOH, if you're installing llama.cpp, you're more than likely planning to run an LLM on your Linux box with an agentic harness, so a curl into bash thing might be the least of your security concerns, :-)
- walrus01 2mo agoOne way I prevent possible catastrophic fuckups is that the 'doing code work' box that runs opencode or pi or whatever, is its entirely own separate VM and desktop environment (running as a xen or kvm guest and with its own LVM logical volume as boot/root and /home disk), than the machine running llama-server itself. The harness gets the openai-compatible endpoint fed into it to talk to llama-server across the network, but the VM has no access whatsoever to my personal files, mail, backups/deep storage, fileserver, Documents folder, etc.
- jakkos 2mo ago> security concerns Yeah I recently tried the coding harness that's recommended here, Pi, in a bubble wrap sandbox and was horrified to learn that it spams multiple warnings at you if you don't give it write access to its own config/extension folder... Everyone else is rawdogging it I guess.
- 2mo ago
- nexawave-ai 2mo agoI think I can probably run Gemma 3 12B on my macbook M3 pro with 18GB. The question is, should I do it? This small model is probably not capable of doing a lot or advanced coding or reasoning. What else could it be used for, since it can run locally and privately?
- kennywinker 2mo agoA quantized Qwen3.6-35B-a3b can run in a similar footprint to gemma 12b, but is smarter. It can do coding tasks, if you specify them at a finer-grain than with bigger models.
- vehemenz 2mo agoContext limit is far too small to do anything serious, tbh.
- whateveracct 2mo ago[flagged]
- bhouston 2mo agoIt seems that llama.app is a direct competitor to ollama.com I can understand the desire for the llama.cpp project to want to own the end user relationship, it is true that previous to this they were a tool provider and not really owning the end user experience.
- ryan_glass 2mo agoOllama uses the llama.cpp backend for inference. I find Ollama noticably slower. Llama.cpp has had a built-in webui (used as llama-server) for a long time now so have owned the user experience too.
- fmajid 2mo agoAnd ollama were sketchy about not providing proper credit to llama.cpp, even though that’s all they are, a wrapper for it.
- pinkmoonx 2mo ago[dead]
- karimf 2mo agoNot sure why it's on the front page now, but I highly recommend using llama.cpp for running AI model locally vs using other inference framework, unless you have a very specific requirement. ggerganov and the team have done a stellar job maintaining the quality while still being fast to implement new models/improvements.
- walrus01 2mo agoAt this point the options are llama-server or vLLM if you're serious about running things at your desk in the under 256GB RAM size class (70B, 120B size models). In addition to, of course, 27B to 35B size things. With of course a ton of compile time build customization options for whatever specific hardware platform you want to run either llama or vllm on.
- embedding-shape 2mo ago> At this point the options are llama-server or vLLM Which last time I checked, both use different formats of the weights, the former GGUF while the latter .safetensors. I mostly end up using vLLM these days and I'm a bit more performance sensitive than what I used to be. Just a shame it's a hassle to share the weights between them with conversion and what not, either batched or on-startup.
- itake 2mo agodoes your comment depend on the OS? I thought MLX has better performance on MacOS than llama.cpp
- quantumleaper 2mo agoThe gap was MUCH larger in the past, but in my tests, oMLX and llama.cpp are now very similar (within 10%) in both prompt processing and generation speed. GGUF ecosystem provides a better selection of quants, in my experience Unsloth ones are excellent.
- MrScruff 2mo ago
- helsinkiandrew 2mo agoI'm confused, is this from Meta? There's no attribution anywhere. Surely releasing an AI tool called llama breaks their trademark if not
- reverius42 2mo agoIt's from https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp -- not associated with Meta, it's been around for years, and surely they know about it -- so I would guess either it's not a trademark violation or they don't care.
- HelloUsername 2mo ago> not associated with Meta, it's been around for year This post (https://news.ycombinator.com/item?id=35100086 https://news.ycombinator.com/item?id=35100086) from march 2023 says in the title "Llama.cpp: Port of Facebook's LLaMA model in C/C++"
- reverius42 2mo agoBy "not associated with Meta" meant, as far as I know, the authors don't work at Meta -- not that llama.cpp is unrelated to Meta's llama model.
- pinkmoonx 2mo ago[flagged]
- imrehg 2mo agollama.cpp works pretty well for me on the Framework 13 laptop, but the current era of "move fast, break things, rarely fix" (sorry, that's how it feels), bites here quite a bit. Two examples: - https://github.com/ggml-org/llama.cpp/pull/25863 https://github.com/ggml-org/llama.cpp/pull/25863 Someone's few lines change broke the native (ROCm) support for the AMD GPU inside Framework (and other integrated systems), and any rollback or proper fix is pending for almost a month. Fortunately there's workaround (switching to Vulkan rather than ROCm devices), but both the way the bug was introduced and the way it is not fixed just doesn't give much confidencen - LM Studio is using llama.cpp internally for GGUF, they ship their own build with their closed source system as "runtimes". Their ROCm runtime does not enable the the AMD GPU inside the Framework, even thought the llama.cpp version would support it. So their runtime keeps telling me that there's no supported AMD GPU -- again, the solution is to use the GPU with the Vulkan devices. Not fixed since Jan at least https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1394 https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1... I guess overall it's the worst runtime I've seen so far, except for all the other runtimes out there... I'm a fan, though in some cases I don't have enough knowledge, or I don't have access to fix things, and that feels like a bummer...
- d3Xt3r 2mo agoSo are there any alternatives which do actually work well with ROCm OOTB?
- imrehg 2mo agoI just switched to Vulkan, and be done with it. :) As much as I can tell, the ROCm version of llama.cpp would be a bit faster on prompt processing, but about the same on the token generation as Vulkan. Real life benchmarks don't seem to give any "ROCm or nothing" sort of vibes. And the difference between the performance of different models are way bigger than the difference between the llama.cpp versions (and versus different runtimes like the llama.cpp/GGUF and the MLX runtimes on Mac for the same models)... I've tinkered enough with the serving, that I'd rather do something with them with, say 10% slower speed, than spending hours on seting things up again... YMMV
- 2mo ago
- prologic 2mo agoIs llama.cpp (and thus llama.app) really that much better than Ollama? I've Only ever played with Ollama, so geniously curious to hear other's real-world experiences.
- hhh 2mo agoollama uses llama.cpp
- hnfong 2mo agoThe dev behind ollama is adamant that ollama doesn't use llama.cpp (based on a technicality -- it uses ggml, which is created by the same people behind llama.cpp and is the backend of llama.cpp) He made such a big fuss about ollama implementing their own kernels and felt slighted about the online comments saying ollama didn't properly credit llama.cpp and it kind of left a bad taste in the mouth among the local inference community. For me personally, it was this that made me avoid them at all costs: https://github.com/ollama/ollama/issues/11714#issuecomment-3172893576 https://github.com/ollama/ollama/issues/11714#issuecomment-3...
- pinkmoonx 2mo ago[dead]
- HelloUsername 2mo ago"Friends don't let friends use ollama" https://sleepingrobots.com/dreams/stop-using-ollama/ https://sleepingrobots.com/dreams/stop-using-ollama/
- pinkmoonx 2mo ago[flagged]
- broodbucket 2mo agollama.cpp is better and it's what everyone switches to after dipping their toes in with ollama, I don't get your point
- mojo-10 2mo ago[flagged]
- pplonski86 2mo agoYesterday I installed llama.cpp to test it with local AI Data Analyst that I'm building. I was also testing other open LLM providers: Ollama, Jan, vLLM, LM Studio. I had older NVIDIA card (RTX 3070) and llama.cpp instalation was smooth, contrary to vLLM which required me to reinstall CUDA drivers because by default it installed the latest one. I'm curious if there is a speed difference between the same open LLM model served with different runners.
- chii 2mo ago> I had older NVIDIA card (RTX 3070) and llama.cpp instalation was smooth what model was it that you were able to run with the rtx 3070?
- pplonski86 2mo agoI was able to fit only small models Qwen3.5-4B in RTX3070 which is not very useful for Python and SQL generation thought. When I wan to test larger open LLM models I often just use cloud resources.
- walrus01 2mo agoJust FYI lm-studio is a GUI wrapper on top of a copy of llama-server that the lm-studio developers compile and distribute
- hypfer 2mo agoOld news by now, but you might not be aware that llama-server can do multi-model for a while now, Meaning that you (and by that I mean your AI agent that has read the llama.cpp code) can write an ini file pointing to your models with parameters optimized for the specific model on your specific hardware. (Optimized by you through testing. Not that AI) Then, any api client can just select a model and the system does the right thing. It's great software. It just works. __ You just need to ignore the cargo culting commandline options on social media. But you should be listening to the devs. Have you already enabled ngram-mod (or rather just spec-default)? It is practically free.
- mrighele 2mo ago> (Optimized by you through testing. Not that AI) Why not optimized by AI through testing ? Give it a test set to work on and let it loose.
- LoganDark 2mo agoAI doesn't necessarily know what feels like a good tradeoff to you. I'm sure it could help guide you though.
- hypfer 2mo ago^ This. Intent is the answer and AI has none.
- Zetaphor 2mo agoThis is all quantifiable. I regularly have my model run benchmarks against all the config permutations and then choose the best based on my criteria, which typically boil down to trading prefill and decode times
- LoganDark 2mo agoWhat model has to trade between those? I have both. You just have different, independently-optimized forward passes for each.
- antonvs 2mo ago> No telemetry Must be tough not to be able to monitor your own models! (The odds that that tagline was AI-generated seem high.)
- tancop 2mo ago[dead]
- halyconWays 2mo agollama.cpp is like the ffmepg of AI, and one of the reasons I so greatly dislike ollama is that the latter completely obfuscates that they're a rebrand of the former. Georgi Gerganov and team did all the hard work; ollama is langchain-like VC-bait with a HF download wrapper.
- kennywinker 2mo agollama.cpp will happily download models from hugging face, btw.
- woadwarrior01 2mo agoHuggingface now owns llama.cpp, btw. https://huggingface.co/blog/ggml-joins-hf https://huggingface.co/blog/ggml-joins-hf
- deleted 2mo ago[deleted]
- woadwarrior01 2mo agoICYMI, llama.cpp was also VC funded. Search for "ggml" on this page: https://aigrant.com https://aigrant.com
- larodi 2mo agoThis site seems scam for not noting origins of llama.cpp and fails to quickly and clearly communicate it NOT being affiliated with GGML org.
- gr_norm 2mo agoFrom https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp: > Visit https://llama.app https://llama.app and follow the instructions It's linked at the start of the README.
- blahblaher 2mo agoand? whats the point of this? Doesn't everyone already know about llama.cpp?
- Maxious 2mo agono, if everyone knew about llama.cpp, ollama would be done already
- HelloUsername 2mo agoThe last time this was posted was in May 2026 (https://news.ycombinator.com/item?id=48325941 https://news.ycombinator.com/item?id=48325941), and before that March 2023 (https://news.ycombinator.com/item?id=35100086 https://news.ycombinator.com/item?id=35100086). Maybe those were the times you learned about it? Nothing wrong with letting new people know about it too now and then.
- Serveurperso 2mo ago[flagged]
- car 2mo agoThis MacOS app used to be called LlamaBarn. Really excellent to see the fast progress being made. Official repo, also has documentation how to configure server parameters: https://github.com/ggml-org/Llama-macOS https://github.com/ggml-org/Llama-macOS Small tip, install llama.cpp with brew before llama.app, which will pick up the existing llama.cpp. That way it's easier to stay up to date with llama.cpp, since llama.app is on a slower release cadence. Also, models installed with the hugging face CLI (hf) are picked up by llama.app automatically. The CLI will keep the model cache updated, e.g. when models get updated. Llama.cpp became part of Huggingface recently.
- TekMol 2mo agoI tried curl -LsSf https://llama.app/install.sh | sh and then llama serve -hf unsloth/Qwen3-4B-GGUF:Q4_0 Then I get: W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden Terminated And the web interface says Server unavailable Maybe it gets killed by the OS because it uses too much RAM? When I try llama serve -hf unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M It seems to work. Nice.
- sylware 2mo agoAny success at transpiling it to C? Using the cfront transpiler improved with coding AI? :)
- redmoonx 2mo agoLlama.cpp team has failed to make their tech easy to install and use for years. Why can’t they figure it out???
- 72deluxe 2mo agoBut it is easy ... clone the git repo, make a build dir, cd into it, run cmake .., run build/bin/llama-server -m /path/to/model.gguf Browse to served web page with chat UI.... By "easy" do you mean "very lazy"?
- pinkmoonx 2mo ago[dead]
- 72deluxe 2mo agoHaha ok, like click & run sort of levels? They do package builds on their github releases depending on your architecture (CUDA or not etc.) so it should be close to "very lazy" levels of ease. I find it useful, but the models I run are pretty rubbish due to my lack of RAM, which is a pity.
- greenmoonx 2mo ago[flagged]
- walrus01 2mo agoThis is like complaining that the mariadb team has failed to make their tech easy to use. It's a server side piece of software. There are plenty of user friendly GUIs or wrappers for it. lm-studio or unsloth are to llama.cpp as something like phpmyadmin is to mysql/mariadb.
- greenmoonx 2mo ago[dead]
- jurgenburgen 2mo agoThere’s now a `llama serve` command? I had to do a double take in case I was reading the `ollama` website.
- equalsione 2mo agoIf I said, “I want a setup that is usable for an agentic coding workflow, and it MUST be local”, what’s the smallest/cheapest option right now? It’s _technically_ possible to get agents running on all kinds of setups but there seems to be an (undefined) floor for useful setups. A lot of the stories people have about getting setups running on relatively low end hardware turn out to have huge compromises or run into issues on anything but trivial cases. I’ve found it hard to find a consensus. Or maybe I just don’t like the multiple thousand dollar price tags people are suggesting…
- timmmmmmay 2mo agostraightforward solution is probably Qwen3.6-27B-Q4 running on a used RTX 3090. price on those is unfortunately high, in fact so high that getting a new Radeon AI PRO R9700 might be a better deal a solid step up from there is anything that can run Deepseek V4 Flash but the hardware ask there is a bit higher
- kamranjon 2mo agoI actually have been exploring this very thing! I think the best option right now, since Apple has raised prices and Mac minis are basically impossible to get your hands on, is to build your own micro-itx machine. I actually built a mini-itx machine, but it does restrict your options a bit. The Arc series Intel GPUs are what I think make this possible. I built a machine with an Arc b50 - it runs Gemma 26b a4b qat at around 30tok/s with their MTP head and prompt processing sits at around 500 tok/s. The really beautiful thing about this setup is the entire energy envelope of this machine sits at 120w at full load - when idle, it's at 40w and i've done some work in ubuntu to basically intelligently hibernate, which drops it to 0 watts when not in use. You can use a raspberry pi and Wake on Lan to wake the machine up for a overall draw of around 5 watts when not in use. All in all this machine cost me 1.4k to build - but if you used micro-itx instead of mini-itx parts you could do it for under 1k - it has just 16gb of ddr5 but you don't really need more if you use models that can fit in vram. I think it's pretty incredible that you can run an actually useful coding agent on a machine with a power envelope that is less than an incandescent light bulb. If you go up to micro-itx you can do even large cards like an intel b60 with 24gb or a b70 with 32gb and run even more powerful models. For all of these intel GPU's you'll want to compile the latest llama.cpp version with SYCL support - they are getting speedups every day, so worth staying on the edge.
- jrm4 2mo agoHey folks, a bit of a hijack; but I've taken to using the Kobold gui, which I'm liking much more than Ollama -- but are there any major benefits to going straight up llama.cpp? Yes, this is somewhat of a question about laziness.
- fintuner 2mo ago[flagged]
- qiine 2mo agoHa the logo is an L made of negative spaces haaa.
- ivorymist_04 2mo agoIvorymist_04
- macwhisperer 2mo agopro tip: if u are building ur own harness with python (recommended) , use "llama-cpp-python"... I started with "llama-server" and custom stuff around it, which is great for single model setups.. but for multi-model harness with quick switching, llama-cpp-python is peak
- teleforce 2mo agoCheck this DLLM D project, a minimal, clean coding agent built directly on llama.cpp without Python, bindings or overhead [1]. The blog on the DLLM development [2]. [1] DLLM: https://github.com/DannyArends/DLLM https://github.com/DannyArends/DLLM [2] Teaching an AI to Know Itself: Building a Local LLM Agent in D: https://blog.dlang.org/2026/06/07/teaching-an-ai-to-know-itself-building-a-local-llm-agent-in-d/ https://blog.dlang.org/2026/06/07/teaching-an-ai-to-know-its...
- bobbylarson 2mo ago[flagged]