13 ms·
Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI
- rvz 8mo agoThis acquisition is almost the same as the acquisition of Bun by Anthropic. Both $0 revenue "companies", but have created software that is essential to the wider ecosystem and has mindshare value; Bun for Javascript and Ggml for AI models. But of course the VCs needed an exit sooner or later. That was inevitable.
- deleted 8mo ago[deleted]
- andsoitis 8mo agoI believe ggml.ai was funded by angel investors, not VC.
- jimmydoe 8mo agoAmazing. I like the openness of both project and really excited for them. Hopefully this does not mean consolidation due to resource dry up but true fusion of the bests.
- mnewme 8mo agoHuggingface is the silent GOAT of the AI space, such a great community and platform
- lairv 8mo agoTruly amazing that they've managed to build an open and profitable platform without shady practices
- al_borland 8mo agoIt’s such a sad state of affairs when shady practices are so normal that finding a company without them is noteworthy.
- geooff_ 8mo agoAs someone who's been in the "AI" space for a while its strange how Hugging Face went from one of the biggest name to not a part of the discussion at all.
- LatencyKills 8mo agoIt isn't necessary to be part of the discussion if you are truly adding value (which HF continues to do). It's nice to see a company doing what it does best without constantly driving the hype train.
- r_lee 8mo agoI think that's because there's less local AI usage now since there's all kinds of image models by the big labs, so there's really no rush of people self hosting stable diffusion etc anymore the space moved from Consumer to Enterprise pretty fast due to models getting bigger
- zozbot234 8mo agoToday's free models are not really bigger when you account for the use of MoE (with ever increasing sparsity, meaning a smaller fraction of active parameters), and better ways of managing KV caching. You can do useful things with very little RAM/VRAM, it just gets slower and slower the more you try to squeeze it where it doesn't quite belong. But that's not a problem if you're willing to wait for every answer.
- r_lee 7mo agoyeah, but I mean more like the old setups where you'd just load a model on a 4090 or something, even with MoE it's a lot more complex and takes more VRAM, right? like it just seems not justifiable for most hobbyists but maybe I'm just slightly out of the loop
- zozbot234 7mo agoWith sparse MoE it's worth running the experts in system RAM since that allows you to transparently use mmap and inactive experts can stay on disk. Of course that's also a slowdown unless you have enough RAM for the full set, but it lets you run much larger models on smaller systems.
- Filip_portive 8mo ago[flagged]
- HanClinto 8mo agoI'm regularly amazed that HuggingFace is able to make money. It does so much good for the world. How solid is its business model? Is it long-term viable? Will they ever "sell out"?
- I_am_tiberius 8mo agoI once tried hugging face because I wanted I worked through some tutorial. They wanted my credit card details during the registration as far as I remember. After a month they invoiced me some amount of money and I had no idea what it was. To be honest, I don't understand what exactly they do and what services I was paying for, but I cancelled my account and never touched it again. For me that was a totally intransparent process.
- dmezzetti 8mo agoThey have paid hosting - https://huggingface.co/enterprise https://huggingface.co/enterprise and paid accounts. Also consulting services. Seems like a pretty good foundation to me.
- julien_c 8mo agoand a lot of traction on paid (private in particular) storage these days; sneak peek at new landing page: https://huggingface.co/storage https://huggingface.co/storage
- microsoftedging 8mo agoFT had a solid piece a few weeks back: "Why AI start-up Hugging Face turned down a $500mn Nvidia deal" https://giftarticle.ft.com/giftarticle/actions/redeem/9b4eca55-1214-4f9e-b85e-58571d8da8d4 https://giftarticle.ft.com/giftarticle/actions/redeem/9b4eca...
- dmezzetti 8mo agoThis is really great news. I've been one of the strongest supporters of local AI dedicating thousands of hours towards building a framework to enable it. I'm looking forward to seeing what comes of it!
- logicallee 8mo ago>I've been one of the strongest supporters of local AI, dedicating thousands of hours towards building a framework to enable it. Sounds like you're very serious about supporting local AI. I have a query for you (and anyone else who feels like donating) about whether you'd be willing to donate some memory/bandwidth resources p2p to hosting an offline model: We have a local model we would like to distribute but don't have a good CDN. As a user/supporter question, would you be willing to donate some spare memory/bandwidth in a simple dedicated browser tab you keep open on your desktop that plays silent audio (to not be put in the background and deloaded) and then allocates 100mb -1 gb of RAM and acts as a webrtc peer, serving checksumed models?[1] (Then our server only has to check that you still have it from time to time, by sending you some salt and a part of the file to hash and your tab proves it still has it by doing so). This doesn't require any trust, and the receiving user will also hash it and report if there's a mismatch. Our server federates the p2p connections, so when someone downloads they do so from a trusted peer (one who has contributed and passed the audits) like you. We considered building a binary for people to run but we consider that people couldn't trust our binaries, or would target our build process somehow, we are paranoid about trust, whereas a web model is inherently untrusted and safer. Why do all this? The purpose of this would be to host an offline model: we successfully ported a 1 GB model from C++ and Python to WASM and WebGPU (you can see Claude doing so here, we livestreamed some of it[2]), but the model weights at 1 GB are too much for us to host. Please let us know whether this is something you would contribute a background tab to hosting on your desktop. It wouldn't impact you much and you could set how much memory to dedicate to it, but you would have the good feeling of knowing that you're helping people run a trusted offline model if they want - from their very own browser, no download required. The model we ported is fast enough for anyone to run on their own machines. Let me know if this is something you'd be willing to keep a tab open for. [1] filesharing over webrtc works like this: https://taonexus.com/p2pfilesharing/ https://taonexus.com/p2pfilesharing/ you can try it in 2 browser tabs. [2] https://www.youtube.com/watch?v=tbAkySCXyp0and https://www.youtube.com/watch?v=tbAkySCXyp0and and some other videos
- beoberha 8mo agoSeems like a great fit - kinda surprised it didn’t happen sooner. I think we are deep in the valley of local AI, but I’d be willing to bet it breaks out in the next 2-3 years. Here’s hoping!
- breisa 7mo agoI mean they already supported the project quite a bit. @ngxson and maybe others? from Huggingface are big contributors to llama.cpp.
- mythz 8mo agoI consider HuggingFace more "Open AI" than OpenAI - one of the few quiet heroes (along with Chinese OSS) helping bring on-premise AI to the masses. I'm old enough to remember when traffic was expensive, so I've no idea how they've managed to offer free hosting for so many models. Hopefully it's backed by a sustainable business model, as the ecosystem would be meaningfully worse without them. We still need good value hardware to run Kimi/GLM in-house, but at least we've got the weights and distribution sorted.
- zozbot234 8mo ago> We still need good value hardware to run Kimi/GLM in-house If you stream weights in from SSD storage and freely use swap to extend your KV cache it will be really slow (multiple seconds per token!) but run on basically anything. And that's still really good for stuff that can be computed overnight, perhaps even by batching many requests simultaneously. It gets progressively better as you add more compute, of course.
- HPsquared 8mo agoAt a certain point the energy starts to cost more than renting some GPUs.
- vardalab 7mo agoYeah, that is hard to argue with because I just go to OpenRouter and play around with a lot of models before I decide which ones I like. But there's something special about running it locally in your basement
- dotancohen 7mo agoI'd love to hear more about this. How do you decide that you like a model? For which use cases?
- fc417fc802 7mo agoAren't decent GPU boxes in excess of $5 per hour? At $0.20 per kWhr (which is on the high side in the US) running a 1 kW workstation 24/7 would work out to the same price as 1 hour of GPU time. The issue you'll actually run into is that most residential housing isn't wired for more than ~2kW per room.
- the__alchemist 8mo agoDoes anyone have a good comparison of HuggingFace/Candle to Burn? I am testing them concurrently, and Burn seems to have an easier-to-use API. (And can use Candle as a backend, which is confusing) When I ask on Reddit or Discord channels, people overwhelmingly recommend Burn, but provide no concrete reasons beyond "Candle is more for inference while Burn is training and inference". This doesn't track, as I've done training on Candle. So, if you've used both: Thoughts?
- csunoser 7mo agoI have used both (albeit 2 years ago, and things change really fast). At the time, Candle didn't have 2d conv backprop with strides properly implemented. And getting Burn running libtch backend was just a lot simpler. I did use candle for wasm based inference for teaching purposes - that was reasonably painless and pretty nice.
- dhruv3006 8mo agoHuggingface is actually something thats driving good in the world. Good to see this collab/
- androiddrew 8mo agoOne of the few acquisitions I do support
- tkp-415 8mo agoCan anyone point me in the direction of getting a model to run locally and efficiently inside something like a Docker container on a system with not so strong computing power (aka a Macbook M1 with 8gb of memory)? Is my only option to invest in a system with more computing power? These local models look great, especially something like https://huggingface.co/AlicanKiraz0/Cybersecurity-BaronLLM_Offensive_Security_LLM_Q6_K_GGUF https://huggingface.co/AlicanKiraz0/Cybersecurity-BaronLLM_O... for assisting in penetration testing. I've experimented with a variety of configurations on my local system, but in the end it turns into a make shift heater.
- xrd 8mo agoI think a better bet is to ask on reddit. https://www.reddit.com/r/LocalLLM/ https://www.reddit.com/r/LocalLLM/ Everytime I ask the same thing here, people point me there.
- zozbot234 8mo agoThe general rule of thumb is that you should feel free to quantize even as low as 2 bits average if this helps you run a model with more active parameters. Quantized models are not perfect at all, but they're preferable to the models with fewer, bigger parameters. With 8GB usable, you could run models with up to 32B active at heavy quantization.
- zargon 7mo agoA large model (100B+, the more the better) may be acceptable at 2-bit quantization, depending on the task. But not a small model. Especially not for technical tasks. On top of that, one still needs room for OS, software and KV cache. 8GB is just not very useful for local LLMs. That said, it can still be entertaining to try out a 4-bit 8B model for the fun of it.
- zozbot234 7mo ago100B+ is the amount of total parameters, whereas what matters here is active - very different for sparse MoE models. You're right that there's some overhead for the OS/software stack but it's not that much. KV-cache is a good candidate for being swapped out, since it only gets a limited amount of writes per emitted token.
- option 8mo agoIsn't HF banned in China? Also, how are many Chinese labs on Twitter all the time? In either case - huge thanks to them for keeping AI open!
- woadwarrior01 8mo agoHF is indeed banned in China. The Chinese equivalent of HF is ModelScope[1]. [1]: https://modelscope.cn/ https://modelscope.cn/
- disiplus 8mo agoI think in the West we think everything is blocked. But for example, if you book an eSIM, when you visit you already get direct access to Western services because they route it to some other server. Hong Kong is totally different: they basically use WhatsApp and Google Maps, and everything worked when I was there.
- embedding-shape 8mo agoBut also yes, parent is right, HF is more or less inaccessible, and Modelscope frequently cited as the mirror to use (although many Chinese labs seems to treat HF as the mirror, and Modelscope as the "real" origin).
- dragonwriter 8mo ago> Isn't HF banned in China? I think, for some definition of “banned”, that’s the case. It doesn’t stop the Chinese labs from having organization accounts on HF and distributing models there. ModelScope is apparently the HF-equivalent for reaching Chinese users.
- segmondy 8mo agoGreat news! I have always worried about ggml and long term prospect for them and wished for them to be rewarded for their effort.
- stephantul 8mo agoGeorgi is such a legend. Glad to see this happening
- jgrahamc 8mo agoThis is great news. I've been sponsoring ggml/llama.cpp/Georgi since 2023 via Github. Glad to see this outcome. I hope you don't mind Georgi but I'm going to cancel my sponsorship now you and the code have found a home!
- superkuh 8mo agoI'm glad the llama.cpp and the ggml backing are getting consistent reliable economic support. I'm glad that ggerganov is getting rewarded for making such excellent tools. I am somewhat anxious about "integration with the Hugging Face transformers library" and possible python ecosystem entanglements that might cause. I know llama.cpp and ggml already have plenty of python tooling but it's not strictly required unless you're quantizing models yourself or other such things.
- periodjet 7mo agoPrediction: Amazon will end up buying HuggingFace. Screenshot this.
- ukblewis 7mo agoHonestly I’m shocked to be the only one I see of this opinion: HuggingFace’s `accelerate`, `transformers` and `datasets` have been some of the worst open source Python libraries I have ever used that I had to use. They break backwards compatibility constantly, even on APIs which are not underscore/dunder named even on minor version releases without even documenting this, they refuse PRs fixing their lack of `overloads` type annotations which breaks type checking on their libraries and they just generally seem to have spaghetti code. I am not excited that another team is joining them and consolidating more engineering might in the hands of these people
- 0xbadcafebee 7mo ago> The community will continue to operate fully autonomously and make technical and architectural decisions as usual. Hugging Face is providing the project with long-term sustainable resources, improving the chances of the project to grow and thrive. The project will continue to be 100% open-source and community driven as it is now. I want this to be true, but business interests win out in the end. Llama.cpp is now the de-facto standard for local inference; more and more projects depend on it. If a company controls it, that means that company controls the local LLM ecosystem. And yeah, Hugging Face seems nice now... so did Google originally. If we all don't want to be locked in, we either need a llama.cpp competitor (with a universal abstration), or it should be controlled by an independent nonprofit.
- zozbot234 7mo agoLlama.cpp is an open source project that anyone can fork as needed, so any "control" over it really only extends to facilitating development of certain features.
- 0xbadcafebee 7mo agoIn practice, nobody does this, because you then have to keep the fork up to date with upstream plus your changes, and this is an endless amount of work.
- raphaelmolly8 7mo ago[dead]
- sheepscreek 7mo agoCurious about the financials behind this deal. Did they close above what they raised? What’s in it for HuggingFace?
- simonw 7mo agoIt's hard to overstate the impact Georgi Gerganov and llama.cpp have had on the local model space. He pretty much kicked off the revolution in March 2023, making LLaMA work on consumer laptops. Here's that README from March 10th 2023 https://github.com/ggml-org/llama.cpp/blob/775328064e69db1ebd7e19ccb59d2a7fa6142470/README.md https://github.com/ggml-org/llama.cpp/blob/775328064e69db1eb... > The main goal is to run the model using 4-bit quantization on a MacBook. [...] This was hacked in an evening - I have no idea if it works correctly. Hugging Face have been a great open source steward of Transformers, I'm optimistic the same will be true for GGML. I wrote a bit about this here: https://simonwillison.net/2026/Feb/20/ggmlai-joins-hugging-face/ https://simonwillison.net/2026/Feb/20/ggmlai-joins-hugging-f...
- ushakov 7mo agoi am curious, why are your comments always pinned to the top?
- carbocation 7mo agoBecause many of us think simonw has discerning taste on this topic and like to read what he has to say about it, so we upvote his comments.
- ushakov 7mo agoi don't doubt this. i just find it questionable that one particular poster always gets in the spotlight when AI is the topic - while other conversations in my opinion offer more interesting angles.
- mattfrommars 7mo agoI don’t know if this warrants a separate thread here but I have to ask… How can I realistically get involved the AI development space? I feel left out with what’s going on and living in a bubble where AI is forced into by my employer to make use of it (GitHub Copilot), what is a realistic road map to kinda slowly get into AI development, whatever that means My background is full stack development in Java and React, albeit development is slow. I’ve only messed with AI on very application side, created a local chat bot for demo purposes to understand what RAG is about to running models locally. But all of this is very superficial and I feel I’m not in the deep with what AI is about. I get I’m too ‘late’ to be on the side of building the next frontier model and makes no sense, what else can I do? I know Python, next step is maybe do ‘LLM from scratch”? Or I pick up Google machine learning crash course certificate? Or do recently released Nvidia Certification? I’m open for suggestions
- fc417fc802 7mo agoI'm not entirely clear what your goals are but roughly, just figure out an application that holds your interest and build a model for it from scratch. Probably don't start with an LLM though. Same as for anything else really. If you're interest in computer graphics then decide on a small scale project and go build it from scratch. Etc.
- breisa 7mo agoMaybe look into model finetuning/distilation. Unsloth [1] has great guides and provides everything you need to get started on Google Colab for free. [1] https://unsloth.ai/ https://unsloth.ai/
- w10-1 7mo agoThe competition for root and branch AI models and infrastructure is intense and skilled. But if you're adjacent to some leaf use-case for AI, you're likely already as good as anyone else at productizing it. And that's who is getting hired: people who show they can deliver product-market fit.
- swyx 7mo agogo thru workshops here https://www.youtube.com/@aiDotEngineer/ https://www.youtube.com/@aiDotEngineer/
- fancy_pantser 7mo agoWas Georgi ever approached by Meta? I wonder what they offered (I'm glad they didn't succeed, just morbid curiosity).
- karmasimida 7mo agoDoes local AI have a future? The models are getting ridiculously big and any storage hardware is hoarded by few companies for next 2 years and nvidia has stopped making consumer GPU for this year. It seems to me there is no chance local ML is going to be anywhere out of the toy status comparing to closed source ones in short term
- rhdunn 7mo agoMistral have small variants (3B, 8B, 14B, etc.), as do others like IBM Granite and Qwen. Then there are finetunes based on these models, depending on your workflow/requirements.
- karmasimida 7mo agoTrue, but anything remotely useful is 300B and above
- dust42 7mo agoI am actually doing now a good part of dev with Qwen3-Coder-Next on an M1 64GB with Qwen Code CLI (a fork of Gemini CLI). I very much like a) to have an idea how much tokens I use and b) be independent of VC financed token machines and c) I can use it on a plane/train Also I never have to wait in a queue, nor will I be told to wait for a few hours. And I get many answers in a second. I don't do full vibe coding with a dozen agents though. I read all the code it produces and guide it where necessary. Last not least, at some point the VC funded party will be over and when this happens one better knows how to be highly efficient in AI token use.
- kristianp 7mo ago> Towards seamless “single-click” integration with the transformers library That's interesting. I thought they would be somewhat redundant. They do similar things after all, except training.
- lukebechtel 7mo agoThank you Georgi <3
- cboyardee 7mo ago[dead]
- forty 7mo agoLooks like someone tried to type "Gmail" while drunk...
- rkomorn 7mo agoLooks like Gargamel of Smurfs fame to me.
- cyanydeez 7mo agoIs there a local webui that integrates with Hugging face? Ollama and webui seem to rapidly lose their charm. Ollama now includes cloud apis which makes no sense as a local.
- moralestapia 7mo agoI hope Georgi gets a big fat check out of this, he deserves it 100%.
- snowhale 7mo ago[dead]
- sbinnee 7mo agoI am happy for ggml team. They did so much work for quantization and actually made it available to everyone. Thank you.
- indiekitai 7mo ago[dead]
- ontouchstart 7mo agoI have played with both mlx-lm and llama.cpp after I bought a 24GB M5 MacBook Pro last year. Then I fell down the rabbit holes of uv, rust and C++ and forgot about LLMs. Today after I saw this announcement and answered someone’s question about how to set it up, when I got home, I decided play with llama.cpp again. I was surprised and impressed: https://ontouchstart.github.io/rabbit-holes/llama.cpp/ https://ontouchstart.github.io/rabbit-holes/llama.cpp/ I am not going to use mlx-lm or lmstudio anymore. llama.cpp is so much fun.
- car 7mo agoSo great to see my two favorite Open Source AI projects/companies joining forces. Since I don't see it mentioned here, LlamaBarn is an awesome little—but mighty—MacOS menubar program, making access to llama.cpp's great web UI and downloading of tastefully curated models easy as pie. It automatically determines the available model- and context-sizes based on available RAM. https://github.com/ggml-org/LlamaBarn https://github.com/ggml-org/LlamaBarn Downloaded models live in: ~/.llamabarn Apart from running on localhost, the server address and port can be set via CLI: # bind to all interfaces (0.0.0.0) defaults write app.llamabarn.LlamaBarn exposeToNetwork -bool YES # or bind to a specific IP (e.g., for Tailscale) defaults write app.llamabarn.LlamaBarn exposeToNetwork -string "100.x.x.x" # disable (default) defaults delete app.llamabarn.LlamaBarn exposeToNetwork
- noisy_boy 7mo agoGithub is showing me unicorn - is there an Linux equivalent? I have a old Thinkpad with a puny Nvidia GPU, can I hope to find anything useful to run on that?
- car 7mo agoBuilding Llama.cpp from source with CUDA enabled should get you pretty far. llama-server has a really good web UI, the latest version supports model switching. As for models, plenty of GGUF quantized (down to 2-bit) available on HF and modelscope.
- am17an 7mo agoOne often overlooked after that is ggml, the tensor library that runs llama.cpp is not based on pytorch, rather just plain cpp. In a world where pytorch dominates, it shows that alternatives are possible and are worthy to be pursued.
- mhher 7mo agoIt's great to see the ggml team getting proper backing. Keeping inference in bare-metal C/C++ without the Python bloat is the only way local AI is going to scale efficiently. Well deserved for Georgi, Johannes, Piotr, and the rest of the team.
- jpcompartir 7mo agoThis is great, brings clear benefits to both sides and the rest of us. Always rooting for Hugging Face
- genie3io 7mo ago[dead]
- 45dsilicon 7mo ago[dead]
- superactro 7mo ago[dead]