25 ms·
Llama.cpp: Full CUDA GPU Acceleration
- rektide 3y agoSuch a pity no one else can compete here presently. Would that others be able to gain a position where their software made them competitive on the free market.
- theaiquestion 3y agoCompete with Llama.cpp? Like transformers llama [0], exllama [1] (really fast), or litllama [2] ? exllama is really memory efficient and really fast [0] https://huggingface.co/docs/transformers/main/model_doc/llama https://huggingface.co/docs/transformers/main/model_doc/llam... [1] https://github.com/turboderp/exllama https://github.com/turboderp/exllama [2] https://github.com/Lightning-AI/lit-llama https://github.com/Lightning-AI/lit-llama EDIT: Or do you mean cuda? Because yeah, it's such a shame AMD's Rocm is so bad even geohot gave up. it's examples don't even run without crashing. https://github.com/RadeonOpenCompute/ROCm/issues/2198#issuecomment-1574383483 https://github.com/RadeonOpenCompute/ROCm/issues/2198#issuec...
- dTal 3y agoThanks for the tip about exllama, I've been on the lookout for a readable python implementation to play with that is also fast and has support for quantized datasets.
- kayvr 3y agoAlso https://github.com/kayvr/TokenHawk https://github.com/kayvr/TokenHawk, a WebGPU implementation of LLaMA. edit: Note that this is my project.
- smoldesu 3y agoThere was free competition here, a while ago. OpenCL was formed by Apple, Khronos et al. to stave off CUDA's dominance. The platform languished from a lack of commitment though, and Apple eventually gave up on open GPU APIs entirely. Nvidia continued funding CUDA and scaling it for industry application, and the rest is history. The landscape of stakeholders is just too bitter to unseat CUDA for what it's used for - your best shot at democratizing AI inferencing acceleration is through something like Microsoft's ONNX[0] runtime. [0] https://onnxruntime.ai/ https://onnxruntime.ai/
- nl 3y agoNote that Llama supports acceleration on both OpenCL and Apple Metal
- chrischen 3y agoThere’s also geohot’s tiny corp betting on AMD gpus.
- pclmulqdq 3y agoNot any more.
- nromiun 3y agohttps://geohot.github.io/blog/jekyll/update/2023/06/07/a-dive-into-amds-drivers.html https://geohot.github.io/blog/jekyll/update/2023/06/07/a-div... AMD gave him a binary blob driver and that fixed his problem. Also, tinygrad is the only Python framework I know that has full OpenCL acceleration.
- vvladymyrov 3y agoWhat do you mean? At least as of June 7 geohot was still working on amd drivers builds and stability. https://geohot.github.io/blog/jekyll/update/2023/06/07/a-dive-into-amds-drivers.html https://geohot.github.io/blog/jekyll/update/2023/06/07/a-div... So far it doesn’t look that AMD is fully on board with Tiny Corp, but they are talking…
- deleted 3y ago[deleted]
- m00x 3y agoWhy not ggml?
- eyegor 3y agoCUDA had a lot of inertia and opencl brought half baked docs and half baked support out of the gate. If they had focused on simplifying their api to be more user friendly for the 80% use case it could've been a success. Opencl always looked nice on the surface but a few hours in and you've exhausted the docs trying to figure out what to do and there's no good example code around. Of course if they really wanted it to succeed they would've built a Cuda to opencl transpiler for the c api or at least a comprehensive migration guide. I'm not convinced anyone involved was trying to make it popular.
- ShamelessC 3y agoThis isn't a market.
- angch 3y agoThere's Fabrice Bellard's textsynth server. https://bellard.org/ts_server/ https://bellard.org/ts_server/ No open source though.
- nl 3y agoUnclear what this is referring to, but if it means CUDA vs other things it is worth noting that: a) CUDA won in a free market because NVidia showed they cared about it b) Llama has support for OpenCL (via CLBlast) and Apple Metal The OpenCL support already has a custom kernel for token generation.
- jfdi 3y agoIs there a legitimate way to get the weights to actually use this without filling in forms?
- wmf 3y agoNot if you want the original LLAMA weights, but now there are other models like RedPajama available.
- csjh 3y agoThey come a dime a dozen on HuggingFace, check out https://old.reddit.com/r/LocalLLaMA/wiki/models https://old.reddit.com/r/LocalLLaMA/wiki/models for a few options
- quickthrower2 3y agoAre these done as LLaMA deltas still? I.e. do I need to apply a patch to LLaMA, and so I still need to source LLaMA?
- Tostino 3y agoMost of them are merged models, so you don't need the base model. It's stupidly simple to get going.
- pram 3y agoThere’s a torrent linked in the Llama.cpp docs, it’s in a merge request on the LLaMA repo. Has all the files.
- hnfong 3y agoIt's almost (or actually is) a "pirated" torrent. So it might not be "legitimate".
- adultSwim 3y agoNo
- 3y ago
- adeon 3y agollama.cpp is great. It started off as CPU-only solution and now looks like it wants to support any computation device it can. I find it interesting that it's an example of an ML software that's totally detached from Python ML ecosystem and also popular. Is Python so annoying to use that when a compelling non-Python solution appears, everyone will love it? Less hassle? Or did it take off for a different reason? Interested in hearing thoughts.
- csjh 3y agoI believe it's moreso for the (actively pursued) speed optimizations it provides. When inference is already computationally expensive any bit of performance is a big plus
- plaguuuuuu 3y agoI've been using it because there are bindings in other languages I know like .NET
- nmfisher 3y agoIt's a great project and an impressive achievement, but I'm also struggling to understand what people use it for that PyTorch wasn't offering. Easy deployment on iOS I guess? I would have thought that's a pretty small use case though. Given the author hand-rolled his own FFT, I'm also guessing it's not as performant?
- civilitty 3y agoEasy deployment anywhere, not just iOS. I haven't used Python in years so I have no idea what package manager is the best now, completely forgot how to use virtualenv, and it only took a few weeks to completely fuck up my local Python install ("Your version of CUDA doesn't match the one used to blah blah") Python is a mess. llama.cpp was literally a git clone followed by "cd llama.cpp && make && ./main" - I can recite the commands from memory and I haven't done any C/C++ development in a long time.
- nmfisher 3y ago
- shon 3y agoNice to see Georgi has started a company: https://twitter.com/ggerganov/status/1666120568993730561?s=46&t=pO499fGQKTiGvvZPpc-cFw https://twitter.com/ggerganov/status/1666120568993730561?s=4... Godspeed
- fnands 3y agoNice. He's obviously a talented engineer who's struck a nerve with the whisper.cpp/llama.cpp projects, so hope he has success with whatever he plans to do.
- csmpltn 3y agoA lot of work going into refactoring proprietary code that can be randomly deprecated and outcompeted without any prior notice by any number of large competitors... problematic business model, in my opinion.
- logicchains 3y agoHe's building a library, ggml; it's generally not hard to add support for new models. For instance llama.ccp already supports the Falcon 7B model (different architecture to llama). And given how politicised AI has become, there's unlikely to be many companies releasing weights for models competitive with the current models (e.g. LLaMA 65B). They may have private models that are better, like GPT3.5 and GPT4, but you can't run these on your own server so they're not competing with ggml.
- csmpltn 3y ago> "there's unlikely to be many companies releasing weights for models competitive with the current models" We're at the very dawn of this technology going mainstream, and you're saying that it's unlikely for new players to release new, competing and incompatible models?
- getcrunk 3y agoAnyone get performance numbers for other 30 series cards? 3060 12gb? I’m curious how it compares to his apple silicon numbers
- speed_spread 3y agoAlso, someone please let me use my 3050 4GB for something else than stable diffusion generation of silly thumbnail-sized pics. I'd be happy with an LLM that's specialized in insults and car analogies.
- wing-_-nuts 3y agoDid you buy that card for cuda? Cause otherwise I have no idea why someone would chose a 3050 over a 6600
- machinawhite 3y agoWhat's a 6600, some AMD card? And why is it better?
- smoldesu 3y agoCheaper Nvidia cards are generally considered to have dubious value. Having seen the benchmarks I agree, but it's not like a game-changing difference really. For CUDA and ML stuff, the 3050 would run circles around the 6600.
- wing-_-nuts 3y agoFor cuda and ML you'd be much better off choosing a 3060. Honestly, if you've only got the money for a 4gb 3050, you're prob better off working in google colab
- smoldesu 3y agoWith layering enabled, I don't necessarily agree. Not being able to load an entire model into memory isn't a dealbreaker these days. You can even layer onto swap space if your drive is fast enough, so there's really no excuse not to use the hardware if you have it. Unless you just like the cloud and hate setting stuff up yourself, or what have you.
- gigel82 3y agoI'm a total newb about the implementation details, but I'm curious if a hybrid is possible (GPU+CPU) to enable inference with even larger models than what fits in consumer GPU VRAM.
- skirmish 3y agollama.cpp does it already. You tell it how many layers to offload to GPU, and it runs remaining ones on CPU.
- rektide 3y agoGot downvoted out of view for saying this once but no less true. Absolutely a pity & a shame no one else has competed with this market dominance by Nvidia. Just a shit world entirely that we are single vendored up. A lot of pissant shitty defenses of monopolization too. Wrong or right, this is a shit world we're in now. https://news.ycombinator.com/item?id=36304225 https://news.ycombinator.com/item?id=36304225
- fragmede 3y agotinygrad is trying to address this problem. We'll see if it's successful.
- m00x 3y agoYou can also use TPUs or other training cards. Nvidia is just the best one that's accessible. But I think you're getting downvoted since it's very off-topic.
- MuffinFlavored 3y agoit’s uh… not that serious
- smoldesu 3y agoCan hardware vendors even put their differences aside to build such a thing? We can't even build a unified open raster graphics API, and now you're asking for machine learning acceleration in that vein?
- OkayPhysicist 3y agoMachine learning would probably be the simpler API. If you can speak Linear Algebra, you're most of the way there.
- smoldesu 3y agoYou're right, and it's why projects like the ONNX runtime exist to unify vendor-specific AI accelerators. Covering the basics isn't too hard. What GP seems to be asking for is an open CUDA replacement, which is kinda like asking someone to fund a Free and Open Source cruise ship to compete with Carnival for you. You'll get somewhere with some effort, luck and good old human intuition, but Nvidia can outspend you 10:1 unless you have funding leverage from FAANG.
- Ono-Sendai 3y agoAnyone know if whisper.cpp is GPU accelerated yet?
- aidenn0 3y agoAFAIK it's not fully, but you can use cuBLAS/clBlast for a pretty good speedup.
- regularfry 3y agoPartially. Keep an eye out for ggml-cuda.cu getting updated.
- aidenn0 3y agoSlightly OT: I have been playing around with whisper.cpp; it's nice because I can run the large model (quantized to 8-bits) at roughly real-time with cublas on a Ryzen 2700 with a 1050Ti. I couldn't even run the pytorch whisper medium on this card with X11 also running. It blows me away that I can get real-time speech-to-text of this quality on a machine that is almost 5 years old.
- moneywoes 3y agoIs it possible to run on apple m1 devices or mobile phones or not yet?
- raihansaputra 3y agoyeah the whisper.cpp github page has a demo for both. Have used it on my M1 MBA for the past few months.
- michelb 3y agoI can recommend the MacWhisper app if you prefer a gui.
- Void_ 3y agoAnd Whisper Memos for iOS https://whispermemos.com/ https://whispermemos.com/
- b33f 3y agoThe really nice part of Whisper is being able to use it offline and on-device, it seems whisper memos is uploading your audio and notes to a server of unknown security, confidentiality etc. I like Aiko for on-device transcription both in macOS and iOS https://apps.apple.com/us/app/aiko/id1672085276 https://apps.apple.com/us/app/aiko/id1672085276
- Void_ 3y agoWhisper Memos uses OpenAI API. The upside is that it uses the largest model - that would take 2GB on your iPhone.
- jokethrowaway 3y agoGreat news but I'd like to know how does it compare with just using torchlib. If this is faster than torchlib this optimizations should flow to torchlib as well Love the idea of not having to deal with python, though; dependency management is just horrible, I'd much rather have ML projects written in cpp.
- hendry 3y agoRTX 3090 isn't cheap, more than 1000GBP new, crikey!
- Tepix 3y agoWhy not get a used one?
- hospitalJail 3y agoAs someone who uses their computers every day, for 7 years... Then I hand them down to my kids.... Then I turn them into servers. I find the cost of computing extremely affordable, even for high end stuff. Whats the amortization on a 2-3k computer over 7 years? How about if I use it 4 hours a day actively and 24 hours passively? I have considered spending 10-30k on a computer given the recent AI craze, but the thing stopping me is that by 2025, a 10-30k computer in the AI space is going to be 2-4x better. Only in the last 1 year are we finding out the importance of absurd amounts of VRAM. I feel like the 4090's 24gb VRAM is going to age alright at best, but most likely poorly. (Not that 4090 buyers are going to have qualms upgrading to the 6090)
- imranq 3y agoThis is pretty cool, do you find the server farm of older computers valuable for your own work?
- hospitalJail 3y agoOh yeah, I have a computer for a minecraft server. A computer hosting my kiddo's website(just for fun, its silly, but randomly he will want me to pull it up from outside of the house). That same computer hosts some listeners/watchdogs for a media computer, but I havent actually used much of that information or features in a year (WFH kind of removed the need for me to use my remote tools). I suppose that's it for now. Oh, I thought of another use, I run a small business on the side and my interns occasionally don't have a laptop, I give them a crappy laptop. (they are basically just using excel/google sheets)
- underdeserver 3y agoCan someone ELI5 why AMD is not in this game? Is it really so much harder to implement this in a non-platform-specific library?
- bilekas 3y agoI'm no expert but if I understand correctly the CUDA cores are the main pull and the API to them. They're supposed to be more optimized and more stable compared to AMD. That's how it was before anyway, not sure today.
- Aardwolf 3y agoIsn't the main component for AI matrix multiplication? What makes it so hard to create a good alternative API for matrix multiplication?
- bilekas 3y agoWell I think there are 2 types right ? Tensor cores (which afaik AMD dont have) which are better for matrix ops, and CUDO which are better for general parallel ops. Maybe someone more clever than me can go into the specifics, I only understand the minimum of the low lvl GPU details. Nice high lvl document [0] https://www.acecloudhosting.com/blog/cuda-cores-vs-tensor-cores/#Difference_Between_CUDA_Cores_and_Tensor_Cores https://www.acecloudhosting.com/blog/cuda-cores-vs-tensor-co...
- marcyb5st 3y agoI think API for matrix multiplication is just a part of the issue. CUDA tooling has better ergonomics, it's easier to set up and treated as first class citizen in tools like Tensorflow and Pytorch. So, while I can't talk about the hardware differences in detail, developer experience is greatly on nVidia side and now AMD has a moat to overcome to catch up.
- dotnet00 3y agoIt's a lot more complicated than just writing a matrix multiplication kernel because there are all sorts of operations you need to have on top of matrix multiplication (non linearities, various ways of manipulating the data) and this sort of effort is only really worthwhile if it's well optimized. On top of that, AMD's compute stack is fairly immature, their OpenCL support is buggy and ROCm compiles device specific code, so it has very limited hardware support and is kind of unrealistic to distribute compiled binaries for. Then, getting to the optimization aspect, NVIDIA has many tools which provide detailed information on the GPU's behavior, making it much easier to identify bottlenecks and optimize. AMD is still working on these. Finally, NVIDIA went out of its way to support ML applications. They provide a lot of their own tooling to make using them easier. AMD seems to have struggled on the "easier" part.
- supermatt 3y agoMy understanding from reading this is that a 3090 GPU is 2x speedup over a decent modern CPU. Is that really the case, or am I reading it wrong? My initial thought was that it would be far higher. Is this typical of inference for these kind of models? If so, why do we need such expensive hardware? Please excuse my lack of knowledge :)
- mFixman 3y agoI don't have a PC with a powerful GPU. What's the easiest way I can play with Llama on AWS, Google Cloud, or somebody else's computer?
- Tepix 3y agoYou can play with Llama on your CPU. Depending on the model you use and the RAM you have available, the performance may be acceptable.
- anentropic 3y agousing llama.cpp it runs on the CPU this news story is that they are now extending GPU support to llama.cpp
- hospitalJail 3y agoDo you know about Oobabooga? You can probably find a google colab link.
- mFixman 3y agoIs that like a LLM Stable Diffusion? Neat.
- sasanaderi 3y ago[dead]
- shgidigo 3y agoExcuse me for my ignorance, but can someone explain why the Llama.cpp is so popular? isn't it possible to port the pytroch lama to any environment using onnx or something?
- v3ss0n 3y agoShould be similar performance but gglm guy did it in what he knows best and biggest selling point is single binary
- jmiskovic 3y agoYou can run it on RPi or any old hardware, only limited by the RAM and your patience. It is a lean code base easy to get up and running, and designed to be interfaced from any app without sacrificing performance. They are also innovating (or at least implementing innovations from papers) different ways to fit bigger models in consumer HW, making them run faster and with better outputs. Pytorch and other libs (bitsandbytes) can be horrible to setup with correct versions, and updating the repo is painful. PyTorch projects require a hefty GPU or enormous CPU+RAM resources, while llama.cpp is flexible enough to use GPU but doesn't require it and runs smaller models well on any laptop. ONNX is a generalized ML platform for researchers to create new models with ease. Once your model is proven to work, there are many optimizations left on the table. At least for distributing an application that relies on LLM it would be easier to add llama.cpp than ONNX.
- ianpurton 3y agoONNX doesn't support the same level of quantization as GGML. So basically GGML will run on hardware with less memory.
- regularfry 3y agoOr alternatively, bigger models with the same memory (just quantised harder).
- naasking 3y agoLlama.cpp runs better than pytorch on a much wider variety of hardware, including mobile phones, Raspberry Pis and more.
- ERUIONKP 3y ago[flagged]
- davidy123 3y agoI'm a bit surprised by the numbers. It's "only" a 2× speedup on a relatively top-end card (4090)? And you can only use one CPU core. With 16+ core CPUs becoming normal and 128GB+ RAM being cheap, that seems like leaving a lot on the table. [edit] realized it's relative to the merged partial CUDA acceleration, so the speedup is more impressive, but still surprised by the core usage.
- LoganDark 3y ago> And you can only use one CPU core. because the core's job is solely to direct the GPU, which is doing all of the work.
- eurekin 3y agoThat cheap ram is about 10x slower than the VRAM. Didn't see any actual figures for latency, but there must be a reason, why newer gpus have memory chips on both sides of the PCB, as close to the GPU as possible
- davidy123 3y agoTo the replies; I think one feature of llama.cpp is it can handle models with more RAM than VRAM provides, this is where I would think more cores would be useful.
- yashthakker 3y ago[dead]
- seventeen17 3y ago[flagged]
- seventeen17 3y ago[flagged]
- Roark66 3y agoI don't quite get it. Where does one get the model from? Or is it just for people that can afford to spend the $ to train the models themselves?
- valyagolev 3y agoyou can easily download llama weights by googling the ipfs link. it doesn't even seem to be illegal to do so, but IANAL
- aaronscott 3y agoYou can download a model from a site like huggingface. Here is a list of models that can be used for inference: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb... This user has some models already compiled for use with GGML (look for models with that in the name): https://huggingface.co/TheBloke https://huggingface.co/TheBloke Or if you want to convert your own model the llama.cpp repo has good instructions. Briefly it's `python3 convert.py <model>` - then if you are using a large parameter model you may need to quantize it to fit in memory `./quantize <source_model> <destination_name> <quantization>`
- Roark66 3y agoThanks
- twobitshifter 3y agoWhat do people see becoming of the non commercial license on llama projects? How much is this type of work wed to using LlaMA?
- brucethemoose2 3y agoTBH people are kinda ignoring the LLaMA license now, as Meta seems to be doing. I see some pseudo commercial (encouraging donations and such) and a few straight up commercial services using a LLaMA backend. llama.cpp specifically has Falcon on their roadmap, and some other quantized implementations already work with it. But the transition will be slow.
- Tepix 3y agoI think other LLMs will be used for commercial purposes in most cases. There are alread a few and i'm sure there's more in the queue.
- Remmy 3y agoI've been using llama.cpp with the python wrappers and it's the speed increase has been great, but it seemed to be limited to a max of 40 N_GPU_LAYERS. Going to have to update and see what sort of improvement I see.
- johnthescott 3y ago> With a boring C project, if it compiles it probably works without hassle. amen.
- mhh__ 3y agoApart from all the graphics drivers and it actually working at all