10 ms·
A brief history of LLaMA models
- simonw 3y agoI'm running Vicuna (a LLaMA variant) on my iPhone right now. https://twitter.com/simonw/status/1652358994214928384 https://twitter.com/simonw/status/1652358994214928384 The same team that built that iPhone app - MLC - also got Vicuna running directly in a web browser using Web GPU: https://simonwillison.net/2023/Apr/16/web-llm/ https://simonwillison.net/2023/Apr/16/web-llm/
- newswasboring 3y agoWith all these new AI models, both stable diffusion and llama specially, I'm considering switching to iPhone. I don't think I fully understand why iPhones and Macs are getting so many implementations but it seems like it's hardware based.
- sp332 3y agoiPhones leaned in to "computational photography" a long time ago. Eventually they added custom hardware to handle all the matrix multiplies efficiently. They exposed some of it to apps with an API called CoreML. They've been adding more features like on-device photo tagging, voice recognition, VR stuff.
- sagarm 3y agoGoogle was the leader on computational smartphone photography. They released their "night sight" mode before Samsung and Apple had anything competitive.
- sp332 3y agoSure, and you can run Stable Diffusion on normal Snapdragon SoCs, and there's a very hacky way to get llama.ccp running on a Pixel phone https://twitter.com/thiteanish/status/1635678053853536256 https://twitter.com/thiteanish/status/1635678053853536256 but I haven't seen any good apps yet.
- bkm 3y agoHomogenized hardware I assume, this is why iOS had so many photography Apps too.
- simonw 3y agoMy understanding is that part of it is that Apple Silicon shares all available RAM between CPU and GPU. I'm not sure how many of these models are actively taking advantage of that architecture yet though.
- int_19h 3y agoThe GPU isn't actually used by llama.cpp. What makes it that much faster is that the workload, either on CPU or on GPU, is very memory-intensive, so it benefits greatly from fast RAM. And Apple is using DDR5 running at very high clock speeds for this shared memory stuff. It's still noticeably slower than GPU, though.
- AnthonyMouse 3y agoMost of these implementations are not platform-specific. I've been running llama.cpp on x86_64 hardware and the performance is fine. The small models are fast and the quantized 65B model generates about a token per second on a system with dual-channel DDR4, which isn't unusable. The tough thing to find is something affordable that will run the unquantized 65B model at an acceptable speed. You can put 128GB of RAM in affordable hardware but ordinary desktops aren't fast. The things that are fast are expensive (e.g. I bet Epyc 9000 series would do great). And that's the thing Apple doesn't get you either, because Apple Silicon isn't available with that much RAM, and if it was it wouldn't be affordable (the 96GB Macbook Pro, which isn't enough to run the full model, is >$4000).
- spudlyo 3y agoIf you want to spend $4800.00 on just the computer, you can get a Mac Studio with 128G of memory with 400GB/s bandwidth. There are sparse reports out there of folks running 65B models on it. I've seen no performance measurements though.
- AnthonyMouse 3y agoIt's interesting that they actually have it but the price is still silly. SP5 system board ~$1000 Epyc 9124 $1083 192GB registered DDR5 (12x16GB) ~$1000 case, power supply, modest storage: ~$300 460GB/s bandwidth from 12 memory channels, 50% more memory and you'd have more than $1000 left over. But >$3000 is not a low price either, it's just lower.
- FloatArtifact 3y agoThere needs to be a slight dedicated to tracking all these models with regular updates.
- mdaniel 3y agoHeh, there is, and you're on it. But a slightly more serious answer is that would be a good feature for Huggingface to add since they're the GitHub of models. I actually suggested to GitHub that they should allow contributions of the repo topics since a lot of developers don't know or don't bother to add topics to their repos, making discoverability harder than necessary. GH ignored it but maybe Huggingface could implement such a thing
- jiggawatts 3y agoIt keeps saying the phrase “model you can run locally”, but despite days of trying, I failed to compile any of the GitHub repos associated with these models. None of the Python dependencies are strongly versioned, and “something” happened to the CUDA compatibility of one of them about a month ago. The original developers “got lucky” but now nobody else can compile this stuff. After years of using only C# and Rust, both of which have sane package managers with semantic versioning, lock files, reproducible builds, and even SHA checksums the Python package ecosystem looks ridiculously immature and even childish. Seriously, can anyone here build a docker image for running these models on CUDA? I think right now it’s borderline impossible, but I’d be happy to be corrected…
- KETpXDDzR 3y agollama.cpp was easy to setup IMO
- jiggawatts 3y agoCan you link to a working Dockerfile? I've heard several people say that it is easy, but then surely it ought to be trivial to set script the build so that it works reliable in a container!
- PostOnce 3y agoNo need to drag a gigabyte of docker stuff into this, just extract the zip file from github and type make into your terminal congratulations, it now works. If you're not a developer, maybe you'll have to type sudo apt install build-essential first. Congratulations, now you too, a non-developer, are running it locally. https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp
- vidarh 3y agoDid that, got a compiler error within seconds within seconds. Looks like it might need a newer version of gcc than is in the distro on my laptop. Which is why people ask for Docker. If it really will work just with make with build-essential on a new enough distro image, a Dockerfile that documents that would be trivial, and does not at all stop people from just typing make if their setup is new enough.
- vessenes 3y agoMost places that recommend llama.cpp for mac fail to mention https://github.com/jankais3r/LLaMA_MPS https://github.com/jankais3r/LLaMA_MPS, which runs unquantized 7b and 13b models on the M1/M2 GPU directly. It's slightly slower, (not a lot), and significantly lower energy usage. To me the win not having to quantize while not melting a hole in my lap is huge; I wish more people knew about it.
- noman-land 3y agoI've been meaning to ask this question as an LLM noob but what exactly is quantizing in this context and why do people do it? I know of quantizing in the digital audio context only.
- dragonwriter 3y agoQuantization is reducing the precision (and size) of values. https://huggingface.co/docs/optimum/concept_guides/quantization https://huggingface.co/docs/optimum/concept_guides/quantizat...
- detrites 3y agoModels in this context are just a big list of numbers. The numbers will have a native "type", for example 32-bit floats. These are numbers like 0.7373663777 or -1.000003663. The 32-bit float type can represent something like 4.3 billion numbers (sort of). It was discovered though, that while models may need this level of precision when creating them ("training"), they don't need it nearly as much after the fact, when simply running them to get results ("inference"). So quantisation is the process of getting that big set of, say, 32-bit floats, and "mapping" them to a much smaller number type. Eg, an 8-bit integer ("INT8"). This is a number in the range 0-255 (or -128 to +127). So, to quantise a list of 32-bit floats, you could go through the list and analyse. Maybe they're all in the range -1.0 to +1.0. Maybe there are many around the value of 0.99999 and 0.998 etc, so you decide to assign those the value "255" instead. Repeat this until you've squashed that bunch of 32-bit values into 8-bits each. (Eg, maybe 0.750000 could be 192, etc.) This could give a saving in memory footprint for the model of 4x smaller, and also makes it able to be run faster. So while you needed 16GB to run it before, now you might only need 4GB. The expense is the model won't be as accurate. But, typically this is on the order of values like 90%, versus the memory savings of 4x. So it's deemed worth it. It's through this process folks can run models that would normally require a 5-figure GPU to run, on their home machine, or even on the CPU, as it might be able to process integers easier and faster than floating point.
- doodlesdev 3y ago> Our system thinks you might be a robot! We're really sorry about this, but it's getting harder and harder to tell the difference between humans and bots these days. Yeah, fuck you too. Come on, really, why put this in front of a _blog post_? Is it that hard to keep up with the bot requests when serving a static page?
- api 3y agoA lot of people just stick cloudflare in front of anything because of cargo cultism. A $5/mo VPS can serve a blog to tens of thousands of people unless you are running something stupidly inefficient. If it’s a static blog make that hundreds of thousands. For millions you might need to splurge on the $10 or $20 per month VPS.
- Spivak 3y agoOr you use the free thing and never think about it?
- hewlett 3y agoYou can either spend $5 per month for VPS for a webserver for your static blog which you now have to secure properly, or you can just stick it on Cloudflare Pages for free
- esquire_900 3y agoCloudflare bot protection adds a slight delay (as in seconds) at best, and completely blocks users like the parent comment at worst. It costs no money, but isn't free either.
- doodlesdev 3y agoWe are talking about different products though, I believe you can have your webpage hosted on Cloudflare Pages or behind Cloudflare CDN without enabling invasive "bot" detection.
- 3y ago
- foobarbecue 3y agoOk I gotta know... what's the art?
- brianjking 3y agoI'll never understand why everyone is spending so much time on a model you cannot use commercially (at all). Secondly, most of us can't even use the model for research or personal use, given the license.
- UncleEntity 3y agoWhy can’t you use it for personal use? I doubt the Facebook Police are going to bust down your door at 3am. …or are they? peeks through curtains
- DustinBrett 3y agoFor me it's because most of what I am learning and trying to do is applicable to LLM's in general. One day the right model will come along, until then I want to play.
- nullc 3y agoThe notion that model weights are copyrightable is absurd on its face. In the US you cannot gain a copyright though sweat of the brow, there must be substantial creative work. Nor does mere collection (e.g. feist v rural) create a copyright. Feeding common crawl to a standard network structure and letting an optimizer do its thing isn't creative, it's just expensive. It requires expertise and skill, sure but so does creating a phone book. The companies working on AI would be foolish to argue for more copyrightability are because it would be hard to conclude the models were copyrightable works without also concluding that the models are unlawful derivatives of the material they were trained on. "Congrats, models can be owned, but regrets: you're bankrupt now because you just committed 4.6 billion acts of copyright infringement carrying statutory damages of $250k each." You might argue that this is far from sure, OKAY-- but parties that take this view will out-compete ones that don't. If it does turn out to be problematic, the people that had something to work from now will pivot to backing their work on something else and will still be ahead of people sitting on their hands. You could see it as a calculated risk, but it seems at least as safe as the one behind the underlying authors of the model weights training on material they're not licensed to distribute.
- Blahah 3y ago
- brucethemoose2 3y agoThere is also CodyCapybara (7B finetuned on code competitions), the "uncensored" Vicuna, OpenAssistant 13B (which is said to be very good), various non English tunes, medalpaca... the release pace maddening.
- acapybara 3y agoAnd let's not forget about Alpacino (offensive/unfiltered model).