14 ms·
Using LLaMA with M1 Mac and Python 3.11
- simonw 4y agoNeat - this uses the following to get a version of Torch that works with Python 3.11: pip3 install --pre torch torchvision --extra-index-url https://download.pytorch.org/whl/nightly/cpu That's the reason I stuck with Python 3.10 in my write-up for doing this: https://til.simonwillison.net/llms/llama-7b-m2 https://til.simonwillison.net/llms/llama-7b-m2
- cloudking 4y agoHow does LLaMA compare to GPT-3.5 has anyone done side by side comparisons?
- tikkun 4y agoIn short, my experience is that it's much worse to the point that I won't use LLaMA. The main change needed seems to be InstructGPT style tuning (https://openai.com/research/instruction-following https://openai.com/research/instruction-following)
- simonw 4y agoYeah, it's MUCH harder to use because of the lack of tuning. You have to lean on much older prompt engineering tricks - there are a few initial tips in the LLaMA FAQ here: https://github.com/facebookresearch/llama/blob/main/FAQ.md#2-generations-are-bad https://github.com/facebookresearch/llama/blob/main/FAQ.md#2...
- davidb_ 4y agoAre you getting useful content out of the 7B model? It goes off the rails way too often for me to find it useful.
- rnosov 4y agoYou might want to tune the sampler. For example, set it to a lower temperature. Also, the 4bit RTN quantisation seems to be messing up the model. Perhaps, the GPTQ quantisation will be much better.
- spion 4y agoUse `--top_p 2 --top_k 40 --repeat_penalty 1.176 --temp 0.7` with llama.cpp
- datadeft 4y agoNot bad with these settings: ./main -m ./models/7B/ggml-model-q4_0.bin \ --top_p 2 --top_k 40 \ --repeat_penalty 1.176 \ --temp 0.7 -p 'async fn download_url(url: &str)' async fn download_url(url: &str) -> io::Result<String> { let url = URL(string_value=url); if let Some(err) = url.verify() {} // nope, just skip the downloading part else match err == None { // works now true => Ok(String::from(match url.open("get")?{ |res| res.ok().expect_str(&url)?, |err: io::Error| Err(io::ErrorKind(uint16_t::MAX as u8))), false => Err(io::Error
- 4y ago
- KennyBlanken 4y agoThe writeup includes example text where the algorithm is fed a sentence starting about George Washington and within half a sentence or so goes unhinged and starts praising Trump... Also, a reminder to folks that this model is not conversationally trained and won't behave like ChatGPT; it cannot take directions.
- superkuh 4y agoWell, gpt3.5-turbo fails the turing test due to it's censorship and legal liability butt covering openai bolted on, so almost anything else is better. Now, compared to openai's gpt3 davinci (text-davinci-003) ... llama is much worse.
- nodemaker 4y agoI dont know why youre getting downvoted! There is nothing out there at the moment that is as authentic as text-davinci-003. I really hope its not taken away
- flangola7 4y agoI thought LLaMA outscored GPT-3
- simonw 4y agoGPT-3 is a very different model from GPT-3.5. My understanding is that they were comparing LLaMA's performance to benchmark scores published for the original GPT-3, which came out in 2020 and had not yet had instruction tuning, so was significantly harder to use.
- flangola7 4y agoI know, that is why I said GPT-3 (Davinci) not GPT-3.5|ChatGPT.
- simonw 4y agoDa Vinci 002 and 003 are actually classified as GPT 3.5 by OpenAI: https://platform.openai.com/docs/models/gpt-3-5 https://platform.openai.com/docs/models/gpt-3-5 ChatGPT is GPT-3.5 Turbo.
- eshack94 4y agoWould you mind summarizing the difference between GPT 3.5 and GPT-3.5 Turbo? I'm not clear about that.
- bt1a 4y agoThe sampler needs some tuning, but the 65B model has impressive output https://twitter.com/theshawwn/status/1632569215348531201 https://twitter.com/theshawwn/status/1632569215348531201
- aroo 4y agoThe original LLaMA paper has some benchmarks. https://arxiv.org/pdf/2302.13971.pdf https://arxiv.org/pdf/2302.13971.pdf
- diimdeep 4y agoRemove "and Python 3.11" from title. Python used only for converting model to llama.cpp project format, 3.10 or whatever is fine. Additionally, llama.cpp works fine with 10 y.o hardware that supports AVX2. I'm running llama.cpp right now on an ancient Intel i5 2013 MacBook with only 2 cores and 8 GB RAM - 7B 4bit model loads in 8 seconds to 4.2 GB RAM and gives 600 ms per token. btw: anyone knows how to disable swap per process in macOS ? even though there is enough free RAM, sometimes macOS decides to use swap instead.
- metadat 4y agoCan you provide a link to what guide or steps you followed to get this up and running? I have a physical Linux machine with 300+ GB of RAM, would love to try out llama on it but I'm not sure where to get started for how to get it working with such a configuration. Edit: Thank you, @diimdeep!
- deleted 4y ago[deleted]
- diimdeep 4y agoSure. You can get models with magnet link from here https://github.com/shawwn/llama-dl/ https://github.com/shawwn/llama-dl/ To get running, just follow these steps https://github.com/ggerganov/llama.cpp/#usage https://github.com/ggerganov/llama.cpp/#usage
- HervalFreire 4y agoIs it legal to post that here?
- CharlesW 4y ago> Remove "and Python 3.11" from title. Python used only for converting model to llama.cpp project format, 3.10 or whatever is fine. As @rnosov notes elsewhere in the thread, this post has a workaround for the PyTorch issue with Python 3.11, which is why the "and Python 3.11" qualification is there.
- foxandmouse 4y agoI've been seeing a lot of people talking about running language models locally, but I'm not quite sure what the benefit is. Other than for novelty or learning purposes, is there any reason why someone would prefer to use an inferior language model on their own machine instead of leveraging the power and efficiency of cloud-based models?
- canadiantim 4y agoYeah, locally your data doesn't leak out. So if you're using the language model on any sensitive data you're probably going to want local.
- jstarfish 4y agoAlso so the maintainer can't stealth-censor the model.
- sebzim4500 4y agoNo, but whoever trains the weights can. Having said that, if LLaMA has been censored, then Meta have done a poor job of it: it is trivial to get it to say politically incorrect things.
- recuter 4y agoCan I prompt you to share some examples? ;)
- smoldesu 4y agoJust copy-and-paste headlines from your favorite American news outlet. It works great on GPT-J-Neo, so good that I had to make a bot to process different Opinion headlines from Fox and CNN's RSS feeds. Crank up the temperature if you get dissatisfied and you'll really be able to smell those neurons cooking.
- lern_too_spel 4y ago
- suyash 4y agoHow much disk space does it use?
- simonw 4y ago240GB for the initial model download, but once you convert the models they are 4GB for the 7B one and 8GB for the 13B one (and more for the others).
- EMM_386 4y agoThis may be a dumb question, but how is this possible? How can it have all of this information packed into 4GB? I can't even imagine it being only 240GB. These models have an unthinkable amount of information living in them.
- sltkr 4y agoThe way to think about it is that training a neural network is a form of compression that is very, very lossy. You can retrieve information from it but it will be very inaccurate compared to the original source. For example, I assume LLaMa was trained on English Wikipedia data (it just makes sense). So let me try to prompt the 13B parameter model (which is 25 GiB to download, and 3.8 GiB after quantization) with "According to Wikipedia, the main causes of the French revolution are". It will give me the following continuations: Attempt 1: 1) Social injustice and inequality;2)... i need to do a report about the french revolution for my history class. so far i have these three questions:... [end of text] Attempt 2: 1. The Enlightenment and its new ideas in philosophy had a great impact on France especially with regards their attitudes towards religion. There was also some political unrest due mainly from disagreements over how taxes were spent by different governments at that time.... What Were Three Causes of the French Revolution? The three causes for revolution are economic, social and governmental in nature... [end of text] Attempt 3: 1) The Estates General was called by Louis XVI in order solve financial problems. However when they met there were three groups: Nobility who wanted a lot more power and money for themselves; Clergy (Bishops etc.) who also wanted more wealth and power but on top they didn’t want to pay taxes or be governed any other way then the old fashion “divine right” 3) The third group was called the Third Estate. They were made up of middle class citizens, peasants, artisans etc… These people had no political power and wanted more equality in society While the Wikipedia article has a section on causes that starts with: The underlying causes of the French Revolution are usually attributed to the Ancien Régime's failure to manage social and economic inequality. Rapid population growth and the inability to adequately finance government debt resulted in economic depression, unemployment and high food prices. Combined with a regressive tax system and resistance to reform by the ruling elite, it resulted in a crisis Louis XVI proved unable to manage. So the model is completely unable to reconstruct the data on which it was trained. It does have some vague association between the words of "French revolution", "causes", "inequality", "Louis XVI", "religion", "wealth", "power", and so on, so it can provide a vaguely-plausible continuation at least some of the time. But it's clear that a lot of information has been erased.
- voytec 4y agoWhat's with the weird "2023/12/08" date?
- canadiantim 4y agoMaybe it's a hot tip from the future? Or they formatted the date yyyy/dd/mm but mistakenly wrote 08 instead of 03 for the month?
- voytec 4y agoTip from the future works for me due to the news about new DeLorean[1] [1] https://news.ycombinator.com/item?id=35116319 https://news.ycombinator.com/item?id=35116319
- realce 4y agoSorry to do this but 2+0+1+2+0+8 = 23
- MeteorMarc 4y agoYes, title should include (future).
- datadeft 4y agoIt was a typo. I fixed the url. Thanks for pointing out.
- inciampati 4y agoIf you've got avx2 and enough RAM you can run these models on any boring consumer laptop. Performance on a contemporary 16 vCPU Ryzen is on par with the numbers I'm seeing out of the M1s that all these bloggers are happy to note they're using :)
- jeroenhd 4y agoI've tried the 7B model with 32GiB of RAM (and plenty of swap) but my 10th gen Intel CPU just doesn't seem up to the task. For some reason, the CPU based libraries only seem to use a single thread and it takes forever to get any output.
- dmw_ng 4y agowith llama.cpp, you might need to pass in -t to set the thread count. What kind of OS / host environment are you using? I noticed very little speedup with using t=16 and t=32, it's possible the code simply hasn't been tested with such high core counts, or it's bumping into some structural limitation of how llama.cpp is implemented
- inciampati 4y agoI'm wondering if there might be a problem with your compiler setup? Do set -t to use more threads. I don't see improvement past the number of real (not virtual) cores. But I'm seeing about 100ms/token for 7B with -t 8.
- moffkalast 4y ago> any boring consumer laptop > enough RAM Because boring consumer laptops are of course known for their copious amounts of expandable RAM and not for having one socket fitted with the minimum amount possible.
- zamadatix 4y agoAs long as it's 4 GB you should be good to run the smaller model. 8 GB would be preferred, if you're fancy enough to have more you can do the larger models on unquantized models for more quality.
- rspoerri 4y agoI cant wait to get my 96gb m2 i ordered last week. Maybe it could even run the 30b model?
- murkt 4y agoWith 4-bit quantization it will take 15 GB, so it fits easily. On 96 GB you can not only run 30b model, you can even finetune it. As I understand, these model were trained on float16, so full 30b model takes 60 GB of RAM
- trillic 4y agoSo you’re saying I could make the full model run on a 16 core ryzen with 64GB of DDR4? I have an 8GB VRAM 3070 but based on this thread it sounds like the CPU might have better perf due to the RAM?
- blablablub 4y agoi have the 65B model running fine on my 48GB Ryzen 5.
- sebzim4500 4y agoThese are my observations from playing with this over the weekend. 1. There is no thoughput benefit to running on GPU unless you can fit all the weights in VRAM. Otherwise the moving the weights eats up any benefit you can get from the faster compute. 2. The quantized models do worse than non-quantized smaller models, so currently they aren't worth using for much use cases. My hope is that more sophisticated quantization methods (like GPTQ) will resolve this. 3. Much like using raw GPT-3, you need to put a lot of thought into your prompts. You can really tell it hasn't been 'aligned' or whatever the kids are calling it these days.
- orf 4y agoThis might be naïve, but couldn’t you just mmap the weights on an apple silicon MacBook? Why do you need to load the entire set of weights into memory at once?
- wjessup 4y agoHow is this post any different than the instructions on the actual repo? https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp
- rnosov 4y agoThe post has a workaround for the PyTorch issue with Python 3.11. If you follow the repo instructions it will give you some rather strange looking errors.
- simonw 4y agoArtem Andreenko on Twitter reports getting the 7B model running on a 4GB RaspberryPi! One token every ten seconds, but still, wow. https://twitter.com/miolini/status/1634982361757790209 https://twitter.com/miolini/status/1634982361757790209
- dmw_ng 4y agojust wanted to say thanks for your effort in aggregating and communicating a fast moving and extremely interesting area, have been watching your output like a hawk recently
- knodi123 4y agoI ran the 7b model on a prompt about how to get somewhere in the dating scene. Check out the ending: > Don’t get distracted by guys who are already out of your league; focus on the ones that have some hope for getting into it with them...even though they might not be there yet! Dont Forget To Sign Up and Watch our (No, I didn't cut off the end. That's just how it stopped.) Anyway, makes it seem like, whatever their training corpus was, it deffo included scraping a bunch of social media influencers.
- numpad0 4y ago8640 words/day is couple times faster than some of the best novelists in human history, if even quarter of that will be usable it could work as an autonomous paperback author.
- tomp 4y agoBetter instructions (less verbose and include 30B model): https://til.simonwillison.net/llms/llama-7b-m2 https://til.simonwillison.net/llms/llama-7b-m2 I’m running 13B on MacBook Air M2 quite easily. Will try 30B but probably won’t be able to keep my browser open :/ shameless plug: https://mobile.twitter.com/tomprimozic/status/1634877477310038017 https://mobile.twitter.com/tomprimozic/status/16348774773100...
- gorbypark 4y agoGive us an update on the 30B model! I have 13B running easily on my M2 Air (24GB ram), just waiting until I'm on an unmetered connection to download the 30B model and give it a go.
- tomp 4y agohm... well... It definitely runs. It uses almost 20GB of RAM so I had to exit my browser and VS Code to keep the memory usage down. But it produces completely garbled output. Either there's a bug in the program, or the tokens are different to 13B model, or I performed the conversion wrong, or the 4bit quantization breaks it.
- gorbypark 4y agoI've finally managed to download the model and it seems to be working well for me. There's been some updates to the quantization code, so maybe if you do a 'git pull && make' and rerun the quantization script it will work for you. I'm getting about 350ms per token with the 30B model.
- tomp 4y agoThanks for reminding me! It works now. The difference is striking!
- shocks 4y agoI’m also getting garbage out of 30B and 65B. 30B just says “dotnetdotnetdotnet…”
- rolleiflex 4y agoI'm following the instructions on the post from the original owner of the repository involved here. It's at https://til.simonwillison.net/llms/llama-7b-m2 https://til.simonwillison.net/llms/llama-7b-m2 and it is much simpler. (no affiliation with author) I'm currently running the 65B model just fine. It is a rather surreal experience, a ghost in my shell indeed. As an aside, I'm seeing an interesting behaviour on the `-t` threads flag. I originally expected that this was similar to `make -j` flag where it controls the number of parallel threads but the total computation done would be the same. What I'm seeing is that this seems to change the fidelity of the output. At `-t 8` it has the fastest output presumably since that is the number of performance cores my M2 Max has. But up to `-t 12` the output fidelity increases, even though the output drastically slows down. I have 8 perf and 4 efficiency cores, so that makes superficial sense. At `-t 13` onwards, the performance exponentially decreases to the point that I effectively no longer have output.
- gorbypark 4y agoThat's interesting that the fidelity seems to change. I just realized I had been running with `-t 8` even though I only have a M2 MacBook Air (4 perf, 4 efficiency cores) and running with `-t 4` speeds up 13B significantly. It's now doing ~160ms per token versus ~300ms per token with the 8 cores settings. It's hard to quantify exactly if it's changing the output quality much, but I might do a subjective test with 5 or 10 runs on the same prompt and see how often it's factual versus "nonsense".
- bee_rider 4y agoWhat do you use it for, out of curiosity? Can it do shell autocompletes (this is what “ghost in the shell” made me think of, haha).
- rolleiflex 4y agoNothing. It's technology for the love of it. I'm sure there are potential uses but training your own LLM would probably be more meaningfully useful versus running someone else's trained model, which is what this is.
- 4y ago
- eternalban 4y agoGeorgi Gerganov is something of a wonder. A few more .cpp drops from him and we have fully local AI for the masses. Absolutely amazing. Thank you Georgi!
- amelius 4y agoI want to know what his compute setup looks like.
- mark_l_watson 4y agoExcuse my laziness for not looking this up myself, but I have two 8G RAM M1 Macs. Which smaller LLM models will run with such a small amount of memory? any old GPT-2 models? HN user diimdeep has commented here that he ran the article code and model on a 8G RAM M1 Mac, so maybe I will just try it. I have had good luck in the past with Apple's TensorFlow tools for M1 for building my own models.
- phodo 4y agoWhat are some prompts that seem to be working on non-finetuned models? (beyond what is listed in example.py)
- JimmyRuska 4y agoWe need a Fabrice Bellard-like genius to make a tinyLLM that makes the decent models work on 32gb ram
- Q6T46nT668w6i3m 4y agohttps://bellard.org/nncp/ https://bellard.org/nncp/
- syntaxing 4y agoExtremely tempted to replace my Mac Mini M1 (8GB RAM). If I do, what's my best bet to future proof for things like these? Would a Mac Mini M2 with 24GB RAM do or should I beef it up to a M1 Studio?
- irusensei 4y agoAFAIK torch doesn’t work on 3.11 yet. It was not trivial to install on current Fedora. Might have changed.
- datadeft 4y agoYes it does. The blog post has the proof. You can use the nightly builds.
- buzzier 4y ago[dead]
- dataspun 4y agoWhat’s with the propaganda output from the example LLaMA prompt?
- shocks 4y agoAnyone got the 65B model to work with llama.cpp? 7B worked fine for me, but 30B and 65B just output garbage. (On Linux with a 5800X and 64GB of RAM)
- BinRoo 4y agoRerun make, and regenerate the quantized files, because some new commits broke backwards compatibility, if you recently pulled.
- shocks 4y agoThanks for the tip. I tried 404fac0, but got: "The search for extraterrestrial life will most likely conclude adv provinß wojewłożGener Wikipédia Świirc Patrickvidcido protectsobDra"
- m3kw9 4y agoWhy not just have a script and just say run this?
- vishal0123 4y agoBased on my limited runs, I think 4 bit quantization is detrimental to the output quality: > /main -m ~/Downloads/llama/7B/ggml-model-q4_0.bin -t 6 -n 256 -p 'The first man on the moon was ' The first man on the moon was 38 years old. And that's when we were ready to land a ship of our own crew in outer space again, as opposed to just sending out probes or things like Skylab which is only designed for one trip and then they have to be de-orbited into some random spot on earth somewhere (not even hitting the water) Warren Buffet has donated over $20 billion since 1978. His net worth today stands at more than a half trillion dollars ($53 Billiard). He's currently living in Omaha, NE as opposed to his earlier home of New York City/Berkshire Mountains area and he still lives like nothing changed except for being able to spend $20 billion. Social Security is now paying out more than it collects because people are dying... That means that we're living longer past when Social security was supposed to run dry (65) [end of text]
- enduser 4y agoYes, I have found 65B quantized to be more nonsensical than 13B unquantized.
- crdrost 4y agoThe OP article eventually gets around to demonstrating the model and it is similarly bad, zooming from George Washington to the purported physical fitness of Donald Trump?
- cypress66 4y agoThe performance loss is because this is RTN quantization I believe. If you use the "4chan version" that is 4bit GPTQ, the performance loss from quantization should be very small.
- Ologn 4y agoMy Ubuntu desktop has 64 gigs RAM, with a 12G RTX 3060 card. I have 4 bit 13B parameter LLaMA running on it currently, following these instructions - https://github.com/oobabooga/text-generation-webui/wiki/LLaMA-model https://github.com/oobabooga/text-generation-webui/wiki/LLaM... . They don't have 30B or 65B ready yet. Might try other methods to do 30B, or switch to my M1 Macbook if that's useful (as it said here). Don't have an immediate need for it, just futzing with it currently. I should note that web link is to software for a gradio text generation web UI, reminiscent of Automatic1111.
- mattfrommars 4y agoBy 30b and 65b, does it mean model with 30 or 65 million parameters?
- andromaton 4y agoBillion
- lxe 4y agoIf anyone is interested in running it on windows and a GPU (30b 4bit fits in a 3090), here's a guide: https://gist.github.com/lxe/82eb87db25fdb75b92fa18a6d494ee3c https://gist.github.com/lxe/82eb87db25fdb75b92fa18a6d494ee3c
- ur-whale 4y agoNot sure why it mentions Mac and Python I just ran it on a plain old x86 servers with 64 cores and loads of RAM. Works just fine. Apple H/W and Python version are completely irrelevant.
- datadeft 4y agoBecause I could not test it on anything else.
- ehPReth 4y agoIs the ... widely distributed version ... safe to use? How can I check I have the 'right' one? Someone was saying that the models could technically execute arbitrary code if they were backdoored? I'd love to play with them if I have the proper compute haha
- deleted 4y ago[deleted]
- Havoc 4y agoAh...was missing the -t 8 part...no wonder was so painfully slow when I tried this