13 ms·
Running a 180B parameter LLM on a single Apple M2 Ultra
- logicchains 3y agoPretty amazing that in such a short span of time we went from people being amazed how powerful GPT3.5 was upon its release to people being able to run something equivalently powerful locally.
- zagfai 3y agohowever, GPT3.5 did not surprised me but GPT4 did. 3.5 just a kid.
- regularfry 3y ago4-bit quantised model, to be precise. When does this guy sleep?
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- ramesh31 3y ago>When does this guy sleep? I don't think he has since July.
- esafak 3y agoWith his name recognition he could easily raise over $10m in funding for a seed round and sleep well, if he wanted.
- swyx 3y agohe already has? lol https://news.ycombinator.com/item?id=36215651 https://news.ycombinator.com/item?id=36215651
- beardedwizard 3y agoWhat ever is he doing, we must protect this man at all costs.
- homarp 3y agohttps://www.reddit.com/r/LocalLLaMA/comments/16bynin/falcon_180b_initial_cpu_performance_numbers/ https://www.reddit.com/r/LocalLLaMA/comments/16bynin/falcon_... has some more data like sample answers with various level of quantizations and https://huggingface.co/TheBloke/Falcon-180B-Chat-GGUF https://huggingface.co/TheBloke/Falcon-180B-Chat-GGUF if you want to try
- sbierwagen 3y agoThe screenshot shows a working set size of 147,456 mb, so he's using the mac studio with 192 gb of ram?
- deleted 3y ago[deleted]
- pella 3y agoIs this an M2 Ultra with 192 GB of unified memory, or the standard version with 64 GB of unified memory?
- zargon 3y ago4-bit quantized 180B will not fit in 64GB. You'll need need over 100 GB for that.
- rvz 3y agoTotally makes sense for C++ or Rust based AI models for inference instead of the over-bloated networks run on Python with sub-optimal inference and fine-tuning costs. Minimal overhead or zero cost abstractions around deep learning libraries implemented in those languages gives some hope that people like ggerganov are not afraid of the 'don't roll your own deep learning library' dogma and now we can see the results as to why DL on the edge and local AI, is the future of efficiency in deep learning. We'll see, but Python just can't compete on speed at all, henceforth Modular's Mojo compiler is another one that solves the problem properly with the almost 1:1 familiarity of Python.
- deleted 3y ago[deleted]
- brucethemoose2 3y agoThe actual inference is not run in Python in PyTorch, and its usually not bottlenecked by it. The problem is CUDA, not Python. LLMs are uniquely suited to local inference in projects like GGML because they are so RAM bandwidth heavy (and hence relatively compute lite), and relatively simple. Your kernel doesn't need to be hyper optimized by 35 Nvidia engineers in 3 stacks before its fast enough to start saturating the memory bus generating tokens. And yet its still an issue... For instance, llama.cpp is having trouble getting prompt ingestion performance in a native implementation comparable cuBLAS, even though they theoretically have a performance advantage by using the quantization directly.
- survirtual 3y agoPython is generally just the glue language for underlying, highly optimized c++ libs. The improvements aren't just about languages. I would imagine facebook is less focused on inference, so didn't bother to make a highly optimized LLM inference engine. There also just isn't a business case for CPU-bound LLMs at an enterprise scale, so why code for that? Additionally, llama.cpp can be called by python and python could still do all the glue. There is no language war. Use whatever tool is necessary to achieve effective results for accomplishing the mission.
- PartiallyTyped 3y ago
- adam_arthur 3y agoEven a linear growth rate of average RAM capacity would obviate the need to run current SOTA LLMs remotely in short order. Historically average RAM has grown far faster than linear, and there really hasn't been anything pressing manufacturers to push the envelope here in the past few years... until now. It could be that LLM model sizes keep increasing such that we continue to require cloud consumption, but I suspect the sizes will not increase as quickly as hardware for inference. Given how useful GPT-4 is already. Maybe one more iteration would unlock the vast majority of practical use cases. I think people will be surprised that consumers ultimately end up benefitting far more from LLMs than the providers. There's not going to be much moat or differentiation to defend margins... more of a race to the bottom on pricing
- tomohelix 3y agoRAM is easy. The hard part is making the unified memory SOC like Apple's. From what I know, Apple performance is almost magic. And whatever Apple is making, they are at peak capacity already and they can't make more even if they want to. Nobody else has a comparable technology. Apple is in its own league.
- AnthonyMouse 3y agoApple is just using a wide memory bus, the same as GPUs and server-class x86 CPUs do. It's not even hard, it's just not something desktop CPUs previously had any use for so the current sockets don't support it. And you could do the same thing without even changing the socket by including RAM on the CPU package as an L4 cache. Some of the Intel server CPUs are already doing this.
- ls612 3y agoFor me the test is; when will a Siri-LLM be able to run locally on my iPhone at at least GPT-4 levels? 2030? Farther out? Never because of governments forbidding it? To what extent will improvements be driven by the last gasps of Moore’s Law vs by improving model architectures to be more efficient?
- 3y ago
- superkuh 3y ago[flagged]
- sbierwagen 3y agoM2 Mac Studio with 192gb of ram is US$5,599 right now.
- deleted 3y ago[deleted]
- superkuh 3y ago[flagged]
- piskov 3y agoYou do understand that you can connect thunderbolt external storage (not just usb3 one)?
- diffeomorphism 3y agoThat does not really make 1tb of non-upgradable storage in a $5k+ device any less ridiculous though.
- yumraj 3y agoThat is true, but a whole separate discussion. It applies to RAM too. My 32GB Mac Studio seemed pretty good before the LLMs.
- yumraj 3y agoIt’s not useless. It seems a Thunderbolt/USB4 external NVME enclosure can do about 2500-3000 MB/s which is about half of internal SSD. So not at all bad. It’ll just add an additional few tens of seconds while loading the model. Totally manageable. Edit: in fact this is the proper route anyway since it allows you to work with huge model and intermediate FP16/FP32 files while quantizing. Internal storage, regardless of how much, will run out quickly.
- randomopining 3y agoIs there any actual usecases to run this stuff on a local computer? Or are most of these models actually suited to run on remote clusters?
- logicchains 3y agoThe use-case is you want to generate pornographic, violence-depicting or politically-incorrect content, and would rather buy a powerful computer than rent a server (or you already own a powerful computer).
- catchnear4321 3y agoit seems infinitely cheaper to jailbreak poorly implemented publicly-facing gimmick LLM “use cases” and “demonstrations” that rely on / thinly veneer commercial apis. (this is not financial advice and i am not a financial advisor.)
- beardedwizard 3y agoYou what? You can run smaller and plenty powerful models on a m1 MacBook. Idk what the porn and violence angle is but maybe keep that one to yourself.
- logicchains 3y agoOne of the largest use-cases for local LLMs is NSFW chatbots, like DIY Replika, AI girl/boyfriends, as the hosted services are too censored to be used for this. Yes there are smaller models, but they're not as intelligent. Similarly people using LLMs as a writing aid need to use local ones if they're writing a story (or .e.g DnD campaign) involving violence, as the hosted ones are generally unwilling to narrate graphic violence, and the smarter the model, the better the story quality. Given that censorship is one of the biggest complaints about the hosted LLMs, it should be no surprise that some of the main use-cases driving local LLMs are those involving creating content that censored LLMs are unwilling to create.
- acdha 3y ago
- tiffanyh 3y agosystem_info: n_threads = 4 / 24 Am I seeing correctly in the video that this ran on only 4 threads?
- wmf 3y agoIt's using the GPU so I guess not that many CPU threads are needed to feed the GPU.
- m3kw9 3y agoOpenAIs moat will soon largely be UX. Anyone can do plugins, code etc but when operating by everyday users the best UX wins after LLM becomes commodified. Just look at stand alone digital cameras vs mobile phone cams from Apple.
- ZoomerCretin 3y agoGPT4 is still leagues ahead of the competition. Open source LLMs will be used more widely, but for the most demanding tasks, there is no alternative for GPT4.
- eurekin 3y agoAnecdata confirmation: I've been toying around with LLMs for simple fun stuff, but when it comes to real work, GPT-4 delivers in spades. I have cut many hours of debugging thanks to it. I could find issues easily, on-call in short conversation, when previously that was reserved as post mortem task. Even reading documentation is nothing like before: once, I was looking for a single command to upload and presign a object in S3. SDK has tens of methods, which require careful scanning, if they do what I want. Going through documentation thoroughly would've taken me hours. GPT-4 simply found, no, there's no operation for that immediately.
- smoldesu 3y ago> but when operating by everyday users the best UX wins Is that not why OpenAI is ahead right now? For free, you can have access to powerful AI on anything with a web browser. You don't need to wait for your SSD to load the model, page it into memory and swap your preexisting processes like it would on a local machine. You don't need to worry about the local battery drain, heat, memory constraints or hardware limitations. If you can read Hacker News, you can use AI. Given the current performance of local models, I bet OpenAI is feeling pretty comfortable from where they're standing. Most people don't have mobile devices with enough RAM to load a 13b, 4-bit Llama quantization. Running a 180B model (much less a GPT-4 scale model) on consumer hardware is financially infeasible. Running it at-scale, in the cloud is pennies on the dollar. I'm not fond of OpenAI in the slightest, but if you've followed the state of local models recently it's clear why they keep coming out ahead.
- doctoboggan 3y agoGeorgi is doing so much to democratize LLM access, I am very thankful he is doing it all on apple silicon!
- growt 3y agoSo how much ram did the machine have?
- ViktorBash 3y agoIt's refreshing to see how fast open LLMs are advancing in terms of the models available. A year ago I thought that besides for the novelty of it, running LLMs locally would be nowhere close to stuff like OpenAI's closed models in terms of utility. As more and more models become open and are able to be run locally, the precedent gets stronger (which is good for the end consumer in my opinion).
- Havoc 3y agoGreat progress, but I also can't help but feel a sense of apprehension on the access front. An M2 Ultra while consumer tech is affordable to a fairly small % of the world population.
- two_in_one 3y agoJust wondering what are local LLMs used for today? So far they look more like a.. promising.