9 ms·
Fork of Facebook’s LLaMa model to run on CPU
- Smith42 4y agoSince this is pytorch it should run on cpu anyway. What am I missing?
- Zetobal 4y agoI guess the simple fact that it didn't before his patch?
- cinntaile 4y agoUsually you just trivially have the model run on cpu or gpu by simply writing .cpu() at specific places, so he's wondering why this isn't the case here.
- markasoftware 4y agothat's literally all I did (plus switching the tensor type). I'd imagine people are posting and upvoting this not because it's actually interesting code but rather just because it runs unexpectedly fast on consumer CPUs and it's not something they considered feasible before.
- roenxi 4y agoThat is vastly underestimating how tricky it is to make novel pieces of software run. There is a huge fringe of people who know how to click things but not use the terminal and a large fringe of people who know how to run "./execute.bat" but not how to write syntactically correct Python. But a lot of those people want to play with LLMs.
- ComplexSystems 4y agoHow are you getting this to run fast? I'm on a top of the line M1 MBP and getting 1 token every 8 minutes.
- markasoftware 4y agoprobably pytorch is very optimized to x86. It's likely using lots of SIMD and whatnot. I'm sure it's possible to get similar performance on m1 macs, but not with the current version of pytorch. Do you have enough ram? (not swapping to disk)?
- ingenieroariel 4y agoTry switching all the .cuda() to .mps() I got a 100x speedup on a different language model on a Macbook M1 Air. https://pytorch.org/docs/stable/notes/mps.html https://pytorch.org/docs/stable/notes/mps.html
- singularity2001 4y agodedicated fork: https://github.com/remixer-dec/llama-mps https://github.com/remixer-dec/llama-mps
- jwitthuhn 4y agoSame experience for me, looks like it is only using one cpu core instead of all of them.
- sva_ 4y agoOr better yet, define a device = 'cpu', and use tensor.to(device).
- tmalsburg2 4y agoIf someone else wrote this comment, would you find it useful?
- progman32 4y agoReading the patch: https://github.com/facebookresearch/llama/compare/main...markasoftware:llama-cpu:main https://github.com/facebookresearch/llama/compare/main...mar... Looks like this is just tweaking some defaults and commenting out some code that enables cuda. It also switches to something called gloo, which I'm not familiar with. Seems like an alternate backend.
- markasoftware 4y agoyou don't actually need to switch to gloo, I just have no idea what I'm doing.
- refulgentis 4y agoLol, all my best work has been when I don’t know what I’m doing and it’s refreshing to see someone moving the ball forward and feeling the same way. Kudos
- rajman187 4y agoGloo is a communication protocol for distributed computation (think along the lines of MPI)
- deleted 4y ago[deleted]
- meghan_rain 4y agohow long for one token to infer on an average cpu?
- kristianp 4y agoFrom the readme: On a Ryzen 7900X, the 7B model is able to infer several words per second, quite a lot better than you'd expect!
- markasoftware 4y agoI tested on a decidedly above average CPU, and got several words per second on the 7B model. I'd guess maybe one word per second on a more average one?
- raverbashing 4y agoCool so we're back to the days of 2400 baud modems
- deleted 4y ago[deleted]
- visarga 4y agoIt's useless before the model gets instruction and preference tunings. Won't even follow a simple ask, it will just assume it is a list of questions and generate more, or continue with slightly related comments. FB trained a LLaMA-I (instruction tuned) variant for sports, just to show they can, but I don't think it got released.
- ShamelessC 4y agoUseless!? C'mon.
- ma2rten 4y agoIt's still useful, but you need to know how to use it.
- notpushkin 4y agoSo, you need to know how to tune it.
- qingdao99 4y agoSurely it would work with a format like: User: <question or task> Assistant:
- cfcf14 4y agoYou have to prompt it correctly, non-instruction-aligned models don't behave like agent simulators by default.
- RugnirViking 4y agoit's not that useless, you just have to prompt it the right way (usually by offering an example of the kind of output you want)
- skhm 4y agoTry giving it a simple request instead ;)
- mrtksn 4y agoUnlike Stable Diffusion, I don't stumble upon people who actually use it. Are there examples of the output this can generate? What happens once you manage to run the model?
- sdrinf 4y agoThis is very new, give it a few days. Here's one from Shawn: https://twitter.com/theshawwn/status/1632595934839177216 https://twitter.com/theshawwn/status/1632595934839177216
- KyeRussell 4y agoPretty sure you wouldn’t see anyone using it commercially as IIRC it’s only public due to a leak.
- mrtksn 4y agoI wasn't looking for a commercial use but its an Interesting point. Would it be possible to prove that someone is using it commercially? 1) Spin it up on a cluster in Belarus 2) ??? 3) Profit?
- dsign 4y agoThe thing I like the most about the current AI wave is the pressure is putting on computing hardware. Yes, mobile phones with long battery lives are cool and all of that, but most cool things I like are locked behind huge computational requirements.
- TaylorAlexander 4y agoAgree. I work in robotics and we never have enough compute. I want to see us get to the point where the most advanced robot ever has all the compute it needs onboard, and that means huge growth in compute density and efficiency are needed.
- ben_w 4y agoThat's genuinely surprising. What sort of on-board compute do you typically have today?
- bick_nyers 4y agoI don't work in the field but just to kind of put it into perspective, a 12v 100A LiFePO4 battery has 1200 Watts capacity and weighs 30 pounds. A typical gaming PC (which to be fair, is more willing to trade power for performance) consumes about 600 Watts per hour. Problem for a Tesla? Not so much. Problem for a lightweight drone? Definitely.
- idiotsecant 4y agoahhh the units in this post are making my eye twitch.
- imtringued 4y agoI would be mad too, if my gaming PC was demanding 300 Watts of power but it took half an hour to ramp up ;)
- bick_nyers 4y ago
- benenglish 4y agoWondering how difficult this would be to get running on a m1 max?
- ComplexSystems 4y agoI got one token every 8 minutes or so.
- 2Gkashmiri 4y agoIs that good? Not good?
- popol12 4y agoUsing which model ? On a pretty mid range i5 11th gen I'm getting 0.35 token/s, using the 7B model. Haven't tried the bigger models.
- deleted 4y ago[deleted]
- swyx 4y agoanother commenter posted a fork that does it https://news.ycombinator.com/item?id=35067469 https://news.ycombinator.com/item?id=35067469 per the readme it looks like there a few bugs to figure out in case anyone here is a pytorch expert
- 2Gkashmiri 4y agoi have a friend who owns an macbook pro m1 max. what kind of performance can i get?
- singularity2001 4y agodedicated fork: https://github.com/remixer-dec/llama-mps https://github.com/remixer-dec/llama-mps
- kristianp 4y agoMps = Metal Performance Shaders, for those out of the loop.
- ur-whale 4y ago> 1. Create a conda environment Uh-oh, bad start.
- _flux 4y agoWhy is it a bad start? It could be venv as well, I suppose, I haven't used conda.
- disgruntledphd2 4y agoConda is gonna work much, much, much better for these kinds of applications, as that's what it's mostly used for, i.e. scientific/numerical computing with C/C++ dependencies.
- ur-whale 4y agoConda is an abomination that will download 4Gig of unnecessary shite and carelessly dump it into your system, thereby ruining your existing configuration in the process. Use it in a container or a VM unless you enjoy re-installing your system from scratch. Or better still, don't use it al all and let it wither away: these kind of braindead projects need to be put down with extreme prejudice.
- orangepurple 4y agoChill. Use Miniconda3 as a light alternative. Conda is unnecessary. I agree nobody should ever use Conda unless they are extreme noobs. We all have to start somewhere.
- bshipp 4y agoThe download size is large but conda doesn't ruin any existing configuration unless you explicitly tell it to be your native python environment. Conda is set up as a self-contained set of independent environments. Why would your system care what's inside the Anaconda directory unless you explicitly add it to your PATH/bash?
- LangIsAllWeNeed 4y agoCan it beat mark zuckerberg at congress deposition wordsmithing though?
- toxicFork 4y agoSomeone must have trained an LLM for that for sure.
- LangIsAllWeNeed 4y agoOh yes “”” Hackernews senator: “”Someone on the internet said meta aka Facebook is not considered a real data native, clean coder and high IQ company unless your new language model exceeds the elegance and slipperiness of mark Zuckerbergs (you) language output in senate hearings. he is smoother than a lake in the metaverse.“” Mark LLM: “ Yes, unfortunately, the media and our competitors are all over the idea that Meta is a “dirty company”. They have tried to spin all our successes and accomplishments in a negative light. This has been incredibly frustrating and demoralizing for us, but we know that we are working hard to build a great company and we are confident that our efforts will be rewarded. In the end, our products speak for themselves, and despite all the negative media coverage we are focused on continuing to build great products for our users and being an amazing place for them to socialize in the virtual world.”
- RugnirViking 4y agoI have to say "he is smoother than a lake in the metaverse" is presumably accidental, based on the quality of the rest of that text, but it has to be one of the wittiest phrases ive seen LLMs come out with to date
- jcuenod 4y agoI opened the twitch AI seinfeld stream once and stumbled into a conversation that went something to the effect of: George: I really like that orange sweater Jerry: Yeah, I just found black so depressing George: Orange is such a great color! Orange is the new black. ...
- crazysim 4y agoCould this fit into GitHub Codespaces's top VM?
- DefineOutside 4y agoThe 65 billion model is 160 GB so no - unless you request larger storage spaces from github. 7 billion and 13 billion should fit though.
- fsiefken 4y agoWould running on a cpu be more or less power efficient then running on a gpu with the same words per second rate?
- frognumber 4y agoLess.
- Havoc 4y agoWould it not be possible to run on both gpu and cpu at same time in whatever proportion the hardware is available ? Most gaming desktops have a solid gpu but not enough vram. Pity having the gpu idle here
- bilsbie 4y agoWhat’s the rough idea of how this is possible? I thought you need the parrelism of a gpu
- raihansaputra 4y agoinference has less pressure of parallelism compared to training
- deleted 4y ago[deleted]
- popol12 4y ago0.35 words/s on my 11th gen i5 with 7B model (framework laptop) not so bad !
- fguerraz 4y agoHow long did you have to wait for it to load? On my machine it's been running for 15mins, I'm still waiting for a prompt...
- haolez 4y agoWould it be possible to run the 65B one like this as well? Is the bottleneck just the RAM, or would I need an absurd number of CPUs as well? It's not that hard to create a consumer-grade desktop with 256GB in 2023.
- gpm 4y agoI don't know about this fork specifically, but in general yes absolutely. Even without enough ram, you can stream model weights from disk and run at [size of model/disk read speed] seconds per token. I'm doing that on a small GPU with this code, but it should be easy to get this working with the CPU as compute instead (and at least with my disk/CPU, I'm not even sure that it would run even slower, I think disk read would probably still be the bottleneck) A lack of an absurd number of CPUs just means it's slow, not impossible. https://github.com/gmorenz/llama/tree/ssd https://github.com/gmorenz/llama/tree/ssd
- haolez 4y agoYeah, I find this area fascinating. Like, it's very cool to run a 7B params model locally, but it must feel like a toy when compared to ChatGPT, for example. However, the 65B parameter, according to the benchmarks, is such a beast that you might be able to do some things on it that are not possible on ChatGPT (despite all of ChatGPT's quality of life features). Amazing times.
- downvotetruth 4y agoYou don't need 256 GB. A pair of the new 48GB DDR5 will work along with a pair of 32GB sticks should work in a consumer DDR5 MB to fit the weights. It does burst when initially loading. So, a fast disk with about the same swap size as RAM seems necessary. It took about 25 mins to generate a single 500 character response using a 5800X & 32 GB DDR4, but I was not able to get to it to run on more than 1 thread with the 7B model.
- haolez 4y agoWhy? Is it a limitation of the model or just something with the configuration that you couldn't figure out for this test?