5 ms·
Since this is pytorch it should run on cpu anyway. What am I missing?
by Smith42 4y ago
Since this is pytorch it should run on cpu anyway. What am I missing?
- Zetobal 4y agoI guess the simple fact that it didn't before his patch?
- cinntaile 4y agoUsually you just trivially have the model run on cpu or gpu by simply writing .cpu() at specific places, so he's wondering why this isn't the case here.
- markasoftware 4y agothat's literally all I did (plus switching the tensor type). I'd imagine people are posting and upvoting this not because it's actually interesting code but rather just because it runs unexpectedly fast on consumer CPUs and it's not something they considered feasible before.
- roenxi 4y agoThat is vastly underestimating how tricky it is to make novel pieces of software run. There is a huge fringe of people who know how to click things but not use the terminal and a large fringe of people who know how to run "./execute.bat" but not how to write syntactically correct Python. But a lot of those people want to play with LLMs.
- ComplexSystems 4y agoHow are you getting this to run fast? I'm on a top of the line M1 MBP and getting 1 token every 8 minutes.
- markasoftware 4y agoprobably pytorch is very optimized to x86. It's likely using lots of SIMD and whatnot. I'm sure it's possible to get similar performance on m1 macs, but not with the current version of pytorch. Do you have enough ram? (not swapping to disk)?
- ingenieroariel 4y agoTry switching all the .cuda() to .mps() I got a 100x speedup on a different language model on a Macbook M1 Air. https://pytorch.org/docs/stable/notes/mps.html https://pytorch.org/docs/stable/notes/mps.html
- singularity2001 4y agodedicated fork: https://github.com/remixer-dec/llama-mps https://github.com/remixer-dec/llama-mps
- jwitthuhn 4y agoSame experience for me, looks like it is only using one cpu core instead of all of them.
- sva_ 4y agoOr better yet, define a device = 'cpu', and use tensor.to(device).
- tmalsburg2 4y agoIf someone else wrote this comment, would you find it useful?
- progman32 4y agoReading the patch: https://github.com/facebookresearch/llama/compare/main...markasoftware:llama-cpu:main https://github.com/facebookresearch/llama/compare/main...mar... Looks like this is just tweaking some defaults and commenting out some code that enables cuda. It also switches to something called gloo, which I'm not familiar with. Seems like an alternate backend.
- markasoftware 4y agoyou don't actually need to switch to gloo, I just have no idea what I'm doing.
- refulgentis 4y agoLol, all my best work has been when I don’t know what I’m doing and it’s refreshing to see someone moving the ball forward and feeling the same way. Kudos
- rajman187 4y agoGloo is a communication protocol for distributed computation (think along the lines of MPI)
- deleted 4y ago[deleted]