4 ms·
Currently running 65B on my 96GB M2 Max.. it's pretty good.
by 19h 4y ago
Currently running 65B on my 96GB M2 Max.. it's pretty good.
- sheepscreek 4y agoNice. I didn’t know Pros could be bumped above 64GB. What did that setup set you back?
- 19h 4y agoGermany -- 5.529,00 € excl. Apple Care which is 149,99 €/Year
- wincy 4y agoFor the super high provisioned laptops it seems like Applecare is a steal.
- satvikpendem 4y agoMeanwhile I just bought 64 GB RAM to try the 65B model on my desktop for like 150 bucks, lol
- snek_case 4y agoWhat seems unclear to me is, does this use the onboard GPU or neural accelerator at all, or is it all CPU-based? Cool if it's able to run 100% CPU-based because that makes portability and deployment a lot easier. Makes this code a lot more accessible.
- ggerganov 4y agoIt currently uses only the CPU via ARM NEON intrinsics - no GPU, no ANE and no Apple Accelerate. Plan is in the future to utilize respective SIMD intrinsics for other architectures (AVX, WASM SIMD, etc) and also add other more accurate quantization approaches. It's actually not a lot of work and I have most of the stuff ready, so hopefully soon! Edit: AVX2 support has just been added
- lxe 4y agoAre you observing higher tokens/second throughout on Apple silicon (you’ve been mentioning 20/second) than running it in PyTorch on a CUDA GPU such as a 3090?
- ggerganov 4y agoI haven't run the original PyTorch model not a single time! I just look at the code and port it. I don't have the hardware to run it.
- simonw 4y agoAre you funded at all? If you need extra hardware I'm sure the community could make that happen.
- MacsHeadroom 4y agoA 4090 gets 30 tokens/second with LLaMA-30B, which is about 10 times faster than the 300ms/token people are reporting in these comments. (20 tokens/second on a Mac is for the smallest model, ~5x smaller than 30B and 10x smaller than 65B)
- lxe 4y agoHow are you getting 20 tokens/second? I'm getting 2.6 tokens/s on 3090 with int4 prequantized model. Is 4090 so much faster?
- 19h 4y agoI couldn't get the mps python version to run, it's insane how much setup it requires... I don't think the Gerganov C++ version uses CoreML or Neural Engine. I previously tried to play with ANE based off prior reverse engineering work [0] but couldn't get it to work nicely. It's actually beyond me how Gerganov's version performs so well -- the output quality of the non-quantised version running on A100 (AWS) isn't noticably better than the one I'm getting. [0] https://i.blackhat.com/asia-21/Friday-Handouts/as21-Wu-Apple-Neural_Engine.pdf https://i.blackhat.com/asia-21/Friday-Handouts/as21-Wu-Apple...
- wjessup 4y agoSame here. 53ms a token. pretty fast!