4 ms·
I have 128 GB in my computer with a 5850x, it allows me to run and load the 180B falcon and 70B llama2 LLMs in llama.cpp, although with different quantization.
by tyfon 3y ago
I have 128 GB in my computer with a 5850x, it allows me to run and load the 180B falcon and 70B llama2 LLMs in llama.cpp, although with different quantization.
Speed is actually not that bad either.
- corn13read 3y ago[dead]
- mosselman 3y agoIs there some documentation on how to run this setup? How fast is your setup?
- rnk 3y agoI'm doing this on a mac studio with 128gb too. I'm using llama.cpp.
- acchow 3y agoSince you get GPU acceleration (because of the unified memory), I imagine this is probably much faster than the PC setup? Edit: Seems some people are getting 1-2.6 tokens/sec on Ryzen (no GPU acceleration), Llama 70B quantized https://www.reddit.com/r/LocalLLaMA/comments/15rqkuw/llama_2_q4_k_s_70b_performance_without_gpu/ https://www.reddit.com/r/LocalLLaMA/comments/15rqkuw/llama_2... Whereas Mac Studio gets 13 tokens/sec https://blog.gopenai.com/how-to-deploy-llama-2-as-api-on-mac-studio-m2-ultra-and-enable-remote-api-access-7c4e6423b2dd https://blog.gopenai.com/how-to-deploy-llama-2-as-api-on-mac...
- stoatmagoats 3y agoFriendly internet stranger’s input: - you don’t get GPU acceleration just by using unified memory. Llama.cpp still only uses the CPU on Apple Silicon chips. - the difference in tokens/sec is likely attributable to memory bandwidth. Mac Studios with the base Max chip have 400 GB/s memory bandwidth compared to around 50 GB/s for the Ryzen 5000 series CPUs
- spott 3y agoLlama.cpp defaults to using metal. [0] [0] https://github.com/ggerganov/llama.cpp#metal-build https://github.com/ggerganov/llama.cpp#metal-build
- acchow 3y agoWhat's your generation speed?