3 ms·
Which consumer gpu runs llama 70B?
by jpeter 3y ago
Which consumer gpu runs llama 70B?
- sroussey 3y agoProsumer gear. MacBook Pro M3 Max.
- mlsu 3y agoA Mac with a lot of unified RAM can do it, or a dual 3090/4090 setup gets you 48gb of VRAM.
- jadbox 3y agoDoes this actually work? I had thought that you can't use SLI to increase your net memory for the modal?
- speedgoose 3y agoIt works. I use ollama these days, with litellm for the api compatibility, and it seems to use both 24GB GPUs on the server.
- benjaminwootton 3y agoI’ve got a 64gb Mac M2. All of the openllm models seem to hang on startup or on API calls. I got them working through GCP colab. Not sure if it’s a configuration issue or if the hardware just isn’t up to it?
- benreesman 3y agoValiant et al work great on my 64Gb Studio at Q4_K_M. Happy to answer questions.
- wahnfrieden 3y agoTry llama.cpp with Metal (critical) and GGUF models from TheBloke Or wait another month or so for https://ChatOnMac.com https://ChatOnMac.com
- brucethemoose2 3y agoA single 3090, or any 24GB GPU. Just barely. Yi 34B is a much better fit. I can cram 75K context onto 24GB without brutalizing the model with <3bpw quantization, like you have to do with 70B for 4K context.
- speedgoose 3y agoCan it produce any meaningful outputs with such an extreme quantisation?
- brucethemoose2 3y agoYeah, quite good actually, especially if you quantize it on text close to what you are trying to output. Llama 70B is a huge compromise at 2.65bpw... This does make the much "dumber." Yi 34B is much better, as you can quantize it at ~4bpw and still have a huge context.
- lossolo 3y agoHow would you compare mistral-7b-instruct 16fp (or similar 7b/13b model like llama2 etc) to Yi-34b quantized?
- brucethemoose2 3y ago34B is better. Quantization hurts some, especially in "pro" non chat use cases like RAG, but the increased parameter count makes models so much smarter in comparison. The perplexity graph here is a pretty good illustration: https://github.com/ggerganov/llama.cpp/pull/1684 https://github.com/ggerganov/llama.cpp/pull/1684 YMMV, as Mistral and Yi are not necessarily comparable like different sizes of llama, and it depends on the task.
- deleted 3y ago[deleted]