3 ms·
I use two 3090s to run the 70b model at a good speed. Takes 32 gigs of vram, more depending on context. I tried CPU+GPU (5900X + 3090) but with extended context
by ImprobableTruth 3y ago
I use two 3090s to run the 70b model at a good speed. Takes 32 gigs of vram, more depending on context. I tried CPU+GPU (5900X + 3090) but with extended context it's slow enough that I wouldn't recommend it (~1 token/s). CPU only gets "let it run over night" slow. Works ok-ish for with a small context though (even if it's still "non-interactive" slow).
- josephg 3y agoWhat’s the difference in output quality between that and the 33b parameter model? That would fit entirely in vram, right?
- brucethemoose2 3y agoThe 33B model is llama V1. Facebook reportedly held back 34B llama v2 because it failed some safety metrics. So... Generally the quality is worse, but the available set of finetunes is totally different. Some llama v1 33b finetunes are not available in 70B, and extremely good at their niche. Also 70B should get more than 1 token/sec on a single 3090 offloaded to CPU. I dunno what framework op is using.
- MuffinFlavored 3y agowhat example niche?
- brucethemoose2 3y ago- roleplaying (chronos merged with airoboros) - theraputic/friend style chat (Samantha) - translation (various single language finetines) - medical advice (can't remember this one) This is non exhaustive. And Llama V2's extended native context does really help some niches (like storytelling) that a few 33B models are still pretty good at.
- cosmojg 3y ago> medical advice (can't remember this one) You're probably thinking of Clinical Camel: https://huggingface.co/augtoma/qCammel-70-x https://huggingface.co/augtoma/qCammel-70-x
- brucethemoose2 3y agoThis is a 70B tune. And its new to me. Looks interesting!
- a20eac1d 3y agoAny chance you could point me in the right direction on how to set something like this up? Right now, I'm using pure CPU Llama but only the 17B version, based on I believe llama.cpp. How do I mix both CPU and GPU together for more performance?
- brucethemoose2 3y agoThe easy way: download koboldcpp. Otherwise you have to compile llama.cpp (or kobold.cpp) with opencl or cuda support. There are instructions for this on the git page. Then offload as many layers as you can to the gpu with the gpu layers flag. You will have to play with this and observe your gpu's vram.