5 ms·
40B is pretty large, right? I expect it would take 70GB or so of RAM to run it. That's some expensive hardware ($10,000 or more).
by amstan 3y ago
40B is pretty large, right? I expect it would take 70GB or so of RAM to run it. That's some expensive hardware ($10,000 or more).
- esperent 3y agoI briefly looked at prices a few days ago, I think you can rent an 80gb GPU for about $2.50 an hour. Then you just pay while you're using it. Someone else might be better able to confirm the pricing, but in any case you don't need to purchase the hardware.
- amstan 3y agoI'm less interested in using someone else's computer (not as much, but similar to how i'm disinterested in an API from someone like OpenAI), would rather pay the upfront hardware cost than worry about how many tokens i generate (kind of hinders creativity and excitment about it).
- gymbeaux 3y ago70GB of RAM would cost around $150 these days depending on how you get there. 64GB (2x32GB) of DDR4 is around $140 then another 8GB stick would be around $15. Used DDR3 ECC would be roughly half that.
- Semaphor 3y agoI think they are talking about VRAM
- gymbeaux 3y agooop you're right... I know once the model is trained they run on the CPU/regular RAM
- drdaeman 3y agoAnything that runs on GPU can be run on CPU, the only question is how much slower it's gonna be.
- Semaphor 3y agoI know, I have a weak GPU but 64 GB RAM. Using those models works, but it’s more "ask a question, then do something else for a while, while your fans spin up" ;)
- amstan 3y agoI've been running vicunda-13B on a workstation that can be aquired from ebay for about $500. It's slow compared to online services, but probably slightly faster than text to speech would recite its output, so plenty. Falcon 40B is probably too much for it, but apparently there's similar cheap hardware that could work.
- GaggiX 3y agoWith 4-bit quantization should take less than 40GB of VRAM/RAM.
- valvar 3y agoThat much good RAM in itself isn't super expensive. So does the rest of the hardware have to be particularly powerful?
- esquire_900 3y agoSome people use second hand P40 GPUs, which go for around 200-300$. Combine 3 of them with SLI and you've got 72GB of VRAM for less then $1000
- gengolas 3y agoI do use a P40 for my machine learning box, but I'm curious how you put three on the same system, given they need a CPU power plug and a pci-e port. Then, to cool them, you need to plug your own cooling system, requiring more specific power plugs to be available. What kind of chassis, motherboard, power unit you use to do that? It'll certainly will cost more than $1000 anyway, especially since you also need a decent amount of RAM to preload the models before you move them to the GPUs.
- esquire_900 3y agohttp://nonint.com/ http://nonint.com/ has some interesting posts about how he build a custom server to house 8 GPU's (3090's in this case). You're right that that will set you back more than $1000, though I was only referring to the GPU's themselves.
- amstan 3y agoWoah, that's a cool direction. Thank you! I'll explore this.
- washadjeffmad 3y agoP40s are kind of a meme. Using ggmls has roughly the same performance at a fraction of the wattage on a dual-channel DDR5 system. I still use GPTQ for 30B, but even CPU generates quickly enough at q5_1 on modern hardware.
- deleted 3y ago[deleted]
- Kelamir 3y ago> https://github.com/ggerganov/llama.cpp/issues/1602#issuecomment-1570827592 https://github.com/ggerganov/llama.cpp/issues/1602#issuecomm... Here somebody quantized it down to 29929.56MB .
- dheera 3y agoStupid question but for feed-forward models why do we not yet have some kind of CPU RAM memory swap mechanism? Why is Pytorch still trying to load the whole damn model into GPU RAM at once and then complaining when it can't, instead of swapping portions of the model to CPU RAM, or hell, even SSD? Sure, it might be a lot slower, but that's a lot better than "I give up, go buy $20K worth of hardware"
- coolspot 3y agoThat’s what llama.cpp does, including offload to a disk. It allows you to run models as big as any combination of your VRAM, RAM or disk. But in the end, if it doesn’t fit into GPU VRAM, it will be slow. For example Guanaco-33B generates ~10 token per second running fully from VRAM of my 3090, the ~1 token/second running from DDR4 RAM of my Ryzen. I would imagine it would do like a token per minute from NVM SSD.
- anonymousDan 3y agoIs it memory/disk bandwidth bound or latency bound?
- mike_hearn 3y agoMemory bandwidth bound.
- sebzim4500 3y agoI haven't run the numbers, but I would expect that doing that would make it slower than just running on the CPU.