5 ms·
I wonder what the memory requirements would be to run such a large model. I'd love to be able to run this model, alas my MacBook can barely run toy models.
by kif 4y ago
I wonder what the memory requirements would be to run such a large model. I'd love to be able to run this model, alas my MacBook can barely run toy models.
- px43 4y agoHell, I'd love to be able to buy a $30k server to run these models. I think to run BLOOM required something more along the lines of a $200k server.
- DeathArrow 4y agoNo need to spend $30k, use Azure or AWS.
- wincy 4y agoYep. It’s expensive to spin up an A100 80GB instance but not THAT expensive. Oracles cloud offering (first thing to show up in google search I know you probably won’t use them and it seems extra expensive) is $4.00 per hour. If you are motivated to screw around with this stuff there’s definitely options.
- gnramires 4y ago> If you are motivated to screw around with this stuff there’s definitely options. Erm, for inference that is. Training is definitely out of question for individuals I believe (unless you use much smaller models?).
- dilyevsky 4y agoGCP spot price for A100 80g gpu is only $1.25 and they give you $300 of credit when you open a new acc
- dragonwriter 4y agoUnless its for something you want to happen whenever and don't mind be dumped in process, shouldn't we look at on-demand, not spot, prices?
- dilyevsky 4y agoIt’s fine for testing it out or serving short queries. Your data can still remain on PV if the vm gets yanked
- vintermann 4y agoI can see some issues with uploading a leaked model to a cloud provider.
- londons_explore 4y agoWith code modifications, it should be possible to run this with a very modest machine as long as you're happy for performance to suck. Transformer models typically need to read all the weights per 'word' output, so if your model is 20GB and you have not enough ram or vram, but have an SSD that reads 1GB/sec, expect 3 words per minute output speed. However, code changes are necessary to achieve that, although they won't be crazy complex.
- moffkalast 4y agoThe most time/cost optimal solution is probably to buy 32 or 64 gigs of ram. That'll still be slow but most people are already half way there.
- esperent 4y agoDoesn't it need to be GPU ram?
- nl 4y agoThey are saying you can run it on a CPU by doing this: > However, code changes are necessary to achieve that, although they won't be crazy complex. This is technically true. It will be very slow though. However, give it 6 months and I think we might see an order of magnitude increase in speed on CPUs. This will still be too slow to be very useful though.
- koheripbal 4y agoThat will be very VERY slow. Pcie bandwidth is way too slow.
- moffkalast 4y agoShould be like an order of magnitude faster than trying to run it from a NVMe still, no? I've ran some small flan models from RAM and it was fine, but yeah it's not exactly realtime.
- 4y ago
- VadimPR 4y agoTrue, and that's why there is a project that is using volunteered, distributed GPUs to run BLOOM/BLOOMZ: https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals, http://chat.petals.ml http://chat.petals.ml.
- sourcecodeplz 4y agoI've tried this but compared to ChatGPT... let's just say it's in a different league.
- permo-w 4y agoyou can - slowly - run Bloom 3b and 7b1 on the free (trial) tiers of Google Cloud Compute if you use the low_cpu_mem_usage parameter of from_pretrained
- q1w2 4y agoYou would need over 65GB of RAM. There are consumer GPUs that have 48GB of RAM, and can be tethered together with NVLink. I wonder if that would work.
- make3 4y agoyou can rent a vm on aws to run it