3 ms·
Happy to see this getting attention, lots of great open models are being worked on and I can't wait to see something people at home with a 3090 could use.
by amrb 4y ago
Happy to see this getting attention, lots of great open models are being worked on and I can't wait to see something people at home with a 3090 could use.
- tarruda 4y agoYou should be able to run Llama 30B q4 on a RTX 3090. Check the table here: https://bellard.org/ts_server/ https://bellard.org/ts_server/ 30B/q4 requires 20GB of RAM while 3090 has 24GB.
- mromanuk 4y agowhich one is the next incarnation of performance/price/goodness after the RTX 3090 for running a mini AI homelab?
- espadrine 4y agoTo increase performance, you need bigger foundation models, like LLaMA 65B. You will always be RAM-bound for that (~40GB for LLaMA 65B INT4). So the next step is to upgrade your motherboard to have multiple GPU slots. If you really wanted to stick to a single-GPU setup, you could upgrade to RTX A6000, but it is more expensive than two RTX 3090 while holding the same amount of RAM.
- yieldcrv 4y agoIts a dual slot gpu What benefit does that offer over 2 GPUs?
- espadrine 4y agoThe two GPUs will need to exchange information in order to complete inference, by having one GPU hold half of the network weights, and the other holding the other half. That transmission of information will be limited by the PCIe bandwidth; for instance, 30 GB/s with v4x16. Meanwhile the A6000 VRAM bandwidth is 768 GB/s (= 16 Gb/s (GDDR6) × 384 bit-width ÷ 8 bits per byte).
- coolspot 4y agoTwo 3090 can be connected using NVLink bridge which is much faster than PCI-E.
- espadrine 4y agoI couldn’t find precise information on what bandwidth you’d get with NVLink on the 3090. To be fair, though, if all we do is inference, using Huggingface pipeline parallelism, the amount of data transferred is pretty small: 8192×2×n_tokens bytes; for most uses, with a recent PCIe setup, that bottleneck will take less than 0.1 ms per token generated, which may not be the dominant latency. (Also, the RTX 3090 has faster VRAM, >900 GB/s, than the A6000, because it is GDDR6X.)
- azeirah 4y agoNot entirely certain about this, but I believe that because Apple M-series GPUs use system memory you could possible run the larger models on accessible (albeit expensive) consumer hardware. Need 64GB for 30B And 128GB for the 65B I'm not sure about the performance, but I think it should be ok? Especially given how much Apple has been investing in what -- I believe they call -- neural cores? Here's some more context: https://news.ycombinator.com/item?id=35105364 https://news.ycombinator.com/item?id=35105364
- michannne 4y agoSeconding this. Llama 30B 4-bit has amazing performance, comparable to GPT-3 quality for my search and novel generating use-cases, and fits on a single 3090. In tandem with 3rd party applications such as Llama Index and the Alpaca LoRa, GPT-3 (and potentially GPT-4) has already been democratized in my eyes.