3 ms·
You can run a GPT-3 sized model on a laptop if you're of the patient type. If you want something more usable you need approximately n GBs of GPU memory for n b
by macrolime 4y ago
You can run a GPT-3 sized model on a laptop if you're of the patient type.
If you want something more usable you need approximately n GBs of GPU memory for n billions of parameters. That is when running all optimizations, like 8-bit Matrix Multiplication.
Looking at some GPU cloud prices it seems with a couple of Nvidia datacenter GPUs that totals over 170GB, you're at around $6-7 per hour.
A GPT-3 sized model should also be possible to run on consumer GPUs with a server mainboard, a GPU mining rig case and 8x24GB cards like RTX 3090 (4090 is a bit overpriced) or maybe the upcoming AMD GPUs. The upcoming AMD GPUs in theory have the advantage that they run on PCIe 5.0 instead of PCIe 4.0, so the connection between the GPUs could be quite a bit faster. AMD cards have sucked for deep learning in the past, but it will be interesting to see how the new generation fares.
- timtom39 4y agoI have a build out similar to the one you describe in my basement :P Few comments: The mining GPU setups have almost no bandwidth to the GPUs and are not suitable. You should use hardware like https://www.ebay.com/itm/255657523530 https://www.ebay.com/itm/255657523530 & https://www.ebay.com/itm/255657523546 https://www.ebay.com/itm/255657523546. On moderately priced (1-3K$) enterprise server you can get ~4x16 pcie slots. If your willing to sacrifice bandwidth that gives you 8 3090s at x8 using birutificaiton. PCIE 5.0 is NOT going to help you because you can't buy many lanes of PCIE. Even PCIE 4.0 isn't really useful because the server hardware that supports it is $$$ The difference between running a 3090 at x16 PCIE 3.0 vs x16 PCIE 4.0 is negligible. This is NOT true for 4090s The mining world did develop great cheap power supply options :) Now with above information, for about $10K you can get about 192GB of GPU memory with OK bandwidth/speed.
- lostmsu 4y ago> The difference between running a 3090 at x16 PCIE 3.0 vs x16 PCIE 4.0 is negligible This is very wrong for communication intensive workloads like large model training. If communication is your bottleneck (which is very possible in that case), PCIE 4.0 is 2x faster than PCIE 3.0.
- timtom39 4y agoYou benchmarked? On what platform? Training what? Can you link to this? Or are you just looking at theoretical? In my experience, for what the 3090 offers, it is not really an issue unless you start using birutificaiton.
- lostmsu 4y agoI did not benchmark. This is theoretical and talks about specific conditions: when your training is bottlenecked on communication between GPUs. PCIE 4.0 is exactly 2x bandwidth, and PCIE 5.0 is another 2x more.