3 ms·
I could really use some advice for making some deep learning hardware purchases for my university lab. It has been a pretty painful experience speaking to the
by chriskanan 3y ago
I could really use some advice for making some deep learning hardware purchases for my university lab.
It has been a pretty painful experience speaking to the companies due to shortages and because of all the limitations of non-H100 chips. For getting a machine with 4 H100s and NVLINK, I was told I'd have to wait one year. With that long of a wait, I might as well wait for the next generation. Without NVLINK, I could potentially get this machine sooner. I struggled to find any reports of the additional complexities and how much slower the machine would be.
I'm basically looking to buy two machines. One for under $40K and one for up to $100K. The $100K machine is intended for training LLMs, and the other would have multiple GPUs where the other would just have a bunch of GPUs for my PhD students to use. For the $40K machine, I'm being told by companies I could only put in 2-3 L40S cards, and they are really pushing these as the only viable option.
In contrast, the last time I had $40K for hardware I got two machines with 4 RTX A5000s and companies were a lot more responsive and seemed more helpful. It is unlikely that I'll have $100K for hardware again, so I'm very reluctant to go with cloud computing since that decreases my hardware budget and I need these to last 3+ years.
I wish I could go with 4090s, but the limited VRAM and that NVIDIA has disabled functionality needed for multi-GPU training makes that pretty much a non-starter.
- sjg 3y agoWe had the same issue for our lab when we were spec'ing up a similar install. The one thing I didn't fully realise is the need of having an Enterprise subscription to run the A100s, H100s and other cards. The drivers for these cards are behind a paywall and for academia it looks like it's around $150 per card, per year to run (if you want to run in a DC - https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/solutions/resources/documents1/Virtual-GPU-Packaging-and-Licensing-Guide.pdf https://www.nvidia.com/content/dam/en-zz/Solutions/design-vi... ). We bought 3, A100s for the servers and 3 L40's (not L40S) this was limited due to space in the severs. The NVLINKs can be added later (it's just a bridge between the cards) so if you can get them without NVLINK I would. (NVLINK works well with cards stacked vertically rather than horizontally). We also bought a desktop system with dual 4090s (£15k - $18.5k) to start on smaller models before scaling up on our servers with the real grunt work happens. This worked well as many problems can be solved with smaller models before going all out needing 4 H100s. Hope this helps with the planning. Feel free to reach out if you want to chat more about our setup. (My HackerNews Username is the same as github and you'll find my email address there)