6 ms·
An answer to: why is Nvidia making GPUs bigger when games don't need them?
I've had this discussion a few times recently including with people who are pretty up-to-date engineers working in IT. It's true that the latest offerings from Nvidia both in the gaming and pro range are pretty mind-boggling in terms of power consumption, size, cost, total VRAM etc... but there is actually a very real need for this sort of thing driven by the ML community, maybe more so than most people expect.
Here is an example from a project I'm working on today. This is the console output of a model I'm training at the moment:
-----------------------------
I1015 20:04:51.426224 139830830814976 supervisor.py:1050] Recording summary at step 107041.
INFO:tensorflow:global step 107050: loss = 1.0418 (0.453 sec/step)
I1015 20:04:55.421283 139841985250112 learning.py:506] global step 107050: loss = 1.0418 (0.453 sec/step)
INFO:tensorflow:global step 107060: loss = 0.9265 (0.461 sec/step)
I1015 20:04:59.865883 139841985250112 learning.py:506] global step 107060: loss = 0.9265 (0.461 sec/step)
INFO:tensorflow:global step 107070: loss = 0.7003 (0.446 sec/step)
I1015 20:05:04.328712 139841985250112 learning.py:506] global step 107070: loss = 0.7003 (0.446 sec/step)
INFO:tensorflow:global step 107080: loss = 0.9612 (0.434 sec/step)
I1015 20:05:08.808678 139841985250112 learning.py:506] global step 107080: loss = 0.9612 (0.434 sec/step)
INFO:tensorflow:global step 107090: loss = 1.7290 (0.444 sec/step)
I1015 20:05:13.288547 139841985250112 learning.py:506] global step 107090: loss = 1.7290 (0.444 sec/step)
-----------------------------
Check out that last line: with a batch size of 42 images (maximum I can fit on my GPU with 24Gb memory) I randomly get the occasional batch where total loss is more than double the moving average over the last 100 batches!
There's nothing fundamentally wrong with this, but it will throw a wrench in convergence of the model for a number of iterations, and it's probably not going to help in reaching the ideal ultimate state of the model within the number of iterations I have planned.
This is in part due to the fact that I have a fundamentally unbalanced dataset, and need to apply some pretty large label-wise weight rebalancing in the loss function to account for that... but this is the best representation of reality in the case I am working on!
--> The ideal solution RIGHT NOW would be to use a larger batch size, in order to minimise the possibility of getting these large outliers in the training set.
To get the best results in the short term I want to train this system (using this small backbone) with batches of 420 images instead of 42, which would require 240Gb of memory... so 3x Nvidia A100 GPUs for example!
--> Ultimately the next step is to make a version with a backbone that has 5x more parameters, and on the same dataset but scaled at 2x linear resolution... requiring probably around 500Gb of VRAM to have large enough batches to achieve good convergence on all classes!
And bear in mind this is a relatively small model, which can be deployed on a stamp-size Intel Movidius TPU and run in real-time directly inside a tiny sensor. For non-realtime inference there are people out there working on models with 1000x more parameters!
So if anyone is wondering why NVIDIA keeps making more and more powerful GPUs, and wondering who could possibly need that much power... this is your answer: the design / development of these GPUs is now being pulled forward 90% by people who need these kinds of solutions for ML ops, and the gaming market is 10% of the real-world "need" for this type of power, where 10 years ago that ratio would have been reversed.
- mikewarot 4y agoA few months ago, I thought... ok, this is amazing, an i7 with 8 cores and 32 GB of RAM and an SSD... I'm set for the next decade! Now I'm just trying to run Stable Diffusion to generate things an it's 30 seconds per iteration at only 512x512 pixels. It looks like I'll have a GPU soon enough if I want to do any development with neural networks. What an amazing ride... heck, I still remember thinking back in 1980 that the 10 Megabyte Corvus hard drive my friend was installing for a business would NEVER be filled... do you have any idea how much typing that is? ;-)
- jsjohnst 4y ago> 10 Megabyte Corvus hard drive my friend was installing for a business would NEVER be filled... do you have any idea how much typing that is? I’ve seen a single PDF which is the internal developer guide to a certain ARM based processor from years ago and that PDF was well over 1,000x in size to that hard drive. Yes, there were some pictures and charts, but most of that was just text, thousand upon thousand pages of detailed specs.
- jwalton 4y agoWhen I bought my 386 sx/16, they were all out of 10mb drives, so I bought a 20mb, and my dad thought I was crazy, because that was more space than anyone could ever need. :)
- izacus 4y agoWhat do you mean "games don't need them"?
- janoc 4y agoHe means that the average gamer that isn't gaming in 4k at 120Hz and/or going crazy with a high end VR rig (these things are still a tiny minority comparatively) will not fully utilize even the capabilities of the current hardware. Whereas the huge machine learning (and also cryptomining) markets can't get enough hardware resources. And those companies are buying these GPUs by truckloads, every month. Gamers are a very vocal but ultimately completely irrelevant market for these manufacturers. They account only for a tiny part of their income.