3 ms·
I disagree with his position on memory (he mentions in the post that anything above 1.5GB should be fine). In my experience, anything below 3GB can be pretty un
by benanne 12y ago
I disagree with his position on memory (he mentions in the post that anything above 1.5GB should be fine). In my experience, anything below 3GB can be pretty uncomfortable these days, if you want to work on serious problems.
It's not just about fitting the parameters into GPU memory, but also all the operations you perform on them, which can require a lot of intermediate storage. The example he gives only has fully-connected layers, but convolutional neural networks tend to require more space, especially some more recent implementations (e.g. FFT-based convolutions or the GEMM approach used by Caffe).
He mentions that his network (fully connected) has 52M parameters and compares it to Krizhevsky's 2012 ImageNet network (convolutional) network, which had 60M. But Krizhevsky actually explicitly mentions in his paper that memory was an issue:
"A single GTX 580 GPU has only 3GB of memory, which limits the maximum size of the networks that can be trained on it. It turns out that 1.2 million training examples are enough to train networks which are too big to fit on one GPU. Therefore we spread the net across two GPUs." (from http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks http://papers.nips.cc/paper/4824-imagenet-classification-wit... )
- deleted 12y ago[deleted]
- PostOnce 12y agoSerious is a spectrum, someone working on more "serious" problems than you might laugh at the idea of doing any work at all on a single consumer GPU.
- benanne 12y agoWell, I did say "in my experience" :) I certainly didn't mean to imply that problems requiring less than 3GB of GPU RAM are laughable, or anything like that. I should have said something like like "problems that people are currently writing papers about", maybe.