3 ms·
For big LSTMs and long-ish sequences, the intermediate gradients can take up a huge amount of memory - often more than the model parameters themselves. In my ex
by kastnerkyle 11y ago
For big LSTMs and long-ish sequences, the intermediate gradients can take up a huge amount of memory - often more than the model parameters themselves. In my experience it is mostly big LSTMs that need the 12GB+ GPUs. You can reduce the batch size to help this a bit, or train using trucated BPTT but RNN training is already a slow, sequential business.
Of course, there are no clear wins (generally, losses) in computation by scaling RNNs horizontally on GPUs - but sometimes you really do need more than 12GB.
- dave_sullivan 11y agoAll true things. But at the heart of the issue: No clear wins in horizontal scaling Reduce batch size Use truncated back prop *search for better hyper params* I usually do 2-4. After those, have you really seen scaling result in significant accuracy gains? And what percent of the time is that necessary? Genuinely interested--and you guys rock btw!
- kastnerkyle 11y agoOne case where I see 1 as being necessary is the softmax size / vocabulary boost in the seq2seq paper (8 GPUs IIRc, and 4 were dedicated to a softmax!). There are other ways to handle this (hierarchical softmax, sampled softmax, skip-thoughts trick of using word2vec vocabulary) but every time someone figures out how to have a larger vocabulary, neural machine translation results seem to improve. In general, 1 is a last resort for me - but maybe this is due to current tooling and availability of hardware as much as anything?