3 ms·
One case where I see 1 as being necessary is the softmax size / vocabulary boost in the seq2seq paper (8 GPUs IIRc, and 4 were dedicated to a softmax!). There a
by kastnerkyle 11y ago
One case where I see 1 as being necessary is the softmax size / vocabulary boost in the seq2seq paper (8 GPUs IIRc, and 4 were dedicated to a softmax!). There are other ways to handle this (hierarchical softmax, sampled softmax, skip-thoughts trick of using word2vec vocabulary) but every time someone figures out how to have a larger vocabulary, neural machine translation results seem to improve. In general, 1 is a last resort for me - but maybe this is due to current tooling and availability of hardware as much as anything?