4 ms·
Looks neat. Why did you bother using <PAD> words to have sentences be the same length, when you're using a bag-of-words (document-term matrix) model anyway? Ea
by primaryobjects 11y ago
Looks neat. Why did you bother using <PAD> words to have sentences be the same length, when you're using a bag-of-words (document-term matrix) model anyway?
Each sentence vector ends up being the length of the vocabulary, so they're already the same length. You can probably drop step #3 in this case.
- dennybritz 11y agoHi! It is not using a BoW model. Each input sentence is a vector of size [sentence_length] (or, in theory, a matrix of size [vocab_size, sentence_length] with one-hot vectors) so the padding is required. There is a way to do it without padding, but it's less efficient from a training point of view. You could instantiate a new network for each possible sentence length then share the paramaters between them, and then batch based on your sentence length. Also, the padding isn't striclty necessary in theory. The feature vector will always end up being the same length, regardless of sentence length, due to the pooling layer. However, Tensorflow forces you to specify the exact size of the pooling operation (you can't just say "pool over the full input"), so you need it if you're using TF.