3 ms·
Yes when I said word vectors I meant rich embeddings not one hot representations. That said, in reviewing the fastext bag of tricks paper on their classificati
by tensor 3y ago
Yes when I said word vectors I meant rich embeddings not one hot representations.
That said, in reviewing the fastext bag of tricks paper on their classification module I’m now second guessing my assumption that they use complex embeddings. Their architecture is otherwise exactly reproducible in vowpalwabbit, and in fact in the paper they claim that it is equivalent to a specific combination of vowpalwabbit flags.
In particular the vowpalwabbit neural network flag is needed. However, vowpalwabbit only uses one hot vectors for their ngram features. The neural network flag just adds a hidden layer.
I had assumed that fasttext uses rich word embeddings in its classifier because it has another module to train them.
If it is actually the same as vowpalwabbit, then I can say that I’ve never had the extra hidden layer really help, though as they note it does make vowpalwabbit quite slow.
- sweezyjeezy 3y agofasttext word embedding is equivalent to adding a hidden layer, as long as you DON'T put a nonlinearity on it if that helps.