3 ms·
Theoretically, any discriminator that maximizes the expected value rather than the underlying distribution will be better. Logistic regression is an example of
by tensor 3y ago
Theoretically, any discriminator that maximizes the expected value rather than the underlying distribution will be better. Logistic regression is an example of a discriminative classifier.
Even better, however, is if you maximize the margin, so a large margin discriminative classifier should be your baseline. I believe that the classification layer in fasttext is equivalent to logistic regression (but not large-margin).
All that said, I would probably use vowpalwabbit as a baseline as it doesn't use word vectors underneath, is extremely fast and easy to use, and has many optimization option. This way you can determine if word vectors help your particular problem or not.
- sweezyjeezy 3y agoAll bag-of-words models use some kind of word vector - for a normal logistic regression / naive bayes, say "spam" and "spammer" are first and second indexed words in the vectorizer, then their word vectors are like [1, 0, 0, ....] and [0, 1, 0, ....] (length = vocab size). For both logistic regression and fasttext you get the 'document vector' by adding up their respective word vectors before applying a final linear projection + softmax to get the class predictions [1] The insight of fasttext is to notice that these high dimension, unit vector word embeddings aren't an ideal way to learn - since they are orthogonal, during training time, a datapoint giving signal between the "spam" word vector and the label gives you no information about "spammer", or indeed any other word. Explicitly modelling and learning word vectors helps with this as now these are two points in a vector space that can be moved relative to one another. [1] this remains true for LR if you rescale the vectors using e.g. TFIDF or use the hashing trick (a la vowpalwabbit).
- tensor 3y agoYes when I said word vectors I meant rich embeddings not one hot representations. That said, in reviewing the fastext bag of tricks paper on their classification module I’m now second guessing my assumption that they use complex embeddings. Their architecture is otherwise exactly reproducible in vowpalwabbit, and in fact in the paper they claim that it is equivalent to a specific combination of vowpalwabbit flags. In particular the vowpalwabbit neural network flag is needed. However, vowpalwabbit only uses one hot vectors for their ngram features. The neural network flag just adds a hidden layer. I had assumed that fasttext uses rich word embeddings in its classifier because it has another module to train them. If it is actually the same as vowpalwabbit, then I can say that I’ve never had the extra hidden layer really help, though as they note it does make vowpalwabbit quite slow.
- sweezyjeezy 3y agofasttext word embedding is equivalent to adding a hidden layer, as long as you DON'T put a nonlinearity on it if that helps.