3 ms·
Bag of Words is not actually a great approach to understand text because it ignores the semantics of the word. For example, 'hotel' and 'motel' which are simila
by ankeshanand 10y ago
Bag of Words is not actually a great approach to understand text because it ignores the semantics of the word. For example, 'hotel' and 'motel' which are similar words have completely different vector representations in the BoW model.
A popular alternative is to use a distributed word embedding such as word2vec[1], where similar words are grouped together in the vectorspace.
Edit: If there are few observations, like in this case, we don't need to train the word2vec model on the dataset itself. We can use pre-trained word embeddings such as the one publicly released by Google which was trained on the Google News dataset.
[1]https://word2vec.googlecode.com/ https://word2vec.googlecode.com/
- autokad 10y agogrouping dont always yield better results, and I think it would probably do pretty poorly in this case because there are few observations. Random Forest won in the samples tried, but I wager a support vector machine with a histogram kernel would do fantastic.
- ankeshanand 10y agoWe could always use pre-trained word embeddings, the few observations won't matter then.
- autokad 10y agothis is very industry specific, finding them that are meaningful for the data at hand seems like a small chance. edit: im not saying dont try it, i certainly would! lets look at the data on github? maybe we could have a wack at the data