4 ms·
Regarding finding similar documents what is the state of the art nowadays, LDA, word2vec, something else? What do you normally use?
by pencilcode 9y ago
Regarding finding similar documents what is the state of the art nowadays, LDA, word2vec, something else? What do you normally use?
- nicklovescode 9y agoHave you heard of word mover’s distance? It works really well!
- zintinio5 9y agoLike everything else, depends on your use-case. I have personally used TF-IDF vectors and token sets with Cosine and Jaccard distances in practice. Some examples of use-cases: are you searching for "semantically similar", or "near duplicate"? You can compare documents under different metrics and different _representations_. Some representations are: LSA, PLSA, LDA, TF-IDF, and Set representations, along with metrics such as Jaccard Distance, Cosine Distance, Euclidean distance, etc. Doc2vec is the Word2vec analog for documents.
- nl 9y agoWord Mover Distance on Word2Vec vectors. There is an implementation in Textacy.