3 ms·
While old, the nice thing about word2vec models is how easily they can be imported and used by any language to create a vector search. I have always wondered (
by boyter 2y ago
While old, the nice thing about word2vec models is how easily they can be imported and used by any language to create a vector search.
I have always wondered (but never done anything) about if you could train on code to achieve a similar result. I suspect there could be some value there even if its just to help identify similar snippets of code.
- maujim 2y agothere is a code2vec as well
- PaulHoule 2y agoYou can do the same with https://sbert.net/ https://sbert.net/ it is not any more work for you to implement except it really gives better results! Vector databases did not become hot in the word2vec era, they became hot once you got embedding that were sensitive to words in context. It's arguable, for instance, there is any value in making an embedding for a word which can have multiple meaning. Take a word like "bat", at least it is specific to match that with the specific word bat. If you vectorize it you're going to have to blend in mammals and blend in sports equipment. In any given situation you care about one of them and don't care about the other, so anything you gain from matching other mammals means you also get spurious matches having to do with sports equipment. BERT sees the context so it will match a particular use of the word "bat" with either mammals or sports equipment so it brings in relevant synonyms but not irrelevant synonyms. That made BERT one of those once-in-a-decade breakthroughs in information retrieval (took a whole decade of conference proceedings to get BM25!) whereas Word2Vec is just a dead end that people wrote too many blog posts about.