4 ms·
The whole embedding thing which converts “tokens” to vectors, which you then store in a vector database so that you can later query by vector distance, seems to
by brabel 1mo ago
The whole embedding thing which converts “tokens” to vectors, which you then store in a vector database so that you can later query by vector distance, seems to be LLM specific technology, no? As far as I know the vectors look a lot like the weights in a LLM itself which is why the vector search also works with some level of intelligence.
- triangle 1mo agoVector embeddings predate LLMs. They have been used as far back as the early 2000s. They are a general machine learning technique, rather than LLM specific
- ozim 1mo agoUnfortunately LLMs made vector search more popular so it seems like something LLM specific. What makes it worse, a lot of people in the thread equate vector search with RAG, whereas RAG is the name for anything that model can query so a user doesn't have to copy/paste feed it to the model manually like access to text files is RAG.
- nilirl 1mo agoSure and that's a new technique for indexing and querying. Where's the new design tension? Indexes always had to be monitored for freshness and queries have always needed cleaning or parsing.
- ewidar 1mo agonot really, vectorising text/books is old school ML by this point. at least to me that seems the same as https://en.wikipedia.org/wiki/Word2vec https://en.wikipedia.org/wiki/Word2vec for e.g.
- Foobar8568 1mo agoWell... Everything new is old "A vector space model for automatic indexing" 1975 - https://dl.acm.org/doi/10.1145/361219.361220 https://dl.acm.org/doi/10.1145/361219.361220
- esafak 1mo agoI wonder who was doing doing semantic search in the last century! "The future is already here—It's just not very evenly distributed..."
- vintermann 1mo agoSure, the idea of making a vector embedding for words, sentences, documents etc. is old, but the meat is in how you construct this embedding. I think embeddings have gotten quite a bit better since word2vec.
- KaseyKim 1mo agoright, it is the foundation of machine learning.