3 ms·
This series of articles from Pinecone seems accessible https://www.pinecone.io/learn/vector-indexes/ https://www.pinecone.io/learn/vector-indexes/ . Stepping ba
by WinLychee 3y ago
This series of articles from Pinecone seems accessible https://www.pinecone.io/learn/vector-indexes/ https://www.pinecone.io/learn/vector-indexes/ . Stepping back a moment, the problem "I want to find the top-k most similar items to a given item from a collection" comes up in a variety of domains, and a common approach is nearest-neighbor search. Depending on the number of dimensions and the number of items, you may need tricks to speed up the search. In recent years we've seen an increase in scale and invention of some new tricks to speed this up.
In a bit more detail the idea is to convert your data into vectors embedded in some d-dimensional space and compute the distance between a query vector and all other vectors in the space. This is an O(dN) computation. ANN covers a number of techniques giving a speedup by letting you reduce number of required comparisons.