3 ms·
The problem with IVF is that you need to find the right centroids. And that doesn't work well if your data grow and mutate over time. Splitting a centroid is a
by ayende 2y ago
The problem with IVF is that you need to find the right centroids.
And that doesn't work well if your data grow and mutate over time.
Splitting a centroid is a pretty complex issue.
As are clustering in an area. For example, let's assume that you hold StackOverflow questions & answers. Now you have a massive amount of additional data (> 25% of the existing dataset) that talks about Rust.
You either need to re-calculate the centroids globally, or find a good way to split.
The posting list are easy to use, but if you are unbalanced, it gets really bad.
- VoVAllen 2y agoHi, I'm the author of the article. Meta have conducted some experiments on dynamic IVF with datasets of several hundred million records. The conclusion was that recall can be maintained through simple repartitioning and rebalancing strategies. You can find more details here: DEDRIFT: Robust Similarity Search under Content Drift https://arxiv.org/pdf/2308.02752 https://arxiv.org/pdf/2308.02752. Additionally, with the help of GPUs, KMeans can be computed quickly, making the cost of rebuilding the entire index acceptable in many cases.