2 ms·
Similarity Search for Multiple Embeddings?
let's say I'm building an app that matches users based on music preferences (ie. their liked songs). I can represent each track (audio file) a user likes by an embedding, but I'm not sure how to use these embeddings to match users. averaging all the tracks' embeddings to create a single embedding for the user, then using it in a KNN search, doesn't seem like a good approach to me, especially in a high dimensional space. especially since I feel like the average of two different tracks may create some meaningless vector.
I'm an ML noob. what are some smarter ways to match users, assuming I have accurate track-level embeddings?
- minimaxir 2y ago> averaging all the tracks' embeddings to create a single embedding for the user, then using it in a KNN search, doesn't seem like a good approach to me, especially in a high dimensional space If your embeddings are good, this is a valid first step. Yes, embeddings are weird like that. You may not want to apply equal weight to all tracks, though.
- jonasbrouthers 2y agothanks, I'll try that. your comment prompted me to look further why averaging is effective, which led me to this article, which made the intuition more clear to me - "The mean is a reasonable way to summarize embeddings thanks to the blessing of dimensionality. Because an exponential number of embeddings are nearly orthogonal in high dimensions, it is unlikely for two independent collections to have similar averages. On the other hand, two related documents are guaranteed to have similar averages. As a result, two collections have similar averages if and only if they share many similar embeddings with very high probability." https://randorithms.com/2020/11/17/Adding-Embeddings.html#:~:text=Spoiler%20Warning%3A%20The%20average%20is,two%20random%20high%2Ddimensional%20vectors https://randorithms.com/2020/11/17/Adding-Embeddings.html#:~....