11 ms·
How does cosine similarity work?
- deleted 2y ago[deleted]
- OutOfHere 2y agoFor those who don't want to use a full-blown RAG database, scipy.spatial.distance has a convenient cosine distance function. And for those who don't even want to use SciPy, the formula in the linked post. For anyone new to the topic, note that the monotonic interpretation of cosine distance is opposite to that of cosine similarity.
- deleted 2y ago[deleted]
- ashvardanian 2y agoSciPy distances module has its own problems. It's pretty slow, and constantly overflows in mixed precision scenarios. It also raises the wrong type of errors when it overflows, and uses general purpose `math` package instead of `numpy` for square roots. So use it with caution. I've outlined some of the related issues here: https://github.com/ashvardanian/SimSIMD#cosine-similarity-reciprocal-square-root-and-newton-raphson-iteration https://github.com/ashvardanian/SimSIMD#cosine-similarity-re...
- OutOfHere 2y agoNoted, and thanks for your great work. My experience with it is limited to working with LLM embeddings, which I believe have been cleanly between 0 and 1. As such, I am yet to encounter these issues. Regarding the speed, yes, I wouldn't use it with big data. Up to a few thousand items has been fine for me, or perhaps a few hundred if pairwise.
- dbfclark 2y agoA good way to understand why cosine similarity is so common in NLP is to think in terms of a keyword search. A bag-of-words vector represents a document as a sparse vector of its word counts; counting the number of occurrences of some set of query words is the dot product of the query vector with the document vector; normalizing for length gives you cosine similarity. If you have word embedding vectors instead of discrete words, you can think of the same game, just now the “count” of a word with another word is the similarity of the word embeddings instead of a 0/1. Finally, LLMs give sentence embeddings as weighted sums of contextual word vectors, so it’s all just fuzzy word counting again.
- deleted 2y ago[deleted]
- heyitsguay 2y agoSorta related -- whenever I'm doing something with embeddings, i just normalize them to length one, at which point cosine similarity becomes a simple dot product. Is there ever a reason to not normalize embedding length? An application where that length matters?
- psyklic 2y agoFor the LLM itself, length matters. For example, the final logits are computed as the un-normalized dot product, making them a function of both direction and magnitude. This means that if you embed then immediately un-embed (using the same embeddings for both), a different token might be obtained. In models such as GPT2, the embedding vector magnitude is loosely correlated with token frequency.
- ashvardanian 2y agoOn the practical side, dot products are great, but break in mixed precision and integer representations, where accurately normalizing to unit length isn't feasible. In other cases people prefer L2 distances for embeddings, where the magnitude can have a serious impact on the distance between a pair of points.
- janalsncm 2y agoIf you’re feeling guilty about it you can usually store the un-normalized lengths separately.
- ashvardanian 2y agoThree separate passes over JavaScript arrays are quite costly, especially for high-dimensional vectors. I'd recommend using `TypedArray` with vanilla `for` loops. It will make things faster, and will allow using C extensions, if you want to benefit from modern hardware features, while still implementing the logic in JavaScript: https://ashvardanian.com/posts/javascript-ai-vector-search/ https://ashvardanian.com/posts/javascript-ai-vector-search/
- seydor 2y agodot product is easiest to understand as the projection of one vector to the other. the rest of it is self explanatory
- acjohnson55 2y agoI might have missed this, but I think the post might bury the lede that in a high dimensional space, two randomly chosen vectors are very unlikely to have high cosine similarity. Or maybe another way to put it is that the expected value of the cosine of two random vectors approaches zero as the dimensionality increases. Most similarity metrics will be very low if vectors don't even point in the same direction, so cosine similarity is a cheap way to filter out the vast majority of the data set. It's been a while since I've studied this stuff, so I might be off target.
- OutOfHere 2y agoEven if two random vectors don't have high cosine similarity, and I have not had this issue in 3000 dimensions, the cosine similarity is still usable in relative terms, i.e. relative to other items in the dataset. This keeps it useful.
- acjohnson55 2y agoMakes sense. I'm guessing it's one of those things where there's significant info in the magnitude of the exponent, in terms of relative similarity?
- DavidSJ 2y agoNitpick: The expected value of the cosine is 0 even in low-dimensional spaces. It’s the expected square of that (i.e. the variance) which gets smaller with the dimension.
- acjohnson55 2y agoThat totally makes a sense, thanks!
- naijaboiler 2y agoImagine 2 points in 3 dimensional space with a vector being the line from the origin to the point. So you have 2 vectors pointing going to the 2 points from the origin. If those points are really close together, then angle between the two vector lines is very small. Loosely speaking cosine is a way to quantize how close two lines with a shared origin is. If both lines are the same, the angle between them is 0, and the cosine of 0 is 1. If two lines are 90 degrees apart, their cosine is 0. If two lines are 180 degrees apart, their cosine is -1. So cosine is a way to quantify the closeness of two lines which share to same origin To go back with 2 points in space that we started with, we can measure how close those 2 points are by taking the cosine of the lines going from origin to the two points. If they are close, the angle between them is small. If they are the exact same point, the angle between the lines is 0. That line is called a vector Cosine similarity measures how closes two vectors are in Euclidean space. That’s we end up using it a lot. It’s no the only way to measure closeness. There are many others
- stouset 2y agoAre all the points in question one unit distant from the origin?
- deleted 2y ago[deleted]
- viciousvoxel 2y agoyes, cosine similarity involves normalizing the points (by the L2 norm) and then dot product. In other words the points lie on the unit (hyper)sphere.
- bee_rider 2y agoI vaguely remember some paper where they didn’t even bother normalizing the vectors, because they expected zeros to be very close to zero, and anything else was considered a one. I have no idea if this a common optimization or if it was something very niche. It was for a heuristic matrix reordering strategy, so I think they were willing to accept some mistakes.
- 2y ago
- derbOac 2y agoMaybe I missed this but I was surprised they didn't mention the connection to correlation. Cosine similarity can be thought of as a correlation, and some bivariate distributions (normal I think?) can be rexpressed in terms of cosine similarity. There's also some generalizations to higher dimensional notions of cosines that are kind of interesting.
- lern_too_spel 2y agoThis doesn't properly explain what it says it explains. To explain it correctly, you have to explain why the dot product of two vectors computed as the sum of the products of the coefficients of an orthonormal basis is a scalar equal to the product of the Euclidean magnitudes of the vectors and the cosine of the angle between them. The Wikipedia article on dot product explains this reasonably well, so just read that.
- esafak 2y agoIt works like the dot product in the case of spherical embeddings. Normalizing the embeddings makes it easier to understand too.
- cproctor 2y agoOne thing I've wondered for a while: Is there a principled reason (e.g. explainable in terms of embedding training) why a vector's magnitude can be ignored within a pretrained embedding, such that cosine similarity is a good measure of semantic distance? Or is it just a computationally-inexpensive trick that works well in practice? For example, if I have a set of words and I want to consider their relative location on an axis between two anchor words (e.g. "good" and "evil"), it makes sense to me to project all the words onto the vector from "good" to "evil." Would comparing each word's "good" and "evil" cosine similarity be equivalent, or even preferable? (I know there are questions about the interpretability of this kind of geometry.)
- marginalia_nu 2y agoDunno if I have the full answer, but it seems in high dimensional spaces, you can typically throw away a lot of information and still preserve distance. The J-L lemma is at least somewhat related, even though it doesn't to my understanding quite describe the same transformation. https://en.m.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_lemma https://en.m.wikipedia.org/wiki/Johnson%E2%80%93Lindenstraus... see also https://en.m.wikipedia.org/wiki/Random_projection https://en.m.wikipedia.org/wiki/Random_projection
- nostrademons 2y agoSo I first learned about cosine similarity in the context of traditional information retrieval, and the simplified models used in that field before the development of LLMs, TensorFlow, and large-scale machine learning might prove instructive. Imagine you have a simple bag-of-words model of a document, where you just count the number of occurrences of each word in the document. Numerically, this is represented as a vector where each dimension is one token (so, you might have one number for the word "number", another for "cosine", another for "the", and so on), and the magnitude of that component is the count of the number of times it occurs. Intuitively, cosine similarity is a measure of how frequently the same word appears in both documents. Words that appear in both documents get multiplied together, but words that are only in one get multiplied by zero and drop out of the cosine sum. So because "cosine", "number", and "vector" appear frequently in my post, it will appear similar to other documents about math. Because "words" and "documents" appear frequently, it will appear similar to other documents about metalanguage or information retrieval. And intuitively, the reason the magnitude doesn't matter is that those counts will be much higher in longer documents, but the length of the document doesn't say much about what the document is about. The reason you take the cosine (which has a denominator of magnitude-squared) is a form of length normalization, so that you can get sensible results without biasing toward shorter or longer documents. Most machine-learned embeddings are similar. The components of the vector are features that your ML model has determined are important. If the product of the same dimension of two items is large, it indicates that they are similar in that dimension. If it's zero, it indicates that that feature is not particularly representative of the item. Embeddings are often normalized, and for normalized vectors the fact that magnitude drops out doesn't really matter. But it doesn't hurt either: the magnitude will be one, so magnitude^2 is also 1 and you just take the pair-wise product of the vectors.
- dheera 2y agoIt's just a normalized dot product. People use "cosine similarity" to sound knowledgeable
- ttoinou 2y agoAgreed. It’s almost like Wildberger‘s rational trigonometry
- youssefabdelm 2y agoHm for me interactive visualizations are more illuminating: https://www.falstad.com/dotproduct/ https://www.falstad.com/dotproduct/ https://wordsandbuttons.online/interactive_mnemonics_for_dot_and_cross_vector_products.html https://wordsandbuttons.online/interactive_mnemonics_for_dot... In essence then, not as confusing to the beginner who might even know what a dot product 'is' operationally but not what it 'does'. So level 1 'it's just a normalized dot product', level 2 more immediately intuitive: 'is arrow 1 pointing in the same direction as arrow 2?' or 'how close is arrow 1's direction to arrow 2's direction?' Now what's left after that is 'Why is it so? Why did we decide on this in embeddings?'
- Izkata 2y ago"Cosine similarity" includes that particular kind of normalization, it does actually impart more information than "normalized dot product".
- janalsncm 2y agoCosine similarity for unit vectors looks like the angle between hands on a high dimensional clock. 12 and 6 have -1 cosine sim. 5 and 6 are pretty close. Cosine similarity works if the model has been deliberately trained with cosine similarity as the distance metric. If they were trained with Euclidean distance the results aren’t reliable. Example: (0,1) and (0,2) have a cosine similarity of 1 but nonzero Euclidean distance.
- anArbitraryOne 2y agoNot enthused about X notating the dot product as opposed to cross product
- p_j_w 2y agoMaybe I missed it but I don’t see the author doing that in the article. They use a dot to denote dot product. They use X for multiplying the magnitude of two vectors, which isn’t my favorite thing but isn’t offensive.
- GuB-42 2y agoI think the use of the term "cosine" here is needlessly confusing. It is the dot product of normalized vectors. Sure, when you do the maths, it gives out a cosine, but since we are not doing geometry here, so it isn't really helpful for a beginner to know that. Especially considering that these vectors have many dimensions and anything above 3D is super confusing when you think about it geometrically. Instead just try to think about what it is: the sum of term-by-term products of normalized vectors. A product is the soft version of a logic AND, and it makes intuitive sense that vectors A and B are similar if there are a lot of traits that are present in both A AND B (represented by the sum) relative to the total number of traits that A and B have (that's the normalization process). Forget about angles and geometry unless you are comfortable with N-dimensional space with N>>3. Most people aren't.
- adw 2y ago> we are not doing geometry here we absolutely are doing geometry here, given we're talking about metrics in a vector space – and this is trigonometry you learned by the first year of high school.
- jameshart 2y agoYou might like to think of vectors in their geometric interpretation but vectors are not inherently geometric - vectors are just lists of numbers, which we sometimes interpret geometrically because it helps us comprehend them. High dimensional vectors grow increasingly ungeometric as we have to wrestle with increasingly implausible numbers of orthogonal spatial dimensions in order to render them ‘geometric’. In the end, vectors (long lists of numbers a1, a2, a3, … an) start looking more like discrete functions f(i) = ai. And you can extend the same concept all the way to continuous functions - they’re like infinite dimensional vectors. For continuous functions over a finite interval the dot product (usually called the inner product in this domain) is just the integral of the product of two functions, and the ‘magnitude’ of a function is its RMS, and that means functions have a ‘cosine similarity’ which is not remotely geometric. There isn’t any geometric sense in which there is an ‘angle between’ cos(x) and sin(x) except it turns out that they have a cosine similarity of 0 so it implies the ‘angle between’ them is 90°, which actually makes a lot of sense. But in this same sense there’s an ‘angle between’ any two functions (over an interval). But we are not doing geometry here.
- anArbitraryOne 2y agoCosine similarly is the epitome of status quo bias. How many DS or ML people actually think through similarity metrics that might be appropriate, then choose cosine? Gods forbid they have to justify using a different measure to their colleagues
- whiterknight 2y agoSeems more like easy to understand and gets the job done bias. Are there common pitfalls you have identified?
- anArbitraryOne 2y agoMagnitude Ignorance: Cosine similarity considers only the angle between vectors, not their magnitude. This is problematic when the magnitude carries important information. For example, if vectors represent term frequencies in documents, cosine similarity treats two documents with vastly different lengths but the same proportion of words as identical. Sensitive to High-dimensional Sparsity: In high-dimensional spaces (e.g., text data), vectors are often sparse (many zeros). Cosine similarity might not provide meaningful results if most dimensions are zero since the similarity could be dominated by a few non-zero entries. No Sense of Absolute Position: Cosine similarity measures the angle between vectors but ignores their absolute position. For example, if vectors represent geographical coordinates, cosine similarity won't capture differences in distances properly. Poor Performance with Highly Noisy Data: If the data has significant noise, cosine similarity can be unreliable. The angle between noisy vectors might not reflect true similarity, especially in high-dimensional spaces. Does Not Handle Negative Values Well: If vectors contain negative values (e.g., sentiment scores, certain word embeddings), cosine similarity may yield unintuitive results since negative values can affect the angle differently compared to positive-only data. Assumes Non-Negative Values: Often, cosine similarity assumes non-negative values. In contexts where vectors have both positive and negative values (e.g., sentiment analysis with positive and negative sentiment words), this assumption can lead to misleading results. Not Ideal for Measuring Dissimilarity: Cosine similarity can be unintuitive when measuring dissimilarity. Two vectors that are orthogonal (90 degrees apart) will have a similarity score of 0, but vectors pointing in opposite directions (-1 cosine similarity) might need a different interpretation depending on the context. Inappropriate Use Cases Data with Magnitude Importance: When the magnitude of vectors is crucial (e.g., comparing sales data, where larger magnitudes indicate higher sales), using cosine similarity would ignore valuable information. Time Series Analysis: For time-series data, the order and distance of data points matter. Cosine similarity does not account for these aspects and may not provide meaningful comparisons for temporal data. Geospatial Data: When working with geospatial coordinates (latitude, longitude), cosine similarity does not account for Earth’s curvature or distance metrics like the Haversine formula. Data Representing Complex Structures: For data representing graphs, trees, or other complex structures where connectivity or sequence matters, cosine similarity may not capture the intricate relationships between nodes or elements. Vectors with Negative Components: In cases where vectors have meaningful negative components (like certain word embeddings or feature vectors in machine learning models), cosine similarity can yield misleading similarity scores. Suggestions for Alternatives Euclidean Distance: When absolute magnitude is important, or when interpreting actual distances between points. Jaccard Similarity: For binary or set-based data, where overlap or presence/absence matters. Pearson Correlation: For datasets where linear relationships are of interest, especially with normally distributed values. Hamming Distance: For comparing binary data, especially for bit strings or categorical attributes. Manhattan Distance (L1 Norm): For high-dimensional data where you want to measure the absolute difference across dimensions. Cosine similarity is effective for certain applications, such as text similarity, but its limitations make it unsuitable for other contexts where magnitude, distance, or data distribution play a critical role.
- throwawaymaths 2y agoI was always curious as to why we don't encode ML parameters and activations using complex numbers. A dot product between two complex numbers naturally encodes confidence in the result in the magnitude. I blame numpy
- SomewhatLikely 2y agoSomething worth mentioning is that if your vectors all have the same length then cosine similarity and Euclidean distance will order most (all?) neighbors in the same order. Think of your query vector as a point on a unit sphere. The Euclidean distance to a neighbor will be a chord from the query point to the neighbor. Just as with the angle between the query-to-origin and the neighbor-to-origin vectors, the farther you move the neighbor from the query point on the surface of the sphere, the longer the chord between those points gets too. EDIT: Here's a better treatment, and it is the case that they give the exact same orderings: https://ajayp.app/posts/2020/05/relationship-between-cosine-similarity-and-euclidean-distance/ https://ajayp.app/posts/2020/05/relationship-between-cosine-...
- dontreact 2y agoCosine similarity is equal to the dot product of each vector normalized
- niemandhier 2y agoIn high dimensional Spaces the distances between nearest and farthest points from query points with respect to normal metrics become almost equal. Cosine similarity still works though, since it only look at how aligned vectors are. The thing that people tend to overlook is, that there is no need for embeddings to be a vector space endowed with an inner product. Words don’t have this structure, we define it on the image of the mapping from words to n-tuples and the embeddings we use coevolved in such a way that we assume the cosine similarity to be meaningful.
- deleted 2y ago[deleted]
- jheriko 2y agoso normalised dot product. the notation here is bad. the bottom of the division looks like a cross product as a games and graphics programmer i find it amazing that this would be a mystery... understanding the dot product is utterly foundational, and is some high-school level basics.
- deleted 2y ago[deleted]
- calderwoodra 2y agoI think it would have been helpful to mention the Pythagorean theroem, as most people are familiar with it, but otherwise the post did a great job explaining and introducing the topic.
- cubacaban 2y agoRather superficial and obfuscating. The article keeps raising the question "why ignore the magnitude" and never answers it. "The important part of an embedding is its direction, not its length. If two embeddings are pointing in the same direction, then according to the model they represent the same "meaning"." This can't be quite right. Any LLM transformer model looks at the embedding of the token sequence, (without normalizing, i.e. including its magnitude) for deciding on the next token. Why would you throw away that information, equivalent to throwing away one embedding dimension? If I had to guess why cosine similarity is the standard for comparing embeddings I suspect it's simply because the score is bounded in [-1, 1], which you may find more interpretable than the unbounded score obtained by the unnormalized dot product or Euclidean distance. In my experience, choice of similarity metric doesn't affect embedding performance much, simply use the one the embedding model was trained with.