4 ms·
Not reading the paper, cosine similarity has little to no semantic understanding of sentences. E.g. the following triple 1: "Yes, this is a demonstration" 2:
by latency-guy2 3y ago
Not reading the paper, cosine similarity has little to no semantic understanding of sentences.
E.g. the following triple
1: "Yes, this is a demonstration"
2: "Yes, this isn't a demonstration"
3: "Here is an example"
<1, 2>, Has "higher" cosine similarity than <1, 3>, structurally equivalent except for one token/word, <1, 2> semantically means the opposite of each other depending on what you're targeting in that sentence. While <1, 3> means effectively the same thing.
If this paper is about persuading people about efficacy with regards to semantic understanding, OK, but that was always known. If its about something with relation to vectors and the underlying operations, then I'll be interested.
- soarerz 3y agoWhat is the cheapest way to capture similarity if not via dot product then?
- deleted 3y ago[deleted]
- gajus 3y agoInterested to know as well
- latency-guy2 3y agoI don't have an answer for this really outside of silly ones like "strict equality check", but I assert that no one else does either, at least today and right now, and its an inherent limitation due to the nature of embeddings and the space it desires to be (cheap, fast, good enough similarity for your use case). You're probably best off using the commercial suggestion, and if its dot product, go for it. I am no expert in this area and my interest wanes every day.
- scotty79 3y agoInstead of sums of multiplications you could for example use sum of squares of differences. Means squared error instead of dot product, it's not cheaper but it's close If you want to go cheaper you could use sum of abs of differences.
- soarerz 3y agoThis is effectively "the same" as dot product. For a lot of embeddings we have today, norm of any embedding vector is roughly of same size, so the angle between two vectors is roughly same size as length of difference that you are saying, and can be expressed in terms of 1 - dot product after scaling
- Tostino 3y agoThat is entirely dependant on the model for the embeddings. You can fine tune for pretty much any outcome you want.
- _t89y 3y agoYou can't fine-tune for understanding or reasoning. You can't "get better performance" on understanding. You're either equipped for it or you're not.
- superkuh 3y agoThat might be true for one-hot vectors but it's not true for learned embedding through the lens of attention. That said, I only made to page 3/9 of the paper before the mark-up for the math went over my head.
- _t89y 3y agoIt is true. And if you want to say anything about meaning this isn't even the right math.
- latency-guy2 3y agoIf we're talking about adding dimensionality and relying on the kernel, sure, but I only get so much incorrect by going from 1 to M dims. I can't of course be certain in all cases, but dimensions are typically (past experience, and using knowledge from word2vec experiments from years ago) derivative of higher dimensions. The kernel still operates on the same concept by applying a norm along with whatever weightings to each dim. Semantic understanding is still not there in my opinion, we might feign it by increasing specificity, but only so much. Largest contributor will likely still the determining factor rather than the series of smaller, more specific dimensions. I tested this using similar sentences as my original comment and failing in more scenarios than passing. I of course am biased since it may be given I did not select the right dimensions or measures.
- vidarh 3y agoWhether your not the cosine similarity of either pair is higher depends on the mapping you create from the strings to the embedding vector. That mapping can be whichever function you choose, and your result will be entirely dependent on that. If you choose a straight linear mapping of tokens to a number, then you'd be right. Extending that, if you choose any mapping which does not do a more extensive remapping from raw syntactic structure to some sort of semantic representation, you'd be right. But hence why we increasingly use models to create embeddings instead of simpler approaches before applying a similarity metric, whether cosine similarity or other. Put another way, there is no inherent reason why you couldn't have a model where the embeddings for 1 and 3 are identical even, and so it is meaningless to talk about the cosine similarity of your sentences without setting out your assumptions about how you will created embeddings from them.
- _t89y 3y agoIt is meaningless to talk about cosine similarity of sentences, or words, at all. Choose whatever mapping you want. You'll still be in Firth Mode.
- _t89y 3y agoUh oh. LOL. Got some angry Firthers out there.
- vidarh 3y agoIt's meaningful to talk about cosine similarity for anything that you can quantify in ways such that the cosine similarity reflects a measure you care about. Same applies for any function. If it works, it's meaningful to talk about it whether or not it has a reasonable interpretation beyond that.
- latency-guy2 3y ago> meaningless to talk about the cosine similarity of your sentences without setting out your assumptions about how you will created embeddings from them. I agree, but from generics POV, you have to settle on a few things to compare between models. If you can't, then benchmarks are useless too outside of extremely narrow measures. I only address structure in the parent, and sure, it can be too generic of a statement by only touching on structure. But I would almost assert structure is still an important feature, and I would almost assert that it is required or otherwise a dominant feature when you want to deliver a product for general use. I don't think I get too much more incorrect going beyond a few dimensions given this.
- deleted 3y ago[deleted]
- _t89y 3y agoNo understanding. Embeddings are a semantically vacuous representation and similarity is a semantically vacuous interpretation.