4 ms·
I have seen this one before somewhere, and what amazes me is that how you can solve problems without a hassle if you get the "trick" right. Another case I read
by clvv 16y ago
I have seen this one before somewhere, and what amazes me is that how you can solve problems without a hassle if you get the "trick" right. Another case I read was that Google use(at least used) two vectors(each consists of many 0s and 1s, which in turn represent whether the web page has the keyword or not) to represent web pages, and calculate the angle between the vectors to figure out the similarity(a value) between web pages.
- l0nwlf 16y agoYou are probably talking about Cosine similarity ( http://en.wikipedia.org/wiki/Cosine_similarity http://en.wikipedia.org/wiki/Cosine_similarity ) , a commonly used technique in the field of IR/NLP.
- maurits 16y agoThat would be the cosine metric then. I don't think it is that simple. The position (or dimension if you will) of a keyword in the data-vector is all important. So equal vectors in terms of the symbolic presence of a number of keywords will generate great cosine distances the moment the keywords are not in equal position.
- bad_user 16y agoThat wouldn't work well, because you need to take into account the relevance of a keyword for those web pages. Check out the Tf-Idf weight for getting an idea on how to do that: http://en.wikipedia.org/wiki/Tf%E2%80%93idf http://en.wikipedia.org/wiki/Tf%E2%80%93idf A useful trick based on the technique you describe: build vectors where each dimension represents a user in the system (1 means the user visited that web page); since people are interested in a narrow array of topics. Furthermore you can extend this to find similar users. Amazon does this to find similar or complementary products.
- tyler 16y agoIt sounds like you're conflating two techniques here. The first (as others have mentioned) is cosine similarity, which measures the angle between the vectors. However, the bit about 0s and 1s sounds like you're talking about locality-sensitive hashing (http://en.wikipedia.org/wiki/Locality_sensitive_hashing http://en.wikipedia.org/wiki/Locality_sensitive_hashing). LSH is often used to estimate cosine similarity, as cosine similarity can be quite expensive to calculate. I know Google and others are using it for such.