6 ms·
I investigated different algorithms for recommendation systems once and was astonished to find that the best one was also the most simple one: Jaccard similarit
by npstr 9y ago
I investigated different algorithms for recommendation systems once and was astonished to find that the best one was also the most simple one: Jaccard similarity. No other similarities nor a proprietary custom built algorithm (which was the actual target of the investigation) could beat it.
- DonaldFisk 9y agoThis implies that you know all the relevant attributes of the items you're recommending. For something like music, books, or movies, there might literally be thousands of them, most of which are unidentifiable, any of which might differ in importance from one item to another. I built a collaborative filtering system (http://web.onetel.com/~hibou/morse/MORSE.html http://web.onetel.com/~hibou/morse/MORSE.html) and didn't have to worry about who directed a film, when it was released, where it was filmed, its budget, who starred in it, or what its genre or plot elements were. All the relevant information was implicit in how the people who saw it rated it out of ten.
- NumberCruncher 9y agoCould be interesting for you: https://en.wikipedia.org/wiki/Netflix_Prize https://en.wikipedia.org/wiki/Netflix_Prize
- scj 9y agoThank you for the link, I intend on studying it carefully this weekend. I may want to implement something similar using board game data from boardgamegeek.com (I've already been working on a recommendation system). BGG has a 1-10 point rating scale (using floating point numbers), so I suspect the method described will fit. My only concern is the number of mutually rated titles between users in board gaming is probably lower than movies. Which I suspect will reduce confidence rates. The current approach (with a Jaccard system), requires a degree of human intervention and only works on some users. It meets the goal of recommending titles for me, but it'd be nice to expose the system externally.
- NumberCruncher 9y agoFunny fact: 4 years ago I was looking for something what I would call today a "transitive similarity index" to measure similarity of shopping baskets of brick-and-mortar stores. Because I didn't find anything I "invented" my own algorithm. This year I implemented the same algorithm in a system which is live since yesterday. Googling up the Jaccard similarity I just realized that I re-invented a transitive version of it. Now my baby has a name. :)