6 ms·
Edit distance of titles?! Do you have a source? I'm very curious about how and why that would help.
by ambition 18y ago
Edit distance of titles?! Do you have a source? I'm very curious about how and why that would help.
- trevelyan 18y agoIndiana Jones and the _______________.
- DaniFong 18y agoHere's a paper on the BellKor solution, from one of the top teams: http://research.att.com/~volinsky/netflix/ProgressPrize2007BellKorSolution.pdf http://research.att.com/~volinsky/netflix/ProgressPrize2007B...
- elq 18y agoYehuda later wrote http://glinden.blogspot.com/2008/03/using-imdb-data-for-netflix-prize.html#c4525777289161540480 http://glinden.blogspot.com/2008/03/using-imdb-data-for-netf... and http://hunch.net/?p=331 http://hunch.net/?p=331 that using movie metadata has produced no measurable improvement in RMSE.
- DaniFong 18y agoThe crux of the argument though, is that if you have a strong CF model with many many ratings, you don't seem to get much benefit with their approach (linear combination of models). That doesn't mean that metadata can't be useful with a different approach. It also doesn't mean that metadata isn't useful for sparse data: in fact, it's incredibly useful, because you don't have much of anything else.
- elq 18y agoI cannot dispute that metadata can be useful. But it appears, at least for prediction tasks similar to the prize, that an ounce of weak or strong explicit user input is worth a ton of rich implicit data (including item metadata).