4 ms·
We have a number of options, with SVM being the fastest so far b/c the library is written in C. Pearson, NaiveBayes and a number of others have all been tried a
by haidut 17y ago
We have a number of options, with SVM being the fastest so far b/c the library is written in C. Pearson, NaiveBayes and a number of others have all been tried and "tested" for acceptance with several thousand users and the majority picked the results from SVM as the most accurate. The recs are computed offline but there is an option to do that on the fly given enough RAM to load the SVM models for all users.
As far as asking users for too much - what exactly do you mean? All we are asking for is for the user to login and if they find article they like (on the front page or from the topics menu) then click on them. What more simple that that? The voting system up/down you talk about is exactly the same except that you explicitly ask the user to vote, while we implicitly do the same.
TechnologyReview ran an article last year how asking the user to explicitly rate stuff is considered overburdening and how the system should quietly monitor the user without asking direct questions. Anyways, your up/down approach and our click/remove approach are the same in my opinion.
But thanks for the comments.
- physcab 17y agoSorry, when I wrote the comment I was on my Iphone and didn't check the website. I just did. I see what you mean. Why don't you just aggregate total click data and see which articles get clicked the most, bin the results to categories (5,4,3,2,1) by topic then compute Pearson? SVM still sounds a bit heavy weight. Sure, I guess it'll handle a few thousand...how about 1-10 million? And how do you know if SVM is the most accurate? What's your criteria?