4 ms·
Wow, this is exactly what I have been searching for for months! I've got a pretty basic grasp on machine learning and collaborative filtering, so I understand
by eggbrain 16y ago
Wow, this is exactly what I have been searching for for months!
I've got a pretty basic grasp on machine learning and collaborative filtering, so I understand some things, but I'm very confused on others, such as:
1)List-wise vs Pair-wise approaches to machine learning. Can you explain in simple terms what is the main differences between them, and in what cases it would be better to use one over the other? I've read a few sources about the differences but it goes over my head a lot of times.
2)When you don't have many users using your site, from my (basic) knowledge, you can't really use KNN algorithms to help with recommendations, because you only have a few people to compare your (lets just say movie preferences) to. What is the best way, then, to get the best recommendations both when your userbase is large and small?
Those are just a few off the top of my head, but I'll be sure to add more later on.
- bravura 16y ago1) Pair-wise approaches, if I understand what you mean, are those in which you are look at pairs of examples. You want to avoid naive implementations, in which you train on a quadratic number of instances (all-pairs). It really depends upon the specific problem how you get around this. For example, if you have a model that operates over pairs, doing example selection in an intelligent way. For example, use all positive pairs, but sample only one negative pair per positive pair. 2) You are talking about the cold start problem. Basically, the problem is how do you deliver good recommendations to new users as quickly as possible. I am actually currently talking to a big company about this problem. Google has talked about how they do this for Google News, especially because they have high item churn (news is not news in a day). http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.80.4329&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.80.... Essentially, they do clustering of users and news items. Another approach is to give the user a hunch-like quiz when they join. Basically, you use a decision tree to cluster the users, getting as much information from them with as few questions as possible. Another approach that requires almost NO user interaction is to use some ambient information about the user. Maybe you have a cookie for the user and other sites will share the user's behavior with you. Maybe you ask them to log in through facebook or twitter, and then can look at their profile and activity history to build a user profile. So getting creative about possible sources of ambient information is another approach.
- eggbrain 16y agoI need some clarification on #1: Could you provide easy examples where you would use a pair-wise approach, and a list-wise approach? Like, let's say I wanted to build a recommendation system for college football pick-ems (Ie, will Iowa beat Michigan State? Etc.). Which method would be preferred, and why? Another question: Let's say (again) that I want to build a recommendation engine for college football picks. How do I know how much (more) data will help for training purposes? IE, will only Win/loss records suffice [probably not], or should I add in the opponent they beat (or were beaten by) as well? What about whether the game was away/at home? What about factoring in the individual players on each team, and their stats? What about the temperature that day? How you do end up knowing when you can stop looking at smaller and smaller variables with your training data, and have something that gives the best results?
- alextp 16y agoYou can do greedy feature selection: make a stable train and development and test (and maybe another truer test set, for use later) sets and define an appropriate quality measure. Then implement the simplest thing you can think of, train on the training data, calibrate the results on the development set (tuning hyperparameters, etc) and look at the performance on test data. That's your baseline. Now think up of a nice new feature you'd like to add, implement it, and see if you can get the error on the development set to go down. If it goes down a lot (you can decide an appropriate threshold, maybe depending on the computational and human cost of using those features), keep it. Otherwise throw it away. Now repeat this process for every set of features you can think of, documenting the combinations you've tried as you go to see where is the best effort/performance tradeoff. Then test your best model on the test set to see if the performance has really improved. If your model is linear, you will probably see diminishing returns as you implement more, different features. It also helps to look at model errors to see which features are pulling things the wrong way, and then you can add other features to compensate. (feature engineering + linear models feels a lot like writing an old-school AI heuristic program, except you have a computer program give a numeric weight the rules you're writing, so you're free to write a lot (millions) of rules and still get a manageable model)
- 16y ago