6 ms·
I've heard that Kaggle data sets encourage people to do "supervised" ML only. Is that true?
by bhnmmhmd 8y ago
I've heard that Kaggle data sets encourage people to do "supervised" ML only. Is that true?
- hideo 8y ago(Not Ben, but - ) outside of academia, the main thing that seems to encourage people to do supervised ML is that it's the only thing that seems to work. I haven't really heard of any success stories with using unsupervised techniques for most common ML applications.
- harias 8y agoWhat about clustering?
- dotancohen 8y agoI used a very simple unsupervised ML built in scikit-learn to find good matches on OK Cupid. Worked very well, it found definite boundaries between the clusters of women. One of the features was a subjective rating of how much I liked some of the women, and scikit-learn then suggested to me other women in the clusters that had my best ratings. It turns out that I like vegetarians, redheads, and left-wingers. Which happens to be true, even though I eat meat and do not identify as left-wing. But those traits correlate with _other_ traits that are more difficult to measure objectively, such as caring about children, liking to hike, and preferring an evening of sex to an evening of television.
- raverbashing 8y agoUnsupervised works, but your ability to measure "does it work or not" is much more dependent on a case by case evaluation rather than a score. (Because if you know a priori what is it that you want to measure - it's supervised)
- atupis 8y agoYeah this my experience too, evaluation ends being almost endless time sink.
- laichzeit0 8y agoI'm not an expert, but I feel that: Unsupervised techniques work really well for language modelling. There is also weakly supervised and distant-supervision, where the labels are "noisy" or not exactly what you want. You're right in that strong supervision, where you basically trust your class label, works really well, because it's probably the easiest case. Combining unsupervised (e.g. pre-trained language models) with a very small set of strongly labeled data, or a larger set of weakly labeled data, seems to work pretty well too.
- taneq 8y agoI think it's more that supervised ML is sufficient for most of the low hanging fruit. It's relatively easy and well-understood, and there are a lot of things out there where we have copious data that we just need to digest into a model to make it useful.
- stuartaxelowen 8y agoNot at all - I released a customer support on Twitter dataset there specifically focused on unsupervised tasks! I think the focus on supervision in what people do with the data shows that there are still a lot of people poking around with the easier supervised tasks. [0]: https://www.kaggle.com/thoughtvector/customer-support-on-twitter https://www.kaggle.com/thoughtvector/customer-support-on-twi...
- benhamner 8y agoThe competitions we host (https://www.kaggle.com/competitions https://www.kaggle.com/competitions) are supervised and always have a target we can create a numeric leaderboard on, but the public datasets (https://www.kaggle.com/datasets https://www.kaggle.com/datasets) are used for everything under the sun. There's some supervised ML use of those, and a lot more open-ended exploration, visualization, cleaning, clustering, language modeling, etc.