3 ms·
When we host competitions on Kaggle, a lot of work goes into asking the right question and structuring the problem. The domain expertise is incorporated in this
by benhamner 14y ago
When we host competitions on Kaggle, a lot of work goes into asking the right question and structuring the problem. The domain expertise is incorporated in this step, as well as in putting the competition results to use in production.
This splits the "domain expertise" and "predictive modeling" components into two separate chunks. While domain expertise is crucial for asking the right questions, we've found that it isn't as necessary for the predictive modeling component. For example, in the essay scoring contest we hosted, none of the winners had touched natural language processing prior to the contest. However, they beat out many experts with decades of experience in NLP.
For an internal data science team, the "domain expertise" component is at least as important, as they are charged with asking the right questions as well. However, this does not mean competition winners cannot develop and learn this - they have already demonstrated their creativity and tenacity in one domain (applied machine learning), and this carries over nicely to other domains from our experience.
- marshallp 14y agoYou're actually from kaggle so I'm going to look like a troll arguing with you so I won't try (though I do privately think the right question is pretty obvious always - will this combination of parameters make profit - or some other obvious single metric. Just give all data to the data scientist and have them build the model).
- svasan 14y agoI'd be curious to know if kaggle measures/publishes model performance, model degradation, etc., for the models that the companies/institutions ended up incorporating in their business activities. edit - for clarity.
- rm999 14y ago>However, this does not mean competition winners cannot develop and learn this - they have already demonstrated their creativity and tenacity in one domain (applied machine learning), and this carries over nicely to other domains from our experience. Strongly agreed. I actually got into the field through a company's data mining contest (pre-kaggle). I think people who are strong at building predictive models are great candidates for data sciences. But it took years of work experience after doing my graduate degree in machine learning to get to a point where I'm comfortable calling myself a decent data scientist. It's easy to think model-building is the only important skill-set, but data and models don't exist in a vacuum; a more holistic view of where your data comes from and how your work will be used is essential. This excellent netflix blog entry illustrates what I'm saying quite well, I think. http://techblog.netflix.com/2012/04/netflix-recommendations-beyond-5-stars.html http://techblog.netflix.com/2012/04/netflix-recommendations-... They make two points that illustrate the divide between a contest and the day-to-day work of a data scientist: * The winning model was not usable in production. Netflix had to gut the 100+ model ensemble to a much simpler 2 model ensemble * Business needs change, the question they were trying to answer changed from the start of the contest