4 ms·
All the evidence from kaggle indicates that deep domain knowledge is not required. Jeremy Howard has some youtube videos discussing this. Pretty much all the sk
by marshallp 14y ago
All the evidence from kaggle indicates that deep domain knowledge is not required. Jeremy Howard has some youtube videos discussing this. Pretty much all the skills you outlined (except for production grade code - which is a software engineer problem, not a data scientist problem) are covered by the contest.
- rm999 14y ago> All the evidence from kaggle indicates that deep domain knowledge is not required That's irrelevant, data scientists don't do data mining contests for a living. In my experience finding the right question to answer is a large chunk of data science, and that is never spelled out for you like it is in a contest.
- marshallp 14y agoYou're saying kaggle is fundamentally different to what data scientists do? I don't understand that. The right question is usually how do I increase profits (or score this essay/image etc). So simply set up your problems that way.
- rm999 14y ago>You're saying kaggle is fundamentally different to what data scientists do Yes. Your view of data science is extremely narrow, there is more to it than creating a predictive model to optimize a single metric. Reread the second paragraph of my first comment.
- marshallp 14y agoMy first comment to your first comment answered those objections. I consider all problems as optimizing a single metric. Anyway, let's let this slide.
- benhamner 14y agoWhen we host competitions on Kaggle, a lot of work goes into asking the right question and structuring the problem. The domain expertise is incorporated in this step, as well as in putting the competition results to use in production. This splits the "domain expertise" and "predictive modeling" components into two separate chunks. While domain expertise is crucial for asking the right questions, we've found that it isn't as necessary for the predictive modeling component. For example, in the essay scoring contest we hosted, none of the winners had touched natural language processing prior to the contest. However, they beat out many experts with decades of experience in NLP. For an internal data science team, the "domain expertise" component is at least as important, as they are charged with asking the right questions as well. However, this does not mean competition winners cannot develop and learn this - they have already demonstrated their creativity and tenacity in one domain (applied machine learning), and this carries over nicely to other domains from our experience.
- marshallp 14y agoYou're actually from kaggle so I'm going to look like a troll arguing with you so I won't try (though I do privately think the right question is pretty obvious always - will this combination of parameters make profit - or some other obvious single metric. Just give all data to the data scientist and have them build the model).
- svasan 14y agoI'd be curious to know if kaggle measures/publishes model performance, model degradation, etc., for the models that the companies/institutions ended up incorporating in their business activities. edit - for clarity.
- rm999 14y ago>However, this does not mean competition winners cannot develop and learn this - they have already demonstrated their creativity and tenacity in one domain (applied machine learning), and this carries over nicely to other domains from our experience. Strongly agreed. I actually got into the field through a company's data mining contest (pre-kaggle). I think people who are strong at building predictive models are great candidates for data sciences. But it took years of work experience after doing my graduate degree in machine learning to get to a point where I'm comfortable calling myself a decent data scientist. It's easy to think model-building is the only important skill-set, but data and models don't exist in a vacuum; a more holistic view of where your data comes from and how your work will be used is essential. This excellent netflix blog entry illustrates what I'm saying quite well, I think. http://techblog.netflix.com/2012/04/netflix-recommendations-beyond-5-stars.html http://techblog.netflix.com/2012/04/netflix-recommendations-... They make two points that illustrate the divide between a contest and the day-to-day work of a data scientist: * The winning model was not usable in production. Netflix had to gut the 100+ model ensemble to a much simpler 2 model ensemble * Business needs change, the question they were trying to answer changed from the start of the contest
- svasan 14y agoWinning a kaggle contest and how a particular statistical model performed under normal business conditions are totally different. (metrics of interest for evaluating the model - how much money was saved, did the model bring about requisite behavior change, etc.) And performance under real life business conditions is what matters, not who won the contest. And to get good models for a specific business need, you do need domain knowledge. Does kaggle publish how the models performed under normal business conditions?
- marshallp 14y agoI'm not following your line of reasoning. Everything is data at the end of the day. All you're doing is creating a predictive model. If the business conditions change, you wold simply send it to the data scientists to reformulate a new model. That's no different to sending it to kaggle again.
- svasan 14y ago>> Everything is data at the end of the day. But you have to interpret the data within the context of the business need/requirement. Building a credit risk model is vastly different from building a (personal) insolvency/bankruptcy model though both may entail the same set of steps in developing the model. The variables that make it to the model depend on the business need. In kaggle, one of the datasets that I messed around with had variable labels as Var_1, Var_2, ..., Var_X. So while fitting a model, I would not know why a particular variable made it into the model. You can see that this kind of variable labeling does not give me any insight into how that variable was generated. I need to know whether the variable was raw/aggregated/transformed etc. And that takes you back to understanding the data in the context of the business/domain.
- marshallp 14y agoWhy not just give all data (or as much as you can afford to) to the data scientist and let them figure it all out.
- micro_cam 14y agoProper featurization is a hugh part of working with real data. Datasets in contests tend to have already been analyzed enough to remove all features that are correlated, caused by or predictive of but irrelevant to the target feature. For example, early analysis of a "fresh" data set involves a lot more "'of course length of hospital stay' correlates with 'had disease'" or "'Phone number' is a unique identifier in this data so of course it shows up most often in random forest predictor" moments. Putting together weighted ensembles in R of a clean data set is basically manual labor but doing the same analysis at scale involves a lot of of nuances that should be considered data science. In particular you need to be able to determine if a problem can be split into mostly independent parallelizable sub problems (which also often requires domain knowledge) or reduced/relaxed into something that has a well established optimized solution (like matrix decomposition). And finally you need to be able to determine if your prototype and the production code are converging to the same results and debug it if it isn't which is non trivial with stochastic algorithms.