3 ms·
Winning a kaggle contest and how a particular statistical model performed under normal business conditions are totally different. (metrics of interest for evalu
by svasan 14y ago
Winning a kaggle contest and how a particular statistical model performed under normal business conditions are totally different. (metrics of interest for evaluating the model - how much money was saved, did the model bring about requisite behavior change, etc.) And performance under real life business conditions is what matters, not who won the contest. And to get good models for a specific business need, you do need domain knowledge.
Does kaggle publish how the models performed under normal business conditions?
- marshallp 14y agoI'm not following your line of reasoning. Everything is data at the end of the day. All you're doing is creating a predictive model. If the business conditions change, you wold simply send it to the data scientists to reformulate a new model. That's no different to sending it to kaggle again.
- svasan 14y ago>> Everything is data at the end of the day. But you have to interpret the data within the context of the business need/requirement. Building a credit risk model is vastly different from building a (personal) insolvency/bankruptcy model though both may entail the same set of steps in developing the model. The variables that make it to the model depend on the business need. In kaggle, one of the datasets that I messed around with had variable labels as Var_1, Var_2, ..., Var_X. So while fitting a model, I would not know why a particular variable made it into the model. You can see that this kind of variable labeling does not give me any insight into how that variable was generated. I need to know whether the variable was raw/aggregated/transformed etc. And that takes you back to understanding the data in the context of the business/domain.
- marshallp 14y agoWhy not just give all data (or as much as you can afford to) to the data scientist and let them figure it all out.
- svasan 14y agoA data dump would not really help because the data could have been influenced by 1) a key company policy 2) specific business activity 3) input coming from another model It is better to a) define the problem, b) collect the data, c) build the variable library, d) and then fit the model rather than jump to step (d) directly because the modeler/scientist has greater understanding of the entire set of data going into the model development. It is very likely that the modeler would uncover any/all of the three influencing factors I mentioned above, during the data collection stage. While kaggle is an interesting concept, from a different perspective it looks like an "effort harvesting" operation. For a pittance, the companies/institutions that are sponsoring the contests are getting a steal. (I am not sure if the million dollar prize is still up for grabs.) However, for folks who do want to break into data sciences/statistics field, kaggle certainly is a good platform to get acquainted with data science/statistics related skills.