3 ms·
A data dump would not really help because the data could have been influenced by 1) a key company policy 2) specific business activity 3) input coming from ano
by svasan 14y ago
A data dump would not really help because the data could have been influenced by
1) a key company policy
2) specific business activity
3) input coming from another model
It is better to
a) define the problem,
b) collect the data,
c) build the variable library,
d) and then fit the model
rather than jump to step (d) directly because the modeler/scientist has greater understanding of the entire set of data going into the model development. It is very likely that the modeler would uncover any/all of the three influencing factors I mentioned above, during the data collection stage.
While kaggle is an interesting concept, from a different perspective it looks like an "effort harvesting" operation. For a pittance, the companies/institutions that are sponsoring the contests are getting a steal. (I am not sure if the million dollar prize is still up for grabs.) However, for folks who do want to break into data sciences/statistics field, kaggle certainly is a good platform to get acquainted with data science/statistics related skills.