3 ms·
The messiness of data is something I even see creating a growing rift between academic ML/DS and real world applications. What makes for a nice paper doesn't n
by huffmsa 6y ago
The messiness of data is something I even see creating a growing rift between academic ML/DS and real world applications.
What makes for a nice paper doesn't necessarily make for a model that will survive contact with new user generated data.
- hef19898 6y agoI remember some similar discussions when I worked in logistics consulting for a while. The data source was a mess, like a mess. Data was even included as screenshots of spreadsheets in other spreadsheets. Some math PhD was in charge of that modelling. Without any domain knowledge concerning the data (logistics, consumption, maintenance) and thus unable to properly interpret the raw data to begin with. Most of the time was spent on writing some Python scripts to analyze the raw data, still full of errors. And build predictive models on top of that mess. Kind of formed my view of data science, unfairly so.
- disgruntledphd2 6y agoJust like almost all software work is maintenance, almost all data science is data cleaning.
- jmatthews 6y agoI liken it to sending a chef to a grocery store. It's not just about being a good cook. Half the battle is in choosing the correct ingredients. Not just a dozen eggs, but free range where the yolks will be a vibrant orange yellow and improve the presentation. The cleanest models with the highest fidelity often fall out as the next obvious transformation of a well groomed and hygienic dataset.