4 ms·
I agree with you, though I actually personally don't find pre-processing data tedious and boring. I kind of like knitting it all together. On a side note... da
by geebee 7y ago
I agree with you, though I actually personally don't find pre-processing data tedious and boring. I kind of like knitting it all together.
On a side note... data pre-processing is often viewed a side job that needs to get done before the real work can begin. I don't think I've ever been able to prepare a data pipeline without making decisions about the data that will impact the outcome. For example
How do you deal with missing data? Interpolate, ignore, use averages, use a machine learning algorithm to plug the gaps?
How do you decide what data sources to include in the pipeline. What if one data source seems more reliable, but another has far more data, too much to use. Should you amplify it, sample from the other data sets to make the volumes equal, keep things proportional?
What if one data set changes more rapidly than another, how should you update the data set used to populate the ML model?
These are just a few that jump into my head, and while there are techniques to deal with them, ultimately, there isn't a correct answer. And honestly, even these examples make it all seem more glamorous than it is, a lot of this is just figuring out why various encoding and formatting errors are breaking the feed, why column headers mysteriously change, why handwriting on form scans gets properly translated into text some of the time and completely garbled in others.
The funny thing is, I do see people fine tuning ML algorithms (kaggle style, seeing if they can wring a bit more predictiveness out of a model), when in the real world projects, decisions upstream about the data pipeline will have an impact perhaps 5-10 times greater than any tweak to the ML parameters (or even which general algorithm to choose). And yet, it's hard to get people to even pay attention to these decisions, probably - as you said - because people find it catastrophically tedious and boring.
- emmanuel_1234 7y agoAs a corollary, "Data Scientists" who can't program their way out of a paper bag (e.g.: write simple SQL, or a scraper in Python) is near useless, and a stress on their peer who can. I'd much rather hire a good programmer with some statistical knowledge than the other way around.
- dawg- 7y agoEither way, you are gonna be paying someone to fill in gaps in their knowledge. Anyone with a laptop and an internet connection can learn to "program their way out of a paper bag" in a weekend.
- geebee 7y agoI kinda disagree, though I suppose it depends on the paper bag.
- omar_a1 7y agoI've seen so much bad math and misunderstood statistics hard coded into widely-implemented software that that's not ideal either. Ultimately, you need a team with a combination of strengths if your product requires multidisciplinary work. Otherwise you end up with hilariously wrong equations/assumptions in your code base. But, full disclosure, my background is math/science.