5 ms·
When I hear "Scientist" I tend to think "PhD level work." Graduate level work in math, Comp Sci and statistics is not something that can be readily compressed
by EzGraphs 14y ago
When I hear "Scientist" I tend to think "PhD level work." Graduate level work in math, Comp Sci and statistics is not something that can be readily compressed into a 9 month program without substantial prerequisites.
Are today's "data scientists" really just software devs who have specialized in digging around in data and using various data mining algorithms with only a superficial understanding of their inner workings?
- achompas 14y agoAre today's "data scientists" really just software devs who have specialized in digging around in data and using various data mining algorithms with only a superficial understanding of their inner workings? Ideally, no. We're witnessing an overuse of the name "data scientist" (which has its own problems, but that's another story). There's a non-trivial difference between a data scientist who understands the theory used for the EM algorithm or belief propagation, and a "data scientist" who is performing large-scale data analysis using various data mining tools. Unfortunately, they're both getting lumped together. To become one of the former, you need graduate-level maths, CS, and statistics, while this certificate caters to the latter.
- Goladus 14y ago> When I hear "Scientist" I tend to think "PhD level work." There's a lot of non-PhD level work that goes into a large research project, often for aspects you may not have considered like animal care. A "Data Scientist" might be a specialized software developer, but it doesn't follow that they have only a superficial understanding of the inner workings, even if that understanding is not good enough on its own to do much original research. This is the first I've heard the term "Data Scientist," though Harvard recently announced a masters-level "Computational Science and Engineering" degree. (http://news.harvard.edu/gazette/story/2012/06/a-new-masters-program/ http://news.harvard.edu/gazette/story/2012/06/a-new-masters-...) The bottom line is that due to the amount of data being generated by research, demand for programmers to help deal with it is rising.
- Homunculiheaded 14y agoWhile I'm sure we have different definitions of "superficial understanding" one thing I've noticed as I've gotten more interested in ML/datamining during the final stages of my master's in CS is that solving real world problems with these techniques is often a very different experience than deeply understanding the theory behind them. For example I couldn't implement an SVM library from scratch to save my life, but I do understand what it means to be a 'maximum margin' classifier, from a high level how the 'kernel trick' works, and why you would tune regularization and cost parameters. However this knowledge has been enough to help me in quite a few interesting problems. Reading accounts of how others have solved real world data mining issues it's amazing how often a very simple model will do the job, and also how often, even among more serious researches, there's a bit of intuition in finding the right combination of parameters, and lots of trial and error in searching for which model/blend of models really does the job. I think there's a lot of room for more people approaching data mining with the 'hacker' mentality. Sure you don't want 'data scientists' using a randomForest whose eyes glaze over when you mention the word "ensemble", or someone who couldn't explain in plain terms what a "maximum margin hyperplane" is. But, there is a growing space for practitioners in this space, that aren't necessarily as strong in the theory as people working in the pure research space.
- EzGraphs 14y agoMuch of what you say resonates with what I have seen from a different vantage point - old school software dev who has dabbled in data mining. Simple models often work and are preferable - easier to explain. More data and simple models generally provides superior results to small amounts of data and complex models. It seems like "ensemble" methods - combining the results of several different algorithms - is generally a less-than-rigorous exercise that involves throwing a bunch of different approaches at the problem and averaging the results. It is good to hear that there is "a growing space for practitioners in this space, that aren't necessarily as strong in the theory." But the term "Data Scientist" seems a bit lofty for folks doing this sort of work.
- Homunculiheaded 14y agoThe thing is that "less-than-rigorous exercise" is true in many areas of ML. Take for example neural nets, which are very popular and successful, even among real expert's there a lot of 'magic' behind why they really work. SVMs are loved partially because they work well, but also they are very sound from a theoretical standpoint, if you know the math you can show that it will work, this is not necessarily true with many other successful techniques. Interesting side note for ensembles: 'averaging' is usually not one of the best methods for blending results. More successful approaches include using either a perceptron or a simply training a linear model to find appropriate weights for predictions from each individual model. I've even had a case where simply picking the MIN of each set of predictions worked surprisingly well for a particular problem. The above btw is something that I think a "Data Scientist" should know, and is well out of the scope of a software engineer who just plugs values into prepackaged algorithms. A "data scientist" should be able to read papers [1] that explain these things, which is more than many software engineers do. Now I'm not a data scientist, but while I can't write an SVM from scratch, when I'm working on data mining problems I am reading several academic papers a week. I really think we're looking at two sincerely distinct areas of expertise and it's not too lofty to look at someone who has to read academic papers to do his job as a "scientist". [1] http://www.edscave.com/docs/Blending_Methods_AusDM2009.pdf http://www.edscave.com/docs/Blending_Methods_AusDM2009.pdf