7 ms·
> why can't data science simply be about applying the scientific method in the realm of data analysis? That's what a statistician do. I've seen these ML and D
by digitalzombie 8y ago
> why can't data science simply be about applying the scientific method in the realm of data analysis?
That's what a statistician do.
I've seen these ML and Datascience people. And the majority the time how they tackle data is radically different from statistician and is more of an art than a science compare to what statistician does.
But this could be my bias opinion and just some small data sample from personal experiences.
---
Actually my last day of internship I've met a few statistician interns some of them are from Cal (UCBerkely) and they came to the same conclusion (we have a lot of complaints). The ML/DS group is really just doing black magic (nicest way of putting it). I wish statistic is better at marketing. Oh well.
- xapata 8y agoThe amazing thing is that the black magic gets results (for certain categories of problems).
- digitalzombie 8y agoI've seen them do random forest on temporal data. They tried to fix it with PCA.
- citation_please 8y agoWhat do you mean by "on temporal data"?. For example, no feature extraction was performed? This sounds pretty amateurish, and completely below the understanding levels of the many machine learning papers that I've read - which is admittedly a small percentage of what has been published.
- xapata 8y agoI've seen "them" test in production. Every field has varying levels of skill.
- natalyarostova 8y agoIn my humble experience as a data scientist at a big tech company, a big differentiating factor is familiarity with the R or Pydata stack. It's not just its own language and library, it's a more general idiomatic way toward approaching and solving problems from a "software first" perspective. I sometimes work with people who are formally trained in stats at a much higher level than myself. But with my lesser ability combined with my capability to write production grade Python code using the pydata stack, my code and solutions are more likely to make it in.
- digitalzombie 8y agoI think we're talking about different roles. I'm coming at it from just a modeler role and it seems like you are stating more of a full stack within the data science. I took OP remark as if it was just within data analyst in term of analysing the data and not including turning it into some kind of application. As for Python, time series isn't as good in python. I do agree that Python production is good. But I don't believe Python have better analysis packages. I'm in the camp of using programing language for their strength and using Apache Thrift or whatever to tie everything in.
- curiousgal 8y ago> As for Python, time series isn't as good in python Time series isn't as good Linear regression isn't as good Mixed models are awful Good luck fitting a spline without diving into scipy and doing it as an interpolation R packages are almost always accompanied with a rigorous paper and a great vignette, whereas in Python, there are almost no talks of the implementation and just documentation about how you can use the library. Every model we develop has assumptions and shortcomings, R offers the tools to diagnose and examine that, whereas in Python, I get the feeling that it's more like "some smart guys figured out this formula, here's an implementation of it".
- digitalzombie 8y agoI agreed, I didn't want come off as a disgruntle statistician and list the weakness of Python. But you're right. I find most Springer, CRC, etc.. books are in R with accompanying R packages. Shumway & Stoffer Time series book comes with the R astra package. Andrew Gelman Bayesian create Stan and their target was R first (rstan). The Sanford people who created lasso and elasticnet created glmnet. The thought that the creators or experts of these subject wrote a book or research paper and then also publish R package to accompany it is reassuring. When it comes to just data analysis, it seem like R is pretty good or better than Python. I think Python is much better for ML stuff such as deep learning and also it seems like Python have better NLP support. Arguably Stanford is very good with NLP and they publish their tool in Java.
- vasili111 8y agoHow good ML and Datascience people that you have met were at statistics and in math in general?
- lottin 8y agoIn social sciences (speaking as an economist), the goal of statistics was always testing theories and hypotheses against empirical evidence. Once you have a solid theory, then you have an understanding of the subject matter, which allows you to make predictions and counterfactual hypotheses. So the primary goal was understanding the data generating process, finding causal relationships between variables, and so on. The goal was specifically not making predictions. One of the first things I learnt was that a bad model can make good predictions (for a while). The ML crowd appears to be taking the opposite approach, focusing on making predictions and disregarding everything else. I suppose it has its uses, but I wouldn't call it science.
- ur-whale 8y ago>> why can't data science simply be about applying the scientific method in the realm of data analysis? >That's what a statistician do. Mmmh. Run that experiment for me next time you meet a statistician: - ask him if he can apply Chi-squared to a decision problem - ask him if he can *explain* how and why Chi-squared works. In my experience, all statisticians can do the first, almost none can do the second. Learning how to use a screwdriver to screw screws without understanding notions of torque and moment doesn't mean you're applying the scientific method.
- digitalzombie 8y ago> In my experience, all statisticians can do the first, almost none can do the second. I think a phd statistician can do this. Master statistic does not touch upon field and measure theory in statistical inference so many questions get unanswered. But I suspect you may be correct. I view chisq as a statistical distance for most of my encounter and learning.
- pimmen 8y agoI would argue that they should be able to understand it to the level that they can at least defend Chi-squared as a tool for the problem at hand. Then, they should be able to evaluate whether or not it works correctly. If a medical researcher is testing a new radio-therapy treatment, but can't mathematically model every fission problem you can throw at them, they're still applying the scientific method.
- citation_please 8y agoThis, and your other responses in this thread, come off as rather disappointing to me, as someone who considers the work they do as "data science". Machine learning as a method clearly has quite a bit to contribute to the business world based on revenue alone. My argument could rest here. But it's also disappointing the way that you belittle all machine learning practitioners, even those with academic credentials, for their work not being worthy of serious consideration. This also sounds a little defensive and projective, and I can't imagine it's easy seeing the forest for the trees with your head in the clouds.
- digitalzombie 8y agoTrying to walk on a rope is pretty hard balancing act. I do want to give a personal view and constructive criticism but any intentional and direct attack toward another domain is not my intention. I think statistic is very well equipped to do just data analysis. The discipline have many weaknesses I do acknowledge that but I don't believe data analysis is one of it when it is the core tenet of what statistic is. I believe data science is too new and is still trying to find it's standing. Also it seems like a jack of all trade and a master of none discipline. I don't believe the discipline can be a master of everything including data analysis with all other things it's trying to incorporate. My critique is that for the original post is that there is already a field for data analysis. Just for it, for a century now, it's named statistic. ML/DS is a new breed that is more than just data analysis. Because of this they can do many things but I don't believe they are an expert at any one thing. Which is fine. But I do understand that this discussion is sensitive since people within the DS/ML since that is how they make money and earn a living.
- autokad 8y agoif you really understand ML/DS, then you would know that ML is founded on all the same concepts. > "... they tackle data is radically different from statistician ..." the short answer is, before we had limited data, compute, and the problems businesses needs increased substantially. now we can do better things, and if you cant do more than what companies were doing in the 1950s, then your value proposition is significantly less than someone that's using industry practices established in 2017 / early 2018. > "I wish statistic is better at marketing. Oh well." as a data scientist, I have to know stats, ML, big data tools (spark/splunk/hdfs), programming, and domain specific knowledge. some data scientists are just rebranded statisticians, which is fine. the role of a data scientist varies to the needs of the org they are in. I mostly do anomaly detection and classification, while others may focus solely on AB testing. > > "> why can't data science simply be about applying the scientific method in the realm of data analysis?" the statement doesnt make sense to me. if you are doing data analysis then by default you are doing observations and measurement on empirical data. As far as experimentation goes, almost no one has the resources or the standing to do that. Some do, like those AB testers I mentioned before. > "Actually my last day of internship I've met a few statistician interns some of them are from Cal (UCBerkely) and they came to the same conclusion (we have a lot of complaints). The ML/DS group is really just doing black magic (nicest way of putting it)" Then they don't appreciate ML and the problems they are trying to tackle. It reminds me of a stats PHD co-worker that was upset about the number of parameters in an image classification problem we had. they didn't think a model could be ran that had more parameters than observations. I was like we are not running linear models ... As an experiment, go to kaggle and try the spooky author classification competition. See how far you can get with basic stats and then see how much further you can get with ML. hopefully it will give you an appreciation for some of the ML tools available. It's not black magic, like many other domains; its just a dense subject that takes a lot of time to understand
- 3pt14159 8y agoYeah, most of this thread has kinda garbage responses. Even if you did a CS degree and then did a BMath with a statistics major you still wouldn't have all the skills you would need, though you'd have a great place to start off from. There is an art in kinda guiding / interpreting the mathematics that's hard to teach and most people in the field kinda pick it up through experimentation. I'm sure it will get more formalized at some point, but I don't think the formalization will be all mathematical or scientific; I think much of it will be making explicit how to think about certain domains or a system of process that are useful. Also I think there is a big difference between Data Engineers, Data Analysts, Data Scientists, and AI researchers. Depending on the field, an AI researcher deals more with idealized forms and is much stronger on the mathematics of machine learning. Data engineers are usually stronger on the CS topics that deal with scale. Data analysts, even the very best ones, tend to be weaker on the CS, the math, but tend to be the strongest at quickly understanding data and making it intelligible to decision makers. I'm not knocking it—I was an analyst at one point—it's just the truth that some positions take more skill than others and most great data analysts tend to treat it as a 3 or 5 year stop before levelling up to management or a more technical role. Data Scientists generally have the skills of a data analyst (though they have trouble dumbing things down at times if they came here from something other than an analyst position) with some of the skills from the other two. The way some application code software developers dismiss a whole sub-discipline is kinda embarrassing.