5 ms·
The State of Data Science and Machine Learning
- technologia 9y agoComparing my own situation, I fit in pretty much in the median for my field, age and salary. It'll be interesting to see what folks dig up over the coming weeks from this dataset.
- willis77 9y agoNothing warms our icy, cold, statistical hearts quite like hearing that a randomly chosen person is near the median. <3
- antgoldbloom 9y agoUnlike the statistician who has his legs in the freezer and his head in the oven but who is on average the right temperature.
- bootcat 9y agoReally worthy insight. Especially for people like me, who wants to get a deeper understanding of the ecosystem before being an actual data scientist !
- antgoldbloom 9y agoFor interest, the raw data is published here: https://www.kaggle.com/kaggle/kaggle-survey-2017 https://www.kaggle.com/kaggle/kaggle-survey-2017 And some early analysis from our community here: https://www.kaggle.com/crawford/analyzing-the-analyzers https://www.kaggle.com/crawford/analyzing-the-analyzers Some things that jumped out at me: 1. more people learn data science and ML from MOOCs than university courses 2. Tensorflow the tech people most want to learn in the next year 3. 40% of people survey spend >1-2 hours per week searching for another job. Surprising given all companies complain about the difficulty in finding data scientists/machine learners.
- PeachPlum 9y agoRe 3. "There's a skills shortage (at the price we want to pay)"
- technologia 9y agoI wish some of these companies would embrace having offices in places other than those with very high cost of living.
- antgoldbloom 9y agoYeh. The median salary for a machine learning engineer, which is definitely higher than what most companies are used to paying (even for software eng roles). My argument is that machine learning also higher leverage than most roles. One algorithm written by one machine learner can generate a huge ROI. Think of an algorithm to predict loan defaults or customer churn for a bank. That algorithm in the hands of a great machine learner can generate a huge ROI.
- rndmwlk 9y ago>3. 40% of people survey spend >1-2 hours per week searching for another job. Surprising given all companies complain about the difficulty in finding data scientists/machine learners. Not too surprised about this. I think a lot of people who go into the field want to do more interesting things than what they find being used in the field. I don't think pay is necessarily the gap here, as others have pointed out, so much as interesting work (or at least the intersection of interesting work and pay).
- huac 9y ago> 1. more people learn data science and ML from MOOCs than university courses More people, out of the subset of people on Kaggle. Lot of selection bias there!
- apohn 9y ago>40% of people survey spend >1-2 hours per week searching for another job. Surprising given all companies complain about the difficulty in finding data scientists/machine learners. I've hired data scientists in the past. One thing I found is that a lot of interviewees want to talk about all the algorithms (e.g. Gradient Boosting) they've used and are not able to describe how they thought through the problems before they applied the algorithms. It's easier to find somebody who downloaded some mostly clean data, then copy/pasted some code than a person actually thinks through the quantification of a problem. There are a lot of buzzword artists out there. This is important because in a lot of organizations the business problems have not yet been quantified in a way that lends itself to getting meaningful and valid results from an algorithm. The Data Scientist has to be able to work with others to quantify a problem. Or at a minimum, recognize that there are issues with the current way the problem is quantified and think of ways to improve it. It's much easier to teach somebody to run a data algorithm than it is to actually understand a business problem. There are issues with people doing the hiring as well. In my last job (not a software company), the VP of the group had pushed to get headcount for a data science team and was fearful of making the wrong hire because he didn't want to say "We hired a data scientist at 2X-3X the cost of a Business Analyst and that was a bad hire." The end result was a massive amount of paralysis, an insanely long and convoluted job description, and complaints about the hiring pipeline.
- kmax12 9y agoIn "What barriers are faced at work?", I really wish they broke down the "dirty data" response into more categories. In particular, I'd love to know if people are dealing with data quality issues, feature engineering issues, or something else all together. In my opinion, this is representative of the problems with data science tools today. There is so much focus on the machine learning algorithms rather than getting data ready for the algorithms. While there is a question that lets respondents pick which of 15 different modeling algorithms they use, there's nothing that talks about what technologies people use to deal with "dirty data", which is agreed to be the biggest challenge for data scientists. I think more formal study of data preparation and feature engineering is too frequently ignored in the industry.
- VHRanger 9y ago> There is so much focus on the machine learning algorithms rather than getting data ready for the algorithms. Generally, once a problem at work has come to the point of being a "kaggle problem", it's trivially easy. The main problem is unstructured data, with infinite ways of specifying similar ways to measure the same attribute, and lots of leeway to build an unmaintainable data pipeline between the data generation process and the model at the end.
- sidlls 9y agoI disagree that a "kaggle problem" style problem is trivially easy, but I strongly agree with the sentiment that dealing with unstructured data is often a much bigger, deeper, and broader problem than the choice of a particular algorithm or ensemble of them. The ability to efficiently and effectively derive insights from such data is scarce.
- VHRanger 9y agoRight, by "kaggle problem" I mean the general case where we roughly know what we're going to want to have on the right hand side of the model we're going to run (plus or minus some feature engineering, model choice and other hyperparameter specification, etc.)
- followmeon 9y agoNext it'd be interesting to see Python 2k vs. Python 3+. My own experience tells me that the majority of top Kagglers still use Python 2k, despite Kaggle Kernels being Python 3+ exclusively. I also am quite amazed with the predominant use of Logistic Regression. I wonder if that is less about interpretability / ease of engineering, and more about the barriers that data scientists face when using more complex methods: lack of data science talent, lack of management support, results not used by decision makers, limitations of tools. If Kaggle results are anything to go by, all businesses that care about best performance on structured data, should be using a form of gradient boosting.
- godelski 9y agoAnyone else find it weird that when you click "other" for gender that the data looks more like garbage? I was trying to actually compare male and female salaries out of interest but have a hard time believing so many people earn <$20k/yr. Even when you switch the filters around. The best I could find is just filtering for the US, but the number of respondents are so low, ~1k total (~200 Females, ~800 males), that it becomes difficult to make accurate comparisons ($22k diff but women had more masters degrees and similar PhDs, by percentage). Has anyone sorted through this data and tried to account for these factors? I'd be interested at the uncertainty and how the information was gathered.
- jerednel 9y agoI pulled a few gender stats here. http://bit.ly/2zjrSJD http://bit.ly/2zjrSJD Accounting for country, education, and industry you really reduce the population you're sampling from but those deviations are huge. You need to account for industry especially.
- godelski 9y agoWell this really doesn't discuss the error associated with the data. Which is what I was trying to get at. There seems to be a lot associated with it, which makes accurate predictions difficult to make.
- denfromufa 9y agoThe results are income do not reflect the location, which is really important in US.
- glial 9y agoWith the rise of Tensorflow and sklearn, the strong Python showing makes sense. However, I wish Python had a solid IDE for interactive work like RStudio. Jupyter notebooks are fine but being able to easily inspect variables is super convenient. Spyder doesn't cut it. Y-hat's Rodeo was still a bit buggy last time I tried it. Any other suggestions?
- reallymental 9y agotried pyCharm?
- jinonoel 9y agoInteresting that the most common models being used are the simpler ones, logistic regression and decision trees. This is despite all the hype for the more complicated techniques like neural nets and GBMs. Is it just because these models are faster to train and easier to interpret or something else?
- knn 9y agoin my experience, doing deep learning is a lot harder than building simpler ml models. training times are killer, need lots of data, overfitting is a challenge, hard to interpret results, lots of things can go wrong. deep learning is the future from a mathematical standpoint (with neural nets you can essentially learn arbitrary functions in some borel space or something whereas simpler ml models are basically a special case of deep learning) but it's definitely harder.