4 ms·
I'm a data engineer for a startup that's trying to hire its first data scientist. The range of candidates that apply with this title is massive. Defining our ex
by red_hare 9y ago
I'm a data engineer for a startup that's trying to hire its first data scientist. The range of candidates that apply with this title is massive. Defining our expectations has been challenging.
My phone screen "fizzbuzz" is having them calculate a standard deviation from an array of data w/out with only basic operators (no numpy.std). Then explain why they choose population/sample and explain the difference.
I studied math in undergrad so one of my requirements is "knows more math than me".
- arca_vorago 9y agoI have to admit this scares me just a little bit. I'm a senior sysadmin who is trying to lateral transition into data science, but I'm no math whiz, I'm just good at pragmatic use of tech stacks and have a generally analytical mind. If you are a math undergrad how could I ever expect to know more math than you? Of course a standard deviation should be easy, but your comment on math just stuck out to me.
- shepardrtc 9y agoMaybe consider being a data engineer or a systems engineer? There's a pretty big demand for people that can set up, maintain, and assist the data scientists with the more complex tech stacks out there. In a former job as a systems engineer, I set up Hadoop clusters and helped manage data going into and out of it. And if you do decide to continue learning to become a data scientist, you'll already have a solid footing on the tech they actually use.
- arca_vorago 9y agoThat might be an option to learn from, but it's not my end goal. As a senior sysadmin who was working for and reporting to PHD execs, I saw directly how what was needed was someone to do the data science and then bring convincing results and reports to the execs, essentially distilling the knowledge and wisdom of what needed to be done. I really want to fill that disconnect. (eg one of my failures as a sysadmin was me focusing too much on the technical, and now I want to expand and play the business board room politics game, but with data science)
- sah2ed 9y agoI think it is a reasonable expectation to require a certain baseline of expertise, considering the first thing the poster admitted to was: "I'm a data engineer for a startup that's trying to hire its first data scientist." Much like how the early hires at Twitter were not deeply experienced in high availability work -- segregating the architecture of a predominantly RoR code base to be resilient at scale, which lead to countless "fail whale" outages, before they eventually landed someone who helped them re-think their architecture to use RoR for what it's good at while introducing the JVM and other languages to handle other aspects of their workload.
- dsacco 9y ago> If you are a math undergrad how could I ever expect to know more math than you? Read through, and do all the exercises in, one textbook each for: 1. Calculus 2. Linear Algebra 3. Abstract Algebra 4. Analysis 5. Topology 6. Probability Theory 7. Number Theory ...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics. Do that, and you have the equivalent of a mathematics undergrad (as far as relevant courses are concerned). You could even do this with something like UIllinois’ NetMath program, or some courses on Coursera. You can swap out Number Theory for Complex Analysis or deeper Probability Theory and it’d be more relevant.
- jeffreyrogers 9y agoThat doesn't seem like a good use of time. I've tried reading through and doing the exercises in an abstract algebra textbook. It's a lot of work and the applicability to real world problems is virtually non-existent. I think a more targeted approach would give you a better return on your time.
- dsacco 9y agoSure, I agree. Abstract algebra isn’t directly helpful. But: 1. The context is knowing more math than someone who has an undergraduate degree in it, 2. Abstract algebra is part of such a degree, and contributes significantly to overall mathematical maturity, and 3. You can avoid some subjects in the short term, but in the long term you can’t progress further without a reasonable mastery of algebra and analysis. Probability theory and linear algebra are heavily used in data science. You won’t be as competitive a candidate for a job if you don’t have a firm grasp of both subjects. At a certain point, linear algebra ceases to be distinct from abstract algebra, and those exercises you were doing become applicable to real world results.
- arca_vorago 9y agoThank you for this list, going on the todo, along with every other relevant comment on this thread. (emacs org mode is my ds notebook and todo app)
- 9y ago
- red_hare 9y agoHonestly, the market is so oversaturated with PHDs who are switching to DS I don't see how anyone can transition into it from a different role. I'm speaking for myself as well as someone with a math background, most people don't consider a math degree and 4 years of applying math models as a data engineering experience enough to be a "data scientist". They just weed out anyone without a PHD. But, what you're describing I would consider "data engineering" (at least how I have been hired to do it). Working through the business problems and pragmatically facilitating data, pipelines, databases, and models to solve those problems. It's less established and less "hot" but, IMO, it's a much more valuable job to most businesses.
- jerednel 9y agoAt least with fizzbuzz you are working through how to logically solve a problem. This is just regurgitating a formula. I don't see how this is helpful.
- shepardrtc 9y agoIts designed to quickly weed out people who don't know the underlying math, just as FizzBuzz is designed to quickly weed out people who don't know programming.
- Xcelerate 9y agoI work as a data scientist, and my graduate research involved harmonic analysis over compact groups, optimization over Riemannian manifolds, and loopy belief propagation. You'd reject me in an interview because I couldn't remember the formula for standard deviation off the top of my head?
- shepardrtc 9y agoPersonally, I don't ask weed-out questions. Never have, never will. I'm just saying that's what they're doing.
- red_hare 9y agoIt's not a hard weed-out for us. But, if you talk through variance for 10 minutes, you get pretty close to the formula for standard deviation. Also, for our role, we're specifically hiring someone with extensive stats background since a large part of the role is learning domain-specific statistics of the industry we're targeting and figuring out how we can adopt those models with our data.
- dsacco 9y agoNo, using numpy would be analogous to a simple formula. Doing it without numpy requires actually understanding what’s going on. It’s a filter that theoretically allows false positives (which is why you continue with other questions), but it really shouldn’t have any false negatives.
- dsacco 9y ago> I studied math in undergrad so one of my requirements is "knows more math than me". What kind of questions are you asking to ensure that they’re correct when they’re speaking about math you don’t know?
- red_hare 9y agoThis is a pretty hard problem I haven’t solved just yet. Generally, my in person interview is based on a set of DS problems I’ve been working on and had to do research myself to solve. What I look for is a strong intuition of the underlying math. Like, I can give them a formula and they can intuitively express what that means and explain it to me and then explain the next place they would take the solution. It’s not a perfect measure, but i’ve found the comfort with core concepts to be the most common trend among great data scientists I’ve worked with in the past.
- peatmoss 9y agoI think regurgitation of math formulas is a terrible way to hire for most data science positions. I've seen a breakdown of data scientists into two categories: 1) People who are great at the mathematics behind the statistical tooling 2) People who are great at conceptualizing a relevant question, operationalizing it, and then using a computer to apply appropriate models. I think in most cases, for businesses needing to solve business problems, the latter kind is probably more useful. There are applications where the former is required, but you probably know if you need this kind of data scientist. I should also add that these traits aren't mutually exclusive, but that individual data scientists typically are stronger or weaker along approximately those axes. In general, I still dislike the term "data science" because it obfuscates meaningful distinctions between math nerds, computer science nerds, and research nerds who happen to do some applied stats.
- red_hare 9y agoI actually agree with your breakdown. But, as a "data engineer" with a math background who's spent 5 years building analytics tools, I already identify as your type-2. We're working in a field that already has a rich history of established statistics that needs to be interpreted and broken down, so I think we're looking for someone who's a type-1. I do, however, think anyone with some lick of statistics background should know the formula for a standard deviation. Considering how fundamental the idea of variance is in statistics.
- Balgair 9y ago> ... of data w/out with only basic operators... (emphasis mine) I'm having a little trouble trying to parse that sentence. Could you explain it better? Based on what I think is being asked, the question is essentially: What is a STD? I think this is a very straightforward and fair question. For less Stat-y HNers: For normally distributed data, the STD is the root of the Variance. The Variance is just the average of the square of the difference between the data points to the mean. Essentially: Take a point, find the distance to the mean, square that, average over all points you've done that to. That's the variance. Root the variance, that's the STD.
- Xcelerate 9y ago> My phone screen "fizzbuzz" is having them calculate a standard deviation from an array of data w/out with only basic operators This is just my n=1 opinion, but this is a terrible test for data science skills. I've had to calculate standard deviation by hand many times in my life, but my short term memory is such that despite doing that dozens of times over the past two decades, I still can't recall the formula off the top of my head. And then there's the whole n vs (n-1) thing in the denominator which has something to do with degrees of freedom, but I would just Google that as soon as I needed to know (depending on exactly what I was trying to do with the data). So I don't understand how your question in any way tests someone's skills at analyzing data to extract valuable business insights. At best, it tests someone's ability to memorize formulas and minutiae (although I'll grant you that understanding the difference between a sample and the population is important). Personally, I think take-home interviews with real data sets are the best way to gauge a candidate's skills. You're actually testing them with a work sample, and they are not under artificial time or memorization constraints.
- charliej2 9y agoThis was useful info for a noob, thanks. It makes sense to me. If you understand SD in principle, you don't need to memorize the formula for a simple exercise like this, so it's a good filter for candidates who can demonstrate they do understand some basic math.