7 ms·
The ML/DS positions highly competitive these days. I don't get why ML positions requires hard preparations for the interviews more than other CS positions while
by mcemilg 5y ago
The ML/DS positions highly competitive these days. I don't get why ML positions requires hard preparations for the interviews more than other CS positions while you do similar things. People expect you to know a lot of theory from statistics, probability, algorithms to linear algebra. I am ok with knowing basic of these topics which are the foundations of ML and DL. But I don't get to ask eigenvectors and challenging algorithm problems in an ML Engineering position at the same while you already proof yourself with a Masters Degree and enough professional experience. I am not defending my PhD there. We will just build some DL models, maybe we will read some DL papers and maybe try to implement some of those. The theory is the only 10% of the job, rest is engineering, data cleaning etc. Honestly I am looking for the soft way to get back to Software Engineering.
- hintymad 5y agoA reason for such requirements is similar to that that software engineers need to leetcode hard: supply and demand. Prestigious companies get hundreds, if not thousands, of applications every day. The companies can afford looking for candidates who have raw talent, such as the capability of mastering many concepts and being able solve hard mathematical problems in a short time. Case in point, you may not need to use eigenvectors directly in the job, but the concept is so essential in linear algebra and I as a hiring manager would expect a candidate to explain and apply it in their sleep. That is, knowing eigenvector is an indirect filter to get people who are deeply geeky. Is it the best strategy for a company? That's up to discussion. I'm just explaining the motives behind such requirements.
- vsareto 5y agoI can’t help but think there’s been a ton of filters used in the past to figure out if someone is deeply geeky, and we’ll continue to invent more in the future. It’s really looking like another rat race. Especially since there’s no central authority, every hiring manager has the potential to invent their own filter, and make it arbitrarily harder or easier based on supply and demand (and then the filter drifts away from the intended purposes).
- hintymad 5y agoIt will be rat race when there are so many interview books and courses and websites. It was a not rat race before 2005, when there were only two reasons that one can solve problems like Pirate Coins or Queen Killing Infidel Husbands: the person is so mathematically mature that such problems are easy for them; the person is so geeky that they read Scientific American or Gardner's columns and remembered everything they read.
- littlestymaar 5y agoYou're missing the third category: people like myself who absolutely love this kind of riddles and destroy them in a few minutes, without any significance on their actual work abilities. I don't think I'm a bad engineer, but I'm certainly not the rock star you absolutely need for your team, but when it comes to this kind of “cleverness” tests, I'm really really good. I've had the “Queen Killing Infidel Husbands" (with another name) in an interview last year and I aced it in a few minutes, and I didn't knew about "Pirate Coins", but when I read your comment HN said your comment was "35 minutes ago" and now it says "40 minutes" which means I googled the problem, figured out the solution and then found the correction online to see if I was right in less than 6 minutes, and so while I'm putting my son to bed! It's really sad because there are many engineers much better at there job than me who will get rejected because of pointless tests like this…
- hintymad 5y agoThe Queen problem first showed up in a Putnam Math Contest. If you solved it in no time, then you're mathematically talented, which puts you in the first category.
- littlestymaar 5y agoI'm not questioning the fact that I'm kind of gifted when it comes to mathematics (I actually ranked #72 in a nation-wide math contest in France when I was 10) but you were talking about “maturity” and not innate skill. Since don't have a math degree and I haven't done math in more than a decade, I'm definitely far from “mature” on any mathematics perspective that can matter for a job. And after ten years working in the industry, I can assure you that it is not a skill I can leverage a lot in my job…
- MontyCarloHall 5y ago> Case in point, you may not need to use eigenvectors directly in the job, but the concept is so essential in linear algebra and I as a hiring manager would expect a candidate to explain and apply it in their sleep. Exactly. Whenever eigenvectors come up during interviews, it’s usually in the context of asking a candidate to explain how something elementary like principal components analysis works. If they claim on their CV to understand PCA, then they’d better understand what eigenvectors are. If not, it means they don’t actually know how PCA works, and the knowledge they profess on their CV is superficial at best. That said, if they don’t claim to know PCA or SVD or other analysis techniques requiring some (generalized) form of eigendecomposition, then I won’t ask them about eigenvectors. But given how fundamental these techniques are, this is rare.
- deleted 5y ago[deleted]
- Der_Einzige 5y agoGiven that PCA is heavily antiquated these days, I'd say that asking your candidates to know algebraic topology (the basis behind many much more effective non linear DR algorithms like UMAP) is far better... But in spite of the field having long ago advanced beyond PCA, you're still using it to gatekeep.
- JustFinishedBSG 5y ago> asking your candidates to know algebraic topology Congratulation, you've eliminated 99% of the ML research community.
- selimthegrim 5y agoThe initialization strategy for UMAP is important enough that asking about that in practice is probably more important than anything out of Ghrist's book as an interview question cf. https://twitter.com/hippopedoid/status/1356906342439669761 https://twitter.com/hippopedoid/status/1356906342439669761
- godelski 5y agoUMAP (and t-SNE) aren't the same as PCA. UMAP is pretty close to t-SNE and I think expanding PCA (Principle Component Analysis) and t-SNE (teacher Stochastic Neighbor Embedding) explain the difference. Neighbor embedding is a visualization technique and not the same as determining principle components. PCA preserves global properties while t-SNE and UMAP don't. They are good techniques for _visual_ dimensional reduction, but they aren't going to tell you the dominant eigenvectors of the data, or _dimensional reduction_. This is a bit of a pet peeve of mine. There's some more in this SE post https://stats.stackexchange.com/questions/238538/are-there-cases-where-pca-is-more-suitable-than-t-sne https://stats.stackexchange.com/questions/238538/are-there-c...
- pasquinelli 5y agoand yet we're also told that tech companies can't get enough people.
- nerdponx 5y agoMaybe "eigenvectors" is a bad example, because it's a pretty foundational linear algebra concept. But there is a threshold where it stops being a test of foundational knowledge and starts being a test of arbitrary trivia, and favors who has the most free time to study and memorize said trivia.
- whimsicalism 5y agoHaving recently completed an MLE interview loop successfully at a top company, I'm wondering where you are getting asked complicated linear algebra questions in interview?
- fault1 5y agoHopefully you aren't equating "eigenvectors" to "complicated linear algebra question". But I agree, a lot of MLE roles don't get asked such things. I think the OP's guide is closer to interviews I've seen for phd programs.
- whimsicalism 5y ago> Hopefully you aren't equating "eigenvectors" to "complicated linear algebra question". They explicitly say something harder than eigenvectors in the GP. I was imagining something involving the spectral theorem or something like that, ie. beyond the most basic linear algebra. OPs guide seems to cover plenty of things I'd expect someone to learn in undergrad, I think I touched on almost all of this - except for stuff involving jax and recent CNN architectures, both of which can easily be supplemented online.
- uoaei 5y agoThe difference between trivia and meaty knowledge is somewhat contextually dependent, but an understanding of how core probability and statistics concepts are integrated into the framework of machine learning by means of linear algebra and the other analytical tools is pretty damn useful to have substantive conversations about ML design decisions. Helps when everyone in the team speaks that language to keep up the momentum.
- uoaei 5y agoIn part because ML fails silently by design. Even if the code runs flawlessly with no errors, the outputs could be completely bunk, useless, or even harmful, and you won't have any idea if that is true just from watching The Number go down during training. It's not enough to know how to build it but also how it works. It's the difference between designing the JWST and assembling it.
- minimaxir 5y ago> In part because ML fails silently by design. That's why there's so much iteration and feedback gathering (e.g. A/B tests) as a part of DS/ML, which incidentally is rarely a part of the interview loop. Anyone who claims they can get a good model the first time they train it is dangerously optimistic. Even the "how it works" aspect has become more and more marginal due to black boxing.
- borroka 5y agoBut the OP was asking something different, that is why someone should excessively focus on theory, when, by the way, DL theory is very far from being solid and trial and error in ML and AI is the common way of operating. The "model is in place, but I have no clue what's doing and so it can fail without me understanding when and how is straw-man". Especially for supervised learning, that is, we have a label for data, it is immediately clear whether the output of the model is "bunk, useless, or even harmful". There is no "fail silently by design". I have been working in the field for almost 20 years in academia and in industry and it is not that I starting every PCA thinking about eigenvectors and eigenvalues and if you ask me now without preparing what are those, I would be between approximately right and wrong. But I fit many, many very accurate models.
- uoaei 5y agoYou are considering only the technical aspects of the model. While of course important to understand, those are less interesting when considering potential harms than the downstream effects of the inference pipeline, particularly when it comes to interpretations of outputs. What is absolutely the worst possible MO is to offload the interpretation portion of a pipeline to a machine using proxy metrics without an exceptional model which justifies the approach unequivocally. For instance, if we put an MSE loss function on a classification NN with sigmoid outputs, and used a classification dataset, we could generate an entire zoo of "many, many very accurate models" as measured by MSE. But once your model returns outputs, how do you interpret them to predict a label for some input data? You could hack some algorithm together (eg argmax of the highest value) which is indistinguishable from the "correct" procedure but the described probabilities are so incorrect that no ML professional would be comfortable trusting anything it says, not least because of the violation of the condition that the probabilities are non-negative and sum to one. But being able to explain why we use MSE or cross-entropy or any other loss function and which output activations (hint: and probability distributions) they are typically associated with actually has a very deep origin in the foundations of probability theory which blows open a whole new way of thinking about statistical modelling that is not made available in any of the programs whose materials I've been exposed to.
- barry-cotter 5y ago> But I don't get to ask eigenvectors and challenging algorithm problems in an ML Engineering position at the same while you already proof yourself with a Masters Degree and enough professional experience. People know pity passes exist for Master's degrees. You can't trust that someone actually knows what they should know just because they have a degree. Ditto professional experience. The entire reason FizzBuzz exists is because people with years of profesional experience can't program.
- vanusa 5y agoWe aren't talking about FizzBuzz here; but rather the fashionable practice of subject people to 4-6 hours of grilling on "medium-to-hard" problems that you absolutely cannot fail, or even be slightly halting in your delivery on. And which can only be effectively prepared for by investing substantial amounts of time on by-the-book cramming. On top of the fact that these problems are often poorly selected, poorly communicated, conducted under completely unrealistic time pressure, often as pile-ons (with 3-4 strangers as if just to add pressure and distraction), and (these days) over video conferencing (so you have to stare in the camera and pretend to make eye contact with people while supposedly thinking about your problem, on top of shitty acoustics), etc, etc. It's just fucking ridiculous.
- vidarh 5y agoI'm quite happy these places makes it so clear they're not places I would be happy to work. I always ask about the interview process and tell the recruiters I'm not interested if they expect really lengthy processes. I'm fine with things dragging out of they have additional questions after initial interviews, but not if their default starting position is that they need that.
- devoutsalsa 5y agoI figure the best way to prepare for an ML job is to pull out the nastiest working rat’s nest of if statements you’ve ever written & claim it was autogenerated by an adversarial network (which was you fighting with your coworkers over your spaghetti code).
- flubflub 5y agoThis really made me laugh, thanks.