6 ms·
People who just use library functions without having any understanding of how they work are not going to do as well as the people who actually understand the ma
by Hotple 6y ago
People who just use library functions without having any understanding of how they work are not going to do as well as the people who actually understand the math. The people who have real understanding will always be called in to figure out the hard stuff.
- dpflan 6y agoCompletely agree, sure you can make superficial progress, but then you will need to further improve, and without deep knowledge of the underlying mechanics you will be lost.
- datameta 6y agoAlso rather important is to know what paths not to pursue in the search for a solution. Basically a solid understand of bias/variance will carry you far and help you avoid doing a week of work for an optimization that doesn't actually pan out.
- omgwtfbbq 6y agoAre you talking specifically about research? Because in my experience the time and $ delta needed to bring a classification model from say 93-96% accuracy is not worth it to most businesses. So all your special "deep knowledge" is irrelevant in most use cases
- dpflan 6y agoIf you're in research I would assume a deeper background or at least an environment where you are encouraged to develop your background. In industry, yes, time to value is important, but it seems clear that in some part of the market competition will require squeezing out those remaining % points of performance -- a company scaling up with poorer accuracy would have more problems...right? Do you have experience in industry? I would like to hear more from your perspective?
- omgwtfbbq 6y ago>In industry, yes, time to value is important, but it seems clear that in some part of the market competition will require squeezing out those remaining % points of performance -- a company scaling up with poorer accuracy would have more problems...right? But I wasn't talking about a poor model vs a good model. Like I said, going from 93-96% accuracy is generally not going to have a lot of value add. >Do you have experience in industry? I would like to hear more from your perspective? Yes and in my experience if your model has lift over random and has a positive roi vs doing nothing it's usually going to be worth building and implementing it. In the fields I've worked in the models are not in direct competition with models from other companies so if Co. A's model for some task is getting 95% accuracy and Co. B's is getting 91% there's not much cause for concern. If Co. B's model is still generating lift over a random guesser it's worth having in production. For most consumers I doubt they could even tell the difference between service like that. If Walmart's product recommendation is 2% better than Amazon's I find it extremely unlikely most consumers could even tell and Amazon's primary concern is whether the model is driving increased sales and not whether it's stealing customers from Walmart(the model, not generally).
- beforeolives 6y agoThis is true but I don't like that the math understanding is often treated like some magic secret that only a select few have access to. As long as you have a decent foundation you can research stuff and get deeper into the theory as necessary. And this inevitably happens when you switch problem domains even if you've gone really deep in a certain type of algorithms. It's not too different from a software developer figuring out a new API/system/technique etc. but for some reason the attitude is that you either know it or you don't. And in the context of this article, I think that the point is that even the deeper understanding isn't that valuable to many organisations.
- mustafa_pasi 6y agoWhat math, anyway? Knowing what a Gaussian is? There's really surprisingly very few math in the whole field and people even refuse to put things in mathematical terms if at all possible.
- bckr 6y agoAgreed. It's really more like playing with Lego. Take residual connections, for instance. The insight was that information wasn't traveling far enough when the networks got too deep. So they just.... Plugged the earlier layers into the later layers. And this has been a very important development. Or things like batch norm. We don't know why it's important. People do math to try to explain what's going on, not so much to figure out where we should go next. Related: Understanding is a poor substitute for convexity https://www.edge.org/conversation/nassim_nicholas_taleb-understanding-is-a-poor-substitute-for-convexity-antifragility https://www.edge.org/conversation/nassim_nicholas_taleb-unde...
- igorkraw 6y agoPosts like this really piss me off, because you can make anything sound small. So don't take the next few paragraphs as me attacking you, I'm just venting a general sentiment I've had for a while. (Looking back after typing it out, I might actually flesh it out and post it on my blog, I hopeit's somewhat thought stimulating for others as well). It's like someone saying "well, electrical engineering is really like playing with building blocks. Take zener diodes for example, the problem was that you can have a lot of power in a circuit, but if there's a power spike it might break. So they just...plugged in a piece that breaks by shunting the power spike into the ground, then reset. And this is now a major piece of electronics." Or "so, everyone is always going off about arabian mathematicians, but one "big development" they did was to invent the zero - just make up a symbol where previously you'd leave a space. It's basically just a change of notation!". Deep learning theory (statistical, information theoretical and optimisation wise) is our process of understanding how to design systems that adapt themselves to feedback, and how to encode tasks in them. Batchnorm was inspired by one thing (internal covariate shift), and that thing was plausible, but as it turns out, in systems as complex (not complicated, complex as in interactions) as universal function approximators, adding one thing can radically change things. As it turns out, batchnorm smooths the function, it decouples parameter magnitude and directions and it positively improves signal propagation. How else would you have figured this out without having systems like neural networks with batchnorm already in place that you can study? And now there are lines of work emerging that do away with batchnorm, but have distilled the positive properties into smaller techniques (https://arxiv.org/pdf/2102.06171.pdf https://arxiv.org/pdf/2102.06171.pdf, Soham De gave a lecture at our lab recently). Same thing with skip connections: Jürgen Schmidhuber will rightfully point out we've had highway networks since his heyday, but details matter. It is really not intuitive before you do it that in such a complex system, skip connections will be beneficial, because before we had them and started studying them on complex system, the ideas of thinking of them as learning small adjustments to a signal, or as an ensemble of shallow learners or the other perspectives that they have been studied under had not been developed. And how would you? Without having them working really well, you'd have to start thinking about them from first principles in the giant design space of nonlinear, nonconvex functions, without being able to prove anything because we don't have the mathematical formalism yet. Deep learning theory and nonconvex optimisation right now is a new physics born out of the marriage of information theory, computer science and computer engineering (and not surprisingly in a menage a trois, a lot of groundwork was laid by the french and other weird europeans /joke). We have a bunch of theory nerds trying to explain what we see in elegant and concise mathematical frameworks and trying to come up with testable predictions, and a bunch of experimentational people actually coming up with ways of testing it, gluing together the bits of understanding we have with soft knowledge to make the learning engine go brr and give feedback to the theorists on what held up, what didn't work predictable and what didn't go according to predictions. And people mouth off about the empirical nature of things. Well, I ask: How else would you figure this stuff out? I think there is a cult of genius at play here, where if you don't start with category theory and platonic ideal conceptions of reality and derive your model without any experiment, you are somehow lesser. Well, despite what people like to sell, disruption is a lie, everything is incremental, and without having the hackers make things work in clunky ways, the theorists would circle jerk themselves in creative dead ends because of a lack of stimulus. And as always, there are a lot of people who make themselves sounds smarter by affecting superiority and disdain on this scientific process, while in the background nerds deepen our understanding of the universe.
- jghn 6y agoI interpret TFA as saying that the space of "the hard stuff" is shrinking over time. So while I agree with you, if the space of the hard stuff is shrinking the demand for people who can figure it out will also be shrinking. A reduction in demand results in a lower valuation.
- commandlinefan 6y ago> the space of "the hard stuff" is shrinking over time That's theoretically true of programming in general - and has been true of programming since programming started. Network programming used to be super specialized, but now the standard for applications is distributed over the web. Graphics programming ability used to be rare, but now it's strange to see an app without a GUI. Yet programming takes longer to learn than it used to - precisely because the state of the art has advanced so much. I've worked with a handful of people who thought they could just use the libraries but didn't understand what they were doing or why they worked - people who could put together fancy UIs in, say, jQuery and such - and inevitably, they would find themselves hopelessly lost because they didn't actually understand what asynchronous callbacks were and couldn't figure out why, when they stepped through their programs with a debugger, the debugger kept "skipping over" their callback function.
- jghn 6y agoThis is roughly what leads me to agree w/ the GP. The people who have the training to understand AI at a fundamental level will have transferable skills that will give them a leg up on whatever the next "hard stuff" might be, even if it's not in the area we currently refer to as AI.
- spaetzleesser 6y agoI remember when the database guys were the high priests of software development and did their secret performance rituals in secret. Nowadays 99% of developers use databases without having a clue how they work and they are fine. I expect the same for AI. 99% of use cases will be commoditized and easily accessible for devs and only a very small number of people who understand it in depth will be needed. You already can do a lot of cool stuff by copying code with some tweaks and I see this trend only continuing.
- itsoktocry 6y ago>People who just use library functions without having any understanding of how they work are not going to do as well as the people who actually understand the math. ...who won't do as well as people who understand the business domain and that "good enough" isn't that hard to achieve with some pretty elementary stuff (regression, xgboost). PhD's have been trying to act as gatekeepers of "Data Science" for the past decade. It's only getting easier for people to apply this stuff. Unless you're doing actual research in these algorithms, there is little need to "understand the math" beyond an undergraduate level.
- usgroup 6y agoI think there's a category error here. The PhD and the tinkerer are not doing the same thing when engaged in the same problem.
- tomkat0789 6y ago+1 to this from someone who learned the math behind ML in a PhD and was looking forward to being a gatekeeper :) My favorite academic paper ever [0] was a comparison against a bunch of dimensionality reduction algorithms and 100 year old PCA was tough to beat! Glad I was able to pivot my career out of AI and ML. My PhD wasn't at Stanford, MIT, et al so I couldn't find any jobs doing the "actual research" - if they existed at all outside academia. EDIT to add another funny "frustration" paper more directly related to ML [1]. I consider DR is more of a data analysis thing. [0]: van der Maaten, et al. Dimensionality Reduction: A Comparative Review https://members.loria.fr/moberger/Enseignement/AVR/Exposes/TR_Dimensiereductie.pdf https://members.loria.fr/moberger/Enseignement/AVR/Exposes/T... [1]: Dacrema, et al. Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches https://arxiv.org/pdf/1907.06902.pdf https://arxiv.org/pdf/1907.06902.pdf
- platz 6y agoWhat is you pivot to?
- tomkat0789 6y ago
- fractionalhare 6y agoThe author's thesis is that the "hard stuff" decreases over time, and gets commoditized as the field matures. My experience aligns with this view. For example: you can theoretically do better than a seasonal ARIMA model for time series analysis. But in practice it's very difficult, and you probably don't have the amount of data you need or even an economic justification. The improvements will usually be marginal, expensive and not worth it. Most teams developing a new product or doing research - even within large and ostensibly sophisticated tech companies - won't usefully or economically outperform something like Box-Jenkins.
- amelius 6y agoDeep learning can be compared to brute-forcing. You can't base your career on brute-forcing because anybody can do it. And being smarter than the competition doesn't help because brute-forcing will solve all your problems anyway.
- deleted 6y ago[deleted]