13 ms·
Ask HN: What maths are critical to pursuing ML/AI?
What maths must be understood to enable pursuit of either of the above fields? are there any seminal texts/courses/content which should be consumed before starting?
- CuriouslyC 9y agoYou absolutely need a solid grounding in multi-variable calculus, linear algebra, probability theory and information theory. It will also be helpful to be well versed in graph theory. In my opinion one of the best starting points is "Information Theory, Inference and Learning Algorithms" by David MacKaye. It's a bit long in the tooth now, but it is still one of the most approachable and well written books in the field. Another old book that stands up very well is "Probability Theory: the Logic of Science" by E. T. Jaynes. "Elements of Statistical Learning" by Tibshirani is also good. "Bayesian Data Analysis" by Andrew Gelman is another great read. "Deep Learning" by Ian Goodfellow and Yoshua Bengio is useful for getting caught up with recent advances in that field.
- moocowtruck 9y agoyou mean i cant just bang out some ipythons and the matrix forms around me? thanks for the list! the only roadblock i've ran into getting into many of these topics are book prices :O usually they are pretty steep
- CuriouslyC 9y agoActually, the MacKaye, Jaynes and Goodfellow books are available for free online. Enjoy!
- popcorncolonel 9y agoI disagree that you need a solid founding in information theory. Almost all that I've seen about IT in ML is minimizing the KL divergence, which can be learned by browsing the wiki page.
- CuriouslyC 9y agoInformation theory is pretty central to model selection.
- srean 9y agoIt depends. All that is essential for an autombile engineer is not essential for a taxi driver.
- sgt101 9y agoMaybe more all that is essential for a molecular biologist isn't necessary for a general practitioner? It's just... those conference calls where you're explaining that because the classifier is working really well now doesn't mean that we can use it in production, those calls can get difficult and annoying, and sometimes the "other side" wins - with predictable results. ha ha ha!
- srean 9y agoYou bring up a very important point and a difficult one which is, if the decision making is in the hands of someone who does not understand the nuances too well nor has the time or inclination, what do you do ? If your salary is going to depend on how many models you pushed out and not how well they continued to perform, many will optimize over the number of models pushed out. A major source of problem (and sometimes a gift) is that you cannot prove a empirical statistical claim true or false in finite time. There is always this non-zero probability that the weirdest thing would happen. It could be just sheer bad luck that the model did so poorly in this cycle.
- eli_gottlieb 9y agoThat's not because you need little background in information theory. That's because KL-divergences are such a universal info-theoretic quantity that if you deeply understand them, you understand much to most of information theory. This is like saying, "You don't need to really know calculus, just integrals."
- jules 9y ago
- mayank 9y agoYou can actually get the latest edition of Elements of Stastical Learning for free as a (legal) pdf from the author! https://web.stanford.edu/~hastie/ElemStatLearn/ https://web.stanford.edu/~hastie/ElemStatLearn/
- cf 9y agoI disagree about the graph theory as well. Unless you are doing things with learning on networks you won't need it. I think a solid background in linear algebra, multivariate calculus, and convex optimization will take you really far.
- CuriouslyC 9y agoA lot of data is best represented graphically, and while you can shoehorn this sort of data into a vector space by projecting using a graph distance metric, the results are likely to be inferior.
- cf 9y agoI agree a lot of data can be represented graphically, but if you look at the literature it mostly is getting shoehorned into vector spaces. This doesn't mean people shouldn't learn about Graph Laplacians and friends, but I don't think it's an entry requirement.
- tptacek 9y agoI'm not super interested in ML but I am very interested in applied mathematics in computer science. I've got a fair bit of linear algebra due to cryptography, but have had virtually no need of any form of calculus (unless I'm relying on it without knowing it) in my career. So beyond just saying that you'd need grounding in multivariable calculus to do serious ML work, I would be super interested in hearing more about why that is and what kinds of problems crop up in ML that demand it.
- eli_gottlieb 9y agoEh, optimization and the occasional bits of analysis show up in ML more often than traditional vector-calculus.
- srean 9y agoAlmost every corner of an ML problem has an optimization problem that needs to be solved: There is a function that you want to minimize subject to constraints. Typically these are everywhere smooth, or sometimes almost everywhere smooth. So calculus shows up in (i) algorithms to find the bottom of these functions (if they exist) or (ii) deriving the location of the minima in closed form. These functions would be "how close am I to the correct parameter", "What losses would these settings rake up on average" etc etc. The reason why this differs from a purely optimization / mathematical programming problem is that we can only approximately evaluate the actual function (the performance of our model on new / unseen data) that we care to optimize. Great optimization algorithms need not be (and often are not) good ML algorithms. In ML we have to optimize a function that's getting revealed to us slowly, one datapoint at a time. The true function typically involves a continuum of datapoints. This is where we can bring probability into the picture (another option is to treat it as an adverserial game with nature). In the probabilistic approach, we make the assumption that functions being revealed to us is in some probabilistic proximity of the true function and the sample is closing onto it slowly. We have to be careful to be not too eager to model the revealed function, our goal is to optimize the function where these revealed functions are ultimately headed. Those things aside, if you have to choose just one prereq, I think it has to be linear algebra and you already have that in your bag. Without it, a lot of multivariate calculus will not make much sense anyway. Then one can push things a little bit and go for the linear algebra where your vectors have infinte dimension. This becomes important because often your data would have far too much information that you can encode in a finite dimensional vector. Thankfully a lot of intution carries over to infinite dimension (except when it does not). This goes by the name functional analysis. Not absolutely essential, but then lack of intution here can rein you in from doing some certain kinds of work. You will just get a better (at times spatial or geometric) understanding of the picture, etc etc. Other than theeir motivating narratives, there is not much difference btween probability/stats and information theory. There is a one to one mapping between many if not all of their core problems. A lot of this applies to signal processing too. Many of the problems that we are stuck at in these domains are the same. Sometimes a problem seems better motivated in one narrative over the other. Some will call it finding the best code for the source, others will call it parameter estimation, yet others will call it learning. Or If I may paraphrase for the CS audience, blame the reals \mathbb{R}. Otherwise it would have been the problem of reverse engineering a noisy Turing machine that we can access only through its input and output. Pretty damn hard even if we dont get into reals. In those situations you could potentially get by without calculus, algebra by itslef should go a long way, but as I said it gets frigging hard. Learning even the lowly regular expression from examples is hard. Calculus would still be helpful because many combinatorial / counting prolems that come up can be dealt with generating function techniques where you would run into integral calculus with complex numbers.
- digitalzombie 9y ago> "Bayesian Data Analysis" by Andrew Gelman is another great read. If you want to read that book you need real analysis more specifically measure theory (unless that subject is in probability theory for you). You cannot get into the last few chapters without it. Dirichlet Process are described using measures. I don't believe you need multivar calc or info theory. Info theory stuff are used but not as often. I believe you're slanted toward researcher phd position. Gini index, entropy, etc... and such are taken as given when needed.
- dtjon 9y agoGreat class and great professor. One of my favorite classes from my degree.
- jmh530 9y agoMy recollection is that you need neither real analysis nor measure theory to appreciate it, but it's been a while since I read it. You might get more out of it if you have studied those. I disagree on multivar calc. Statistics often makes use of matrix derivatives. I have found it helpful to know.
- kgwgk 9y agoYou don't really need measure theory. It's true that the last chapter in the book (in the 3rd edition) uses measure theory, but it's the only one. http://andrewgelman.com/2017/08/02/seemingly-intuitive-low-math-intros-bayes-never-seem-deliver-hoped/#comment-537801 http://andrewgelman.com/2017/08/02/seemingly-intuitive-low-m...
- mindcrime 9y agoWhat's required as a prereq to Measure Theory? Any suggestions on good resources for learning Measure Theory? I have a vague notion that Probability and Measure Theory are intertwined / related somehow, but have never studied the latter specifically.
- catnaroek 9y agoI'm taking a measure theory course right now, and we primarily use some set theory and some topology of R^n.
- dtjon 9y ago+1 for "Elements of Statistical Learning", this is the basis for most rigorous intro to ML classes
- septimus111 9y agoThese are all brilliant books, but I feel like anyone who is ready for them wouldn't need to be asking this question.
- misiti3780 9y agoProbability Theory: the Logic of Science is mindblowing, not a page turner, but if you can digest it is is very good.
- rdudekul 9y agoFree PDFs of some of the books mentioned: "Information Theory, Inference and Learning Algorithms" by David MacKaye http://www.inference.org.uk/itprnn/book.pdf http://www.inference.org.uk/itprnn/book.pdf "Probability Theory: the Logic of Science" by E. T. Jaynes http://www.med.mcgill.ca/epidemiology/hanley/bios601/GaussianModel/JaynesProbabilityTheory.pdf http://www.med.mcgill.ca/epidemiology/hanley/bios601/Gaussia... "Elements of Statistical Learning" by Tibshirani https://web.stanford.edu/~hastie/Papers/ESLII.pdf https://web.stanford.edu/~hastie/Papers/ESLII.pdf "Bayesian Data Analysis" by Andrew Gelman http://hbanaszak.mjr.uw.edu.pl/TempTxt/(Chapman%20&%20Hall_CRC%20Texts%20in%20Statistical%20Science)%20Andrew%20Gelman,%20John%20B.%20Carlin,%20Hal%20S.%20Stern,%20David%20B.%20Dunson,%20Aki%20Vehtari,%20Donald%20B.%20Rubin-Bayesian%20Data%20Analysis-Chapman%20and%20Hall_CRC%20(2014).pdf http://hbanaszak.mjr.uw.edu.pl/TempTxt/(Chapman%20&%20Hall_C...
- kgwgk 9y agoNote that only MacKay (that’s the correct spelling) and Hastie/Tibshirani/Friedman are legally available online. edit: Goodfellow/Bengio/Courville, not mentioned in the previous comment, is also available online: http://www.deeplearningbook.org http://www.deeplearningbook.org
- deleted 9y ago[deleted]
- capkutay 9y agoDo you have any good guides for the calculus required to do ML? Is it just the basic Calc AB from high school?
- VT_Drew 9y agoGame Theory would probably be more valuable to understand than graph theory. Just my 2 cents.
- rocqua 9y agoFor calculus, I'd skip the more physics like finding of integrals and derivatives. What matters is understanding the concepts of integrals and derivatives, and knowing properties like the chain rule. It pays much less to know that the integral of 1/x is ln(x) (or the other way round). The linear algebra and probability theory are most important imho. I'd also distinguish between probability theory and statistics. Both are important, but they are distinct disciplines.
- rhaps0dy 9y agoWhat about "Pattern Recognition and Machine Learning" by Christopher Bishop?
- mendeza 9y agoWhat about the Pattern Recognition book by Bishop? I am reading it now and its more approachable than the Elements of Statistical Learning book
- adamnemecek 9y agoStanford EE263 is very spicy http://ee263.stanford.edu http://ee263.stanford.edu
- e19293001 9y agoYou can learn the required maths along the way through Andrew Ng's deep learning course at coursera.
- mindcrime 9y agoIt depends on how deep you want to go and what your goals are, but I'd say that CuriouslyC pretty much nailed it. Multi-variable calculus, linear algebra, and probability / stats are definitely the core. If you're interested in finding more "freely available online" maths references, check out: http://people.math.gatech.edu/~cain/textbooks/onlinebooks.html http://people.math.gatech.edu/~cain/textbooks/onlinebooks.ht... http://www.openculture.com/free-math-textbooks http://www.openculture.com/free-math-textbooks https://open.umn.edu/opentextbooks/SearchResults.aspx?subjectAreaId=7 https://open.umn.edu/opentextbooks/SearchResults.aspx?subjec... https://ocw.mit.edu/courses/online-textbooks/#mathematics https://ocw.mit.edu/courses/online-textbooks/#mathematics https://aimath.org/textbooks/approved-textbooks/ https://aimath.org/textbooks/approved-textbooks/ There's also a TON of high-quality maths instructional content on Youtube, Videolectures.net, etc. For example, there's some really good stuff by David McKay (also mentioned in CuriouslyC's post) here: http://videolectures.net/david_mackay/ http://videolectures.net/david_mackay/ Be sure to check out Professor Leonard: https://www.youtube.com/user/professorleonard57 https://www.youtube.com/user/professorleonard57 Gilbert Strang: https://www.youtube.com/results?search_query=gilbert+strang https://www.youtube.com/results?search_query=gilbert+strang and 3blue1brown: https://www.youtube.com/channel/UCYO_jab_esuFRV4b17AJtAw https://www.youtube.com/channel/UCYO_jab_esuFRV4b17AJtAw as well.
- torbjorn 9y ago3blue1brown is great. I also recommend Siraj Raval's Youtube course the Math of Intelligence: https://www.youtube.com/watch?v=xRJCOz3AfYY&list=PL2-dafEMk2A7mu0bSksCGMJEmeddU_H4D https://www.youtube.com/watch?v=xRJCOz3AfYY&list=PL2-dafEMk2...
- jonahx 9y agoAnother upvote for 3blue1brown. I just watched his linear algebra series and it's probably the most outstanding math instruction I've encountered.
- zimzim 9y agohttps://www.youtube.com/user/EugeneKhutoryansky https://www.youtube.com/user/EugeneKhutoryansky another nice yt channel about math and physics.
- amrrs 9y agoStatistics and Probability - For non-math background, Openintro.org with R and Sas lab is a good one. Khan academy videos on the same again makes a lot of concepts easier. http://www.r-bloggers.com/in-depth-introduction-to-machine-learning-in-15-hours-of-expert-videos/ http://www.r-bloggers.com/in-depth-introduction-to-machine-l... Introduction to Statistical Learning http://www-bcf.usc.edu/~gareth/ISL/ http://www-bcf.usc.edu/~gareth/ISL/ (Rob S and by Trevor H, Free I guess) for more in depth, Elements of Statistical Learning by the same. Linear Algebra (Andrew Ng's this part in Introduction to Machine Learning is a short and crisp one) If you're not scared by Derivatives, you can check them. But you can easily survive and even excel as a data scientist or ML practitioner with these.
- gtani 9y agoFor going all in, https://ocw.mit.edu/courses/mathematics/18-657-mathematics-of-machine-learning-fall-2015/syllabus/ https://ocw.mit.edu/courses/mathematics/18-657-mathematics-o... But a good number of people that are doing work haven't taken real analysis, or it's been awhile and so you should be current on multivariable and vector calculus. Calculus of variations shows up from time to time. For math reviews, look at the following (there's others if you want more refs, ping me): http://www.deeplearningbook.org/ http://www.deeplearningbook.org/ https://metacademy.org/roadmaps/ https://metacademy.org/roadmaps/ http://www.cs.huji.ac.il/~shais/UnderstandingMachineLearning/understanding-machine-learning-theory-algorithms.pdf http://www.cs.huji.ac.il/~shais/UnderstandingMachineLearning...
- pveierland 9y agoThe following is a concise and good explanation of necessary knowledge of information theory: http://colah.github.io/posts/2015-09-Visual-Information/ http://colah.github.io/posts/2015-09-Visual-Information/
- ivan_ah 9y agoProbability theory and linear algebra are pretty much the core. Learning LA will help you become comfortable with multi-dimensional quantities, vector spaces, and give you some powerful computational techniques, e.g., SVG==PCA. Here is a short tutorial on linear algebra: https://minireference.com/static/tutorials/linear_algebra_in_4_pages.pdf https://minireference.com/static/tutorials/linear_algebra_in... and a preview of the full book: https://minireference.com/static/excerpts/noBSguide2LA_preview.pdf https://minireference.com/static/excerpts/noBSguide2LA_previ...
- soniakross 9y agoYour partner for a credit fast, sure and adapted to your needs. Your partner on a daily basis.We are proud to help people from all walks of life, no matter what their financial situation might be. https://www.carroncredits.com/en/index.html https://www.carroncredits.com/en/index.html
- blt 9y agoIf you want to understand SVMs deeply, a course in convex optimization. In general, proving maximum likelihood estimation for a lot of classic machine learning models involves using the method of Lagrange multipliers. But not deep neural networks :)
- Myrmornis 9y agoIt depends whether you want to work more as an engineer / data analyst, or more as a "ML researcher". For the latter, then, yes, as everyone says below, you need to be totally comfortable with multivariable calculus, linear algebra, probability and statistics, numerical optimization etc. But many jobs are more practical in nature, in which the main case essential skill is, being able to run a bunch of different models with different parameter values and collect and interpret the results, efficiently and reproducibly, and be able to talk about them and make recommendations for the way forwards. In those jobs you're not actually going to need to be able to derive updates for backpropagation, even though it's certainly satisfying to understand it.
- mindcrime 9y agoYep. We have to keep in mind the distinction between "applied ML" and "ML research" (while realizing that this is a continuum, not a binary distinction). Not everybody is doing cutting edge original research... some people really can just get by with downloading DL4J, reading a few tutorials, and then applying a basic network to their problem, and create some value in the process. I think cars are a good analogy. In the early days of automobiles, you needed to be something just short of a mechanical engineer to keep one going for any length of time, and it was routine to need to carry around tools and spare parts to perform significant repairs. You really needed to know a pretty good bit about how the car worked to use it effectively. But over time cars developed better abstractions and became more dependable and it became possible to operate a car without caring one lick about how it works, beyond know that it needs gas (or electricity!) and taking it in for the occasional tuneup /tire change / alignment / etc. I wouldn't say we're at the point yet where ML afford one the opportunity to be completely divorced from caring about the underlying details, but I think we are at a point where you can legitimately get useful stuff done without needing to be able to, say, derive the equations for backprop by hand.
- charlescearl 9y agoMichael I. Jordan's suggested reading list has been posted here a few times https://news.ycombinator.com/item?id=1055389 https://news.ycombinator.com/item?id=1055389
- framebit 9y agoA thorough, intuitive grounding in statistics is crucial, IMO. Doing any kind of ML means questioning all the assumptions that go into your results and understanding how those assumptions could affect the outcome. That process starts in stats.
- irchans 9y agoBasic probability is very helpful: Expectation, Standard deviation, P(A and B) = P(A)*P(B) is A and B are independent, P(A or B) = P(A)+P(B) if A and B are mutually exclusive. Also, knowing algebra is very helpful. In a way, you don't really need to know much more because there is a lot of good software out there. If you want to learn more math, learn Linear Regression, Logistic Regression, p-values, probability density functions, cumulative density function, the Central Limit Theorem, Gaussian Distributions, Exponential Distributions, Binomial Distribution, (maybe) Student-T distribution. If you want to learn even more, first learn matrices (adding, multiplying, inverting, rank, span, matrix decomposition (SVD, and eigendecomposition are the most important)). If you want to learn even more, it's time to learn calculus. Integral calculus is needed for continuous probability distributions and information theory. Differential calculus is needed to understand back propagation. There are a lot of other good suggestions written by the other commentators.
- septimus111 9y agoThis is a great list of the main concepts to know.
- dtjon 9y agoProbability, and thus multivariate calculus and partial differential equations. Linear algebra. Convex Optimization, and thus multivariate and partial differential equations. Some principals of statistics is usually helpful
- jules 9y agoWhy do you need partial differential equations? I don't think you necessarily need any knowledge of differential equations to do ML, though the top ML people certainly would know it because of their general math education.
- septimus111 9y agoI spent a lot of time messing with PDEs as a student but sadly that knowledge hasn't been very useful - I've only seen them come up in quite specialised areas like optical flow...
- jules 9y ago* Calculus * Linear algebra * Optimisation * Probability Various universities have very good course content freely available online, often including textbook recommendations, course notes, exercises, sample exams, and video lectures. Realistically it is probably going to be quite difficult to learn this on your own.
- leecarraher 9y agoIt will depend on the level you plan to engage in the ML/AI space. If you just want a job in ML/AI , you are in luck. Due to the growing assortment of available, mostly to fully automated, solutions like Datarobot, H2O, sckit-learn, keras(w/ tensorflow) the only math you will absolutely 'need' is probably just Statistics. Regardless of what's going on behind the scenes with whatever automatically tuned and selected algorithm your chosen solutions uses, you will still need some stats in the end to show the brass that 'your' model works. the upside is that then you can spend time, learning feature extraction, data engineering, and the aforementioned toolkits, in particular what models they make available. If you want to develop new techniques and algorithms, the the skies the limit, you'll of course want Stats too though.
- rcarrigan87 9y agoCan you recommend a Stats course that would be most relevant for people trying to be more practitioners (not researchers)?
- samstave 9y agoPlease just recommend the best online stats course you know of as a general toolbelt-notch.
- mindcrime 9y agoThere's a series of courses on Coursera, part of a Specialization from Duke titled something like "Statistics and Probability with R" or something like that. I've taken the first few classes in that series and have found them pretty good. The class on Bayesian Statistics is a little more difficult, but not too bad. I'll just say that you might want to complement the class with another book or other references on Bayesian stats. I've used this book: https://www.amazon.com/Bayes-Rule-Tutorial-Introduction-Bayesian/dp/0956372848 https://www.amazon.com/Bayes-Rule-Tutorial-Introduction-Baye...
- leecarraher 9y agoThis https://www.amazon.com/Probability-Statistics-Engineers-Scientists-MyStatLab/dp/0134468910 https://www.amazon.com/Probability-Statistics-Engineers-Scie... is the newer version of the stats book i had in undergrad, But @anst makes a good point about scikit learn. there is alot of good math to learn just from the docs and you can then investigate further on wiki, quora, stackexchange. for the what's up in Data Science i like datatau.com. and there are some great podcasts too, like datascienceathome and partiallyderivative (there are lists).
- mtzet 9y agoHonestly, I don't think having to learn some stuff before starting anything is nessecary, especially for learning a field as wide as ML/AI. It's much better to start out trying to learning something you're interested in, and then trying to fill in the gaps. This will also help you understand and motivate the underlying theory you're reading. So for example, start with some source in ML/AI you'd like to read. If you get stuck, ask somewhere (possibly an online forum like this) what field you're having trouble with and how to get started there.
- pramalin 9y agoYou can also dive in first and then cover the math behind ML, by taking Andrew Ng's courses. https://www.coursera.org/learn/machine-learning https://www.coursera.org/learn/machine-learning https://www.coursera.org/specializations/deep-learning https://www.coursera.org/specializations/deep-learning
- septimus111 9y agoNot knowing anything about you, I'll assume that - you are starting with the equivalent of a high school level of maths - you want to take a ML course or read an ML book without feeling totally lost As some commenters have said, Calculus, Probability and Linear Algebra will be very helpful. Some people like to recommend the "best" or "most important" books which you "should" read, but there is a strong chance these will end up sitting on a bookshelf, barely touched. So I will recommend some books which are perhaps more accessible. - Calculus by Gilbert Strang - Linear Algebra by Gilbert Strang For Probability: I don't have any favourites, sorry.
- bluetwo 9y agoNot a mention so far about game theory or Nash equilibrium. I'm no expert but does anyone think these apply?
- mindcrime 9y agoNot a mention so far about game theory or Nash equilibrium. It depends on what you're doing. I was literally just watching a video on Generative Adversarial Networks this morning, and game theory did come up there, at least in passing. If one sat down and started reading the papers on this subject and trying to implement / improve stuff in this area, I suspect game theory would be at least moderately important. There is also the field of Competitive Learning where game theory has some application. See, for example: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.71.3023&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.71....
- bluetwo 9y agoInteresting. Thanks.
- srean 9y agoIt very much does. Boosting, a one of the best off the shelf ensemble classifier is derived from a game theoretic formulation. Besides that there is this huge body of literature about prediction under non-probabilistic sequence of test cases. This line of work is primarily held up by game theoretic arguments and that of online convex optimization.
- bluetwo 9y agoI wasn't familiar with Boosting. Now off the read some articles.
- WhitneyLand 9y agoSurprising level of disagreement here on a few items for a sub field that has its own degree tracks. Multiavariable calc you either "abolsutely" need or don't really need. Should be well versed in graph theory, or don't need it much. Surely some of the contradiction is caused by different assumptions of what the goal is. But some of its hard to relate to as a reader. For example, I haven't been in the field but but have tried to read enough to understand the concepts, and having studied graph theory I don't see how it's a top 5 recommendation. I don't doubt anyone's experience, would just be nice to know which assumption is behind a suggestion.
- PeterisP 9y agoTo apply known methods in cases where they mostly work, you don't need to know the math behind them, you just need to know basic stats and basic probability to interpret the results. So if the assumption is that you'll simply be solving your problems by applying the known methods using the (great!) tooling made by others, then you don't need the math background; you can certainly train undergrads to solve quite nifty problems with the powerful tools without going into much if any detail about the underlying math, treating it as an engineering problem of following best practices. After all, the choice of e.g. a particular gradient descent optimization algorithm is not based on their mathematical properties (the proven bounds are so far away from practical results, and a better proven bound doesn't correlate that much with having better results) but on empirical evaluation, and in most cases you're not going to implement any of the low-level structures/formulas on your own anyway, in practical solution development you're just going to choose them from a list by name in the framework of your choice. On the other hand, if the assumption is that your particular problem is not solvable easily and reliably with the current approaches, then quite a lot of the math background helps - if you want to improve on the current results, or debug/understand why your solution doesn't work as intended, or why the conceptual solution can't work on your problem because of incompatible assumptions, then these areas of math are useful. If you want to use a new bleeding-edge construct, or a rare niche construct that's not yet implemented in the framework of your choice, then you're going to need to write it yourself, and then you need to understand how it works. There's a large distance between using and applying ML techniques and researching and improving ML techniques; it's a continuum, but there's space for many people standing purely in the applied end.
- KirinDave 9y agoIf you care about actually reading the nournals, as I do, and you had a very poor math education (as mine was abysmally opposed to both math and science as enemies of religion) then here are things I've determined I need to know to read journals: - Core statistics. You need to be familiar with how statisticians treat data, because it comes up a lot. - Calculus. You do not need to be a wizard at working the numbers but you do need to understand how to describe the process of differentiation and integration over multiple variables comfortably. - Linear algebra. It's essentially the basis for everything, even more than statistics. - Numerical nethods for computing. I constantly have to refer to references to understand why people make the choices they do. - Theory of computation and the research clustered around it. Familiarity here helps a lot. Sometimes I even catch errors or am able to recognize improvements available. Also there is a lot of crossover, as one would expect. An example: everyone is remembering how good automatic differentiation is! And given that properly combined differentiable equations are also differentiable, AD let's you optimize over your optimization process. It's differentiable turtles all the way down. My next big challenge is nonparametric statistics. Many researchers tell me that this is a very fruitful place to be and many methods there are increasingly making improvements in ML.
- cynicaldevil 9y agoHow did you learn these topics? Did you solve problems for each of them?
- KirinDave 9y agoReading, study, and textbooks. Tbh, I'm not where I want to be with them. So maybe next year I can talk about 2017 and my math oddessy.
- rdrey 9y agoI previously attempted Andrew Ng's old course, but didn't complete the tutorials. Now I would start in this order: Watch the course.fast.ai lectures quickly, just to see a lot of practical ML/AI applications. You'll see how effective you can be just by knowing the tools with very little math background. Next I'd look at the NEW Andrew Ng introduction on Coursera. It is much more approachable than his first course. You might still feel a little overwhelmed by a few equations, but then you'll implement them yourself in numpy. (And the ipython/jupyter notebooks are really well written, walking you through every step.)
- mendeza 9y agoI wish the people who answer this question are people that are current deep learning engineers or data scientist that use deep learning in real world settings, I am worried that people who are not credible are giving advice, which is not valuable. I am a masters student taking a PhD class in Bayesian machine learning to figure this out as well. I hope to have a better answer for this by the end of the course!
- xenihn 9y ago>I wish the people who answer this question are people that are current deep learning engineers or data scientist that use deep learning in real world settings There simply aren't very many people in those roles because the number of ML/AI/DL jobs out there are still limited, I think.
- mindcrime 9y agoI wish the people who answer this question are people that are current deep learning engineers or data scientist that use deep learning in real world settings, Why do you want answers only from people doing deep learning? Deep learning is just a subset of the overall field (albeit an incredibly popular and useful one). Anyway, the simple solution is just to use some simple machine learning of your own to analyze the data set which these threads constitute. Look for patterns... are certain answers being repeated over and over again, by different posters? Then I'd argue that your Bayesian posterior for "this is legitimately important" should go up. Take Linear Algebra for example... given the sheer number of people saying "linear algebra" in their answers, it seems a reasonably bet to me that LA is really, truly useful. Either that or there's some really freaking group-think shit going on. :-)
- mendeza 9y agoI guess what I am looking for is advice from practitioners who wont lead people astray who are really interested in diving deep into ML. I have attempted to read the Statistical Learning book, and its so daunting because the book expects a lot of background knowledge, and it takes a while to really wrap your head around these concepts. I think people should learn from a lighter book, before diving into these books if you are lacking the background. My current approach to pursuing a career in DL and ML is going to graduate school, taking a graduate ML course, and trying to apply my knowledge to different problems I am interested in. I am reading the Bishop book Pattern Recognition now. I think from the perspective of having to re-learn a lot of calculus and probability, that book is more approachable than Statistical learning. My advice (which I am attempting now) to dive deep into ML is follows: 1. Taking Bayesian ML class (at Cornell) 2. Read/Study Pattern Recognition by Bishop, for 5hrs/day 3. Try exercises, if fail, review solutions 4. If lost(which is usually), review missing concepts from MIT OCW scholar courses
- proofofstake 9y ago> What maths must be understood to enable pursuit of either of the above fields? None. > Are there any seminal texts/courses/content which should be consumed before starting? No. You don't need to know binary to start being a programmer/developer either. Just start already. As long as you are not in charge of a medical diagnosis or financial model, you don't get any drawback in experimenting (and failing miserably). Assuming applied ML, the most difficult part will be the human-political business element of it: People not understanding your model or using its output correctly, bias, feedback loops, acquiring enough resources, etc. The more you can explain to them, without resorting to heavy maths, the better communicator you are. That said, it can't hurt to do Ng's Coursera course (a lot of top performers started out with this course). Learning from Data by Caltech's Abu-Mostafa goes very wide on machine learning. "Programming Collective Intelligence" is a, somewhat dated, good book. As for seminal texts, the field is too wide for this. A better bet is: Find a professor in the field you are interested in. Say "Deep Learning", you could have a look at LeCun, Hinton, Schmidhuber, Bengio, ... Now look at their PhD-students, their papers, their courses, their conference talks, their software, their current research. Basically become a student under the most authoritative professor in the subfield you can find and resonate with, without ever paying any university tuition or them knowing you exist. This is very possible these days. But by all means: Just start out. Machine learning is fun. Learning about dry 100 year old maths not so much. Make mistakes. Learn to detect and avoid overfit. Find out if you are passionate and curious about parts of the field, then the theory will come eventually. A lot of the time these questions seem to demand answers like: "You need a PhD-level understanding of mathematics" Just so your brain can go: "I am not good enough for this, so let's look at something easier". Don't use this as an excuse. Start making intelligent stuff. There are 16-year-olds on Kaggle routinely beating maths PhD's. Also remember that, despite the current trend of calling everything "AI", that AI is a very wide field, of which mathematics is only a small part. There is philosophy, linguistics, cognitive science, physics, neuroscience, psychology, computer science, robotics, logic, ... all these parts vary wildly in their prerequisite maths knowledge.
- graycat 9y agoPart I (1) Calculus Generally should have college freshman and sophomore calculus. (1.1) Functions So, there can understand better what a function is. E.g., function f(x) = 3x^2 + 1. (1.2) Derivatives Then will learn how to find the slope of the graph of a function. That is the derivative of the function. E.g., for function f with f(x) = 3x + 2, as in high school algebra, the slope is 3. Then for each x, the derivative of f at x is just 3. The derivative of function f is denoted by either of f'(x) = d/dx f(x) E.g., for function f(x) = 3x^2 + 1 it turns out that f'(x) = 6x. (1.3) Integration For function g(x) = 6 x maybe we want to know what function f(x) will give us f'(x) = g(x) Finding such a function f is anti-differentiation, that is, undoes differentiation. So, sure, f(x) = 3x^2 + C for any constant C. Such anti-differentiation is also the way to find the area under a curve. So, can use that to find the area of a circle, volume of a cylinder, etc. Doing that the anti-differentiation is integration. The fundamental theorem of calculus shows how differentiation and integration are related. (1.4) Analytic Geometry Commonly taught at the beginning of a calculus course is analytic geometry. So, take a cone an cut it. Then the cut surfaces will be one of a circle, an ellipse, a parabola, a hyperbola, or just two crossed straight lines. So, those curves are from a cone and are the conic sections. There is some simple associated algebra. Conic sections are important off and on; e.g., applied math is awash in circles; the planets move in ellipses; a baseball moves in a parabola or nearly so; an electron moving toward a negative charge will turn away from that charge in a hyperbola. It turns out that in linear algebra (below) circles and ellipses are important. (1.5) Role of Calculus Calculus was invented by Newton as part of working with force and acceleration for understanding the motion of the planets. E.g., if at time t function d(t) gives distance traveled, then function v(t) = d'(t) is the velocity at time t and function a(t) = v'(t) is the acceleration at time t. Then Newton's second law is F(t) = m a(t) where F(t) is the force at time t applied to mass m. Calculus is the first approach to the analysis of continuous change and is a pillar of civilization. Knowledge of calculus will commonly be assumed in work in ML/AL, data science, statistics, optimization, applied math, engineering, etc. E.g., a lot in ML, AI, and data science is getting best fits to data; best fitting is to minimize errors in the fit; such minimization is mostly a calculus problem; one of the main steps in ML is steepest descent, and that is from a derivative. Probability theory (e.g., evaluating coin tossing, poker hands, accuracy in ML) will be important in ML/AI, etc.; two of the basic notions in probability are cumulative distributions and density distributions; the cumulative is from an integration, and the density is from a differentiation.
- EternalData 9y agoSome people have had a more comprehensive view on this -- if I were to focus on one field of math to understand really well though, it'd be statistical reasoning and the understanding of probability and uncertainty.
- sumitgt 9y agoI won't really comment about ML/AI in general. But, if you specifically care about getting into Deep Learning, I would say only bother looking into: - Basic linear algebra and matrix algebra. Since you would rely on frameworks like Tensorflow to handle figuring out the derivatives for you, you don't really need to know much calculus. Just read up on what derivative of a function at a particular point signifies. This should give you enough intuition to understand things initially. A skill that would really come in useful would be ability to look at a function and think how increasing/decreasing one of the variables would affect it's value. This would help develop intuition around a lot of concepts used in Deep Learning topologies.
- ChadyWady 9y agoFor ML, the other users gave a good coverage of topics. But AI is an incredibly broad field, and each specialty uses different math topics. Learning all of the math would be infeasible. What are your particular interests? Russell and Norvig have a good book at http://aima.cs.berkeley.edu http://aima.cs.berkeley.edu that covers many different topics in AI, although it is definitely not comprehensive. I would say that whatever you learn in an undergraduate CS degree would give you a good starting point for learning any particular AI topics.
- wickedgamer 9y agoCalculus Functions Derivatives Integration Analytic Geometry That's all I think *http://shrugemojis.com/shrug-emoji/ http://shrugemojis.com/shrug-emoji/
- wickedgamer 9y agoYet it depends. Theres a lot out there on google one could learn. http://fitnessjab.com/ http://fitnessjab.com/
- ratsimihah 9y agoMake sure to differentiate between AI researcher and applied AI software engineer, or whatever that is called. The former needs the mathematical background mentioned here to develop groundbreaking algorithms or improve on existing ones, while the latter merely implements them and requires a much smaller mathematical toolset.
- mindhash 9y agodemystified has good series on calculus and linear algebra. Its light weight
- dekhn 9y agoLinear algebra, probability, and tree/graphs.
- blubb-fish 9y agojust start learning and you will see what math comes along ;)
- bjourne 9y agoIt depends on what "pursuing ML/AI" means. I've written a recommendation engine with barely understanding linear algebra and a spam filter without knowing Bayes theorem. A programmer can work on ML systems without having a solid foundation in higher maths. However, if you want to develop your own solutions then you surely need the math. I would recommend reading Toby Segaran's Programming Collective Intelligence: http://shop.oreilly.com/product/9780596529321.do http://shop.oreilly.com/product/9780596529321.do
- jochenleidner 9y ago1. You can get a long way with high school calculus and probability theory. 2. Regarding books I second the late David McKay's "Information Theory, Inference and Learning Algorithms" and the second edition of "Elements of Statistical Learning" by Tibshirani et al. (there's also a more accessible version of a subset of the material targeting MBA students called James et al., An Introduction to Statistical Learning). Duda/Hart/Stork's Pattern Classification (2nd ed.) is also great. The self-published volume by Abu-Mostafa/Magdon-Ismail/Lin, Learning from Data: A Short Course is impressive, short and useful for self-study. 3. Wikipedia is surprisingly good at providing help, and so is Stack Exchange, which has a statistics sub-forum, and of course there are many online MOOC courses on statistics/probability and more specialized ones on machine learning. 4. After that you will want to consult conference papers and online tutorials on particular models (k-means, Ward/HAC, HMM, SVM, perceptron, MLP, linear and logistic regression, kNN, multinomial naive Bayes, ...).
- bootcat 9y agoProbability and Statistics to begin with !
- wadams19 9y agoTotally depends on where you want to land on the engineering-AI-products to pure-AI-research spectrum. So, what do you mean by "pursuing"? But even still, I would caution against trying to upload a bunch of new math concepts into your brain without first understanding the ML/AI context. I would say go through both of Andrew Ng's ML and DL courses on Coursera. Then, pick a domain/ problem that you're interested in. Then, read papers about how ML/AI is applied in that domain. Then, try to reproduce a paper that you understand and are interested in.
- xchip 9y agoAll I needed to write my conv net library was to understand the chain rule and some basic multiplication. People like to make this look harder than it is.
- bitL 9y agoCalculus (preferably both multi-variate and discrete), probability, statistics, operations research, graph theory, topology, computational complexity. All depends on how deep you'd like to go.
- clircle 9y agoDiscrete calculus? I think you mean univariate.
- ronald_raygun 9y agoI'd say everything in this list is good to know http://pages.cs.wisc.edu/~tdw/files/cookbook-en.pdf http://pages.cs.wisc.edu/~tdw/files/cookbook-en.pdf
- roumenguha 9y agoI believe there's a more recent version of this document available here: http://statistics.zone http://statistics.zone
- deepnotderp 9y agoBasic high school calculus and linear algebra is really the only required thing. I would recommend probability theory and statistics as well.
- coconut_crab 9y agoMaybe unrelated to OP's question but I have always felt that it is impossible to get a job in AI/ML without a PhD in that field (by getting a job I mean do something new/useful and not just coding algorithm devised by other people). I studied mechatronics in university and fairly comfortable with math (calculus, linear algebra and stats), I have even written a small neural network back then to optimise parameters for lathe machining. But that's no where near enough for a job in AI/ML. Unlike writing a web page, which someone can learn within a week to produce something usable, I feel like you need years and years of studying to barely get a start in ML/AI and there is no hope for us non computer scientist at all. [Added] Of course writing webpages pays well enough, but I still can't shake off this feeling that I am missing something by not jumping on the AI/ML train though.
- sunsu 9y agoNone. You can be a productive ML engineer without understanding the math. Many elitist engineers here will downvote me, but its true. ML libraries that allow you to quickly get productive have come a long way. BUT, you have to have a solid understanding of WHICH algorithms/tools to use WHEN. There is also a lot of "voodoo" knowledge to gain that isn't well documented or explained (unrelated to maths).
- master_yoda_1 9y agoThere are two way to approach ML/AI 1) First read all the prerequisites and then work on a problem 2) Start working on a problem and learn all the math ML/AI as you need The second option works best.
- pramalin 9y agoI found this study plan very useful for me. https://www.analyticsvidhya.com/blog/2017/01/the-most-comprehensive-data-science-learning-plan-for-2017/ https://www.analyticsvidhya.com/blog/2017/01/the-most-compre... Provides a very good idea of the courses required and their time frame. I roughly followed along this path but took "Analytics Edge" https://www.edx.org/course/analytics-edge-mitx-15-071x-3 https://www.edx.org/course/analytics-edge-mitx-15-071x-3 for introduction into ML algorithms.