3 ms·
The way I like to explain it, which is how I have seen it explained several times, is that all AI can be reduced to search/optimization. ML is just applying the
by imaltont 4y ago
The way I like to explain it, which is how I have seen it explained several times, is that all AI can be reduced to search/optimization. ML is just applying the search over the function that will search for the final answer over a dataset (either generated on the fly or prepared beforehand). For neural networks the hypothesis space (all the solutions you are searching through to find the best ones) is the weights for the neural network, and your search strategy/optimization is (usually) backpropagation. If you translate the weights to something traversable by other algorithms they could do just "fine" (assuming infinite time and space) in it's place. It really opens the mind up for experimentation on every bit of the process. The book that really hammered it in for me was Intelligence Emerging by Keith Downing, short, great book on bio-inspired AI.
- tehsauce 4y agoThe search strategy/optimizer is actually gradient descent, backpropagation is just an efficient way to compute gradients.
- imaltont 4y agoBoth Tom M. Mitchell's "Machine Learning" as well as Russel & Norvig's "Artificial Intelligence: A Modern Approach" define the whole process from propagating the input until you have an output, calculate the gradient and update the weights.
- tehsauce 4y ago[4] Goodfellow, Bengio & Courville 2016, p. 200, "The term back-propagation is often misunderstood as meaning the whole learning algorithm for multilayer neural networks. Backpropagation refers only to the method for computing the gradient, while other algorithms, such as stochastic gradient descent, is used to perform learning using this gradient."
- imaltont 4y agoAs good as Ian Goodfellow's work has been, I think I have to disagree with him on what is part of the back propagation algorithm. Rumelhart et. al 1986 also describe it is a single algorithm from start to end. It is true you can use other things than stochastic gradient descent though, Mitchell also points this out by saying their algorithm was a backpropagation example using gradient descent.