4 ms·
Reinforcement Learning: An Introduction (2018) [pdf]
- keeptrying 8y agoThe authors , Barto and Sutton take such a complicated subject and explain it in such simple prose. I don’t think think I’ve read any other work that does this as well. Also RL is only going to grow in use and popularity. Highly recommend it for mL practitioners.
- abhgh 8y agoI hope it grows in popularity if only because its an interesting take on learning. I did a course on RL in 2007 and our textbook was the 1st edition of this book - back then, it was perceived to be a very niche area and a lot of ML practitioners (there weren't many of those either :) ) had only just about heard of RL. I am happy that it's popular today.
- deleted 8y ago[deleted]
- colmvp 8y agoI also recommend interested people to watch David Silver's RL lectures at UCL on YouTube. He covers material from the book. https://www.youtube.com/watch?v=2pWv7GOvuf0&list=PL7-jPKtc4r78-wCZcQn5IqyuWhBZ8fOxT https://www.youtube.com/watch?v=2pWv7GOvuf0&list=PL7-jPKtc4r...
- deleted 8y ago[deleted]
- nafizh 8y agoWith all its hype in RL, I am yet to see significant real life problems solved with it. I am afraid with all the funding going into it, and nothing to show for except being able to play complex games, this might contribute to the mistrust in proper utilization of research funds. Also the reproducibility problem in RL is many times worse than in ML.
- imh 8y agoI'd bet that sample efficiency is a factor in translating they most hyped bits of RL into solving IRL problems. So many business problems translate to "Learn which of these things to do, as quickly and cheaply as possible."
- wenc 8y agoSo I can see RL augmenting traditional optimal control in regimes outside of previously modeled spaces. For instance, a machine would operate via optimal control in regimes that are known and characterized by a model, but if it ever gets into a new unmodeled situation, it can use RL to figure stuff out and find a way to proceed suboptimally (subject to safety constraints, etc.). An illustrative example is Roomba. Roomba is probably based on some form of RL, and it does a decent job. But suppose we have a map of the room that Roomba can use -- this would let it plot the optimal path. However suppose the map of the room is incomplete. Roomba can still operate near optimally within the mapped area, but will have to learn the environment outside the map. Or if the layout of the room has changed since the map was created (new furniture), Roomba's RL can kick in.
- milaresearcher 8y agoI agree with you that it's early days for RL. I think some companies are using it in their advertising platforms, but it's not really my field. That said, I strongly disagree about what constitutes the proper utilization of research funds. IMO, society should invest in basic research without the expectation of solutions to significant real-world problems. Still, I'd be really surprised if I don't see advances from the field of reinforcement learning used in a ton of applications during my lifetime.
- closed 8y ago
- svalorzen 8y agoIf you ever feel like trying out the algorithms contained in the book without going to the trouble of reimplementing everything from scratch feel free to come over to https://github.com/Svalorzen/AI-Toolbox https://github.com/Svalorzen/AI-Toolbox. This is a library I have maintained during the past 5 years and implements quite a lot of RL algorithms, and can be used with both C++ and Python. It's very focused on being understandable and having a clear documentation, so I'd love to help you out starting up :)
- clickok 8y agoCool! I'd also like to plug my own RL-related repositories: https://github.com/rldotai/rl-algorithms https://github.com/rldotai/rl-algorithms and https://github.com/rldotai/mdpy https://github.com/rldotai/mdpy . The first one implements some of the more "exotic" temporal difference learning algorithms (Gradient, Emphatic, Direct Variance) with links to the associated papers. It's in Python and heavily documented. The second one (mdpy) has code for analyzing MDPs (with a particular focus on RL), so you can look at what the solutions to the algorithms might be under linear function approximation. I wrote it when I was trying to get a feel for what the math meant and continue to find it helpful, particularly when I'm dubious about the results of some calculation.
- wenc 8y agoMy understanding is RL is a reasonable attack for situations where the environment is either (1) mathematically uncharacterized (2) insufficiently characterized (3) characterized, but resulting model is too complex to use, and therefore RL simultaneously explores the environment in simple ways and takes actions to maximize some objective function. However, there are many environments (chemical/power plants, machines, etc.) where there are good mathematical/empirical data-based models, where model-based optimal control works extremely well in practice (much better than RL). I'm wondering why the ML community has elected to skip over this latter class of problems with large swaths of proven applications, and instead have gone directly to RL, which is a really hard problem? Is it to publish more papers? Or because self-driving cars?* (* optimal control tends to not work too well in highly uncertain, non-characterized, changing environments -- self-driving cars are an example of one such environment, where even the sensing problem is highly complicated, much less control)
- computerphage 8y agoDo you have an example of a self driving car company that uses RL?
- wenc 8y agoNope.
- svalorzen 8y agoRL is actually quite an umbrella term for a lot of things. There's policy gradient methods, which improve directly on the policy to select better actions, there's value based methods which try to approximate the value function of the problem, and get a policy from that, and there's model based methods which try to learn a model and do some sort of planning/processing in order to get the policy. Using model based methods can allow you to do some pretty fancy stuff while massively reducing the number of data samples you need, but on the other side there's a trade off. Using the model usually tends to require lots of not-very-parallelizable computations, and can be more costly computationally. Very large problems can get out of hand pretty quickly, and there's still a lot of work to do before there is something which can be applied in general quickly and efficiently.