5 ms·
Deep Reinforcement Learning in Depth in 60 Days
- lkhatter 8y agoI’ll do this, please keep it updated as the weeks go by!
- minimaxir 8y agoAre the only resources you're referencing those by others, or do you plan to include projects/lessons you yourself have made? There has been a rise lately in Machine Learning/Deep Learning resources which have zero original projects or original ideas, just a list of papers/blog posts (or worse, MOOC teachers/YouTubers who do that and obfuscate the source of the original ideas). While that's an educational option, it's, in my opinion, lazy and avoids furthering the ecosystem as a whole.
- curiousgal 8y ago> While that's an educational option, it's, in my opinion, lazy and avoids furthering the ecosystem as a whole. I've missed HN's cynical take on things lately. There's absolutely nothing wrong with using those ressources to learn. You say it's an educational "option" whereas the entire purpose is educational.
- minimaxir 8y agoTrue, there's nothing wrong with using these resources, although I'm a bit disappointed when a "Awesome List of ML/DL" pops up every other week on HN with similar content/topics. I apologize for being overly cynical.
- curiousgal 8y agoI am equally disappointed as well but I realized that my disappointment should be directed towards the people upvoting such lists and not towards whomever's creating them, because, on a personal level, such lists are useful. For the entire field however, I agree with you.
- andri27 8y agoMy goal is to put together, for each week, theoretical material done by experienced people (e.g. video,papers,books ecc..) with projects done by myself.
- DrNuke 8y ago> There has been a rise lately in Machine Learning/Deep Learning resources which have zero original projects or original ideas Nah, not lazy imho. On one side, the barrier to just messing around is pretty low these days but the barrier to original projects or ideas is quite steep without a strong domain expertise and a nicely assembled dataset; on the other side, the arXiv repository and the most relevant conferences are possibly more suited to originality than the average list by Noob McNobody from his/her basement in Nowhere, Planet Earth?
- nafizh 8y agoWith all of its excitement surrounding RL, I am yet to see substantial practical applications of RL in real life apart from games, and some articles I read on how companies use RL for recommendation or ad suggestion. So, it is indeed kind of puzzling to understand as an outsider what generates this excitement.
- tikhonj 8y agoReinforcement learning is a solid fit for a number of traditional operations research problems. (In fact, that's pretty much what motivated research into RL in the first place, as I understand it.) One concrete example I heard about was using reinforcement learning to price airline tickets. Behind the scenes, airlines break up the tickets on a single plane into a large number of distinctly priced types. The question of how much of each type of ticket to offer at what price and how to change this over time (as the actual flight is coming closer and closer) is a massive optimization problem that's too large to solve exactly. Reinforcement learning coupled with simulation can find good solutions if you set up the feature space correctly. (In this case, I remember that the only feature that ultimately mattered was either total profit or total revenue for the mix of tickets being offered.) One thing to note here is that this is using "normal" reinforcement learning, not "deep" reinforcement learning. You can get away with having a simple functional approximation of the state instead of reaching for a neural network. This seems true for most operations research problems where you would reach for reinforcement learning—figuring out a way to model the state by hand works well enough and has the important benefit of being easier to understand and interpret. The "deep" part becomes useful when your state space is so large and complex that other techniques become infeasible.
- kleiba 8y agoSpoken dialog systems. Check out Steve Young's group at Cambridge, it's considered the state of the art (well, at least by some ;-)).
- gaius 8y agoSerious A/B testing is usually done with a bandit algo these days, you can converge with far fewer trials than a statistically meaningful A/B. News or other articles recommendations are often bandits too - MSN being a prime example.
- guard0g 8y agoDRL in financial derivative pricing, risk modeling, HFT. Check out Igor Halperin on Coursera.