4 ms·
Decision Transformer; Transformer for RL partially ditching dynamic programming
- conview 5y agoThe "Decision Transformer: Reinforcement Learning via Sequence Modeling" : by Chen L. et al explaining how to use transformer to replace dynamic programming for context length reference.
- JamilD 5y agoA tl;dr, because the idea behind the Decision Transformer is really cool and elegant; they model reinforcement learning tasks in a similar way to autoregressive language modeling. In language modeling, you want to learn, for example, the next word in a sentence. Given a sequence of tokens t₀, t₁, t₂, predict t₃. The authors here model RL as a similar autoregressive task, where we want to predict which action to take, given a sequence of previous actions, states, and "rewards-to-go", or the estimated remaining rewards in the trajectory. For example, given (s₀, a₀, R₀), (s₁, R₁), we want to predict the best a₁. This allows us to train in a supervised way by accumulating trajectories from either training data or random walks. Then, at inference time, all we do is input into the model (a) the start state s₀, and (b) the desired "rewards-to-go" R₀ we'd like to have. The output at the next step is an action a₀. We then calculate the next state s₁ caused by taking a₀, and the remaining rewards-to-go R₁, to find the best a₁. We can repeat this with partial trajectories until the rewards-to-go hit 0, or until the trajectory is complete.
- radarsat1 5y agoIt all makes sense but it strikes me as odd (that is, not very RL-like) that it doesn't allow to actually "maximize" reward. What do you do, just specify a very high target? Then it seems you are depending very strongly on out of distribution generalization, which seems dangerous (but apparently works)
- omegalulw 5y agoWhile I have yet to read the paper, I don't see how this is a good idea. One of the fundamental ideas of RL is MDPs, and that is still around because problems that aren't approximable as MDPs are very hard. With MDPs, your goal is to understand the state and action values so that you don't have care about the history or sequences. If you throw MDPs away you are getting rid of a very useful inductive bias and might just be making your problem a lot harder.
- jononor 5y agoMDP is Markov Decision Process, a process that has the Markov property - future state depends only on current state and new inputs, not the history/sequence of states that got you there.
- d4rkp4ttern 5y agoAny process can be Markovian given a sufficiently complex state :)
- visarga 5y agoAlso recommended watch - Yannic Kilcher: Decision Transformer: Reinforcement Learning via Sequence Modeling (Research Paper Explained) https://www.youtube.com/watch?v=-buULmf7dec https://www.youtube.com/watch?v=-buULmf7dec
- bitL 5y agoWhy restrict oneself to RL though? One could potentially escape Markov curse (no memory) with this one.
- zwaps 5y agoHaha dang it I was working on this also :< :> kudos to the authors