5 ms·
Q-learning blew my mind when I took an AI course in college. One thing I never fully understood, why would anyone choose SARSA when they could use Q-learning?
by 2bitencryption 7y ago
Q-learning blew my mind when I took an AI course in college.
One thing I never fully understood, why would anyone choose SARSA when they could use Q-learning? I believe they use the same inputs, and are nearly the same algorithm, but Q-learning is off-policy while SARSA is on-policy (if I remember right?)
- tlb 7y agoDouble Q-learning can learn from off-policy data. But it's fairly tricky to get all the tuning parameters right. It's a good choice if you can start up 100 or more runs with a range of parameters and pick the one that worked best.
- fnord77 7y agowas it the AI Pacman? I did that project, it really was mind-blowing
- BlahBoy3 7y agoFrom my understanding, SARSA could be more ideal when there is a greater cost associated with making a mistake whilst learning. SARSA is more conservative, as it takes into account possible large negative rewards during the exploratory phase. The classic example problem is "cliff walking."[0] [0] https://github.com/cvhu/CliffWalking https://github.com/cvhu/CliffWalking
- loehnsberg 7y agoSARSA follows the current policy. Suppose you're minimizing cost. Then, if the value function of the MDP is a lower bound, it will explore interesting states simply because SARSA underestimates the cost-to-go. If updates of the value function then tighten this lower bound, SARSA will converge. In this case, SARSA is more efficient than Q-learning. Apart from that using Q-factors does not scale well. If your action space is a game controller, things may still look ok, but not if your action space is multi-dimensional and continuous.