2 ms·
They use Proximal Policy Optimization, which is pretty much the go-to algorithm for RL at OpenAI. It was developed by the last author of procgen, John Schulman,
by rivesunder 7y ago
They use Proximal Policy Optimization, which is pretty much the go-to algorithm for RL at OpenAI. It was developed by the last author of procgen, John Schulman, et al.
(https://openai.com/blog/openai-baselines-ppo/ https://openai.com/blog/openai-baselines-ppo/)