3 ms·
The author mentioned AlphaGo and Alpha Zero without mentioning OpenAI gym and OpenAI Five. Those products show OpenAI was innovating and leading in RL at that
by paradite 1y ago
The author mentioned AlphaGo and Alpha Zero without mentioning OpenAI gym and OpenAI Five.
Those products show OpenAI was innovating and leading in RL at that stage around 2017 to 2019.
https://github.com/openai/gym https://github.com/openai/gym
https://en.wikipedia.org/wiki/OpenAI_Five https://en.wikipedia.org/wiki/OpenAI_Five
- bitpush 1y agoThis is the first I'm hearing about it.
- paradite 1y agoI forgot to mention that OpenAI also invented PPO, which is the default algorithm that everyone uses for RL since 2017: https://en.wikipedia.org/wiki/Proximal_policy_optimization https://en.wikipedia.org/wiki/Proximal_policy_optimization DeepSeek's GRPO is also just a minor variant of PPO.