3 ms·
This seems like a solid example of practicing Deep RL with a few lines of code. With the OpenAI thing you run a program from command-line. Curious if you have a
by robbiemitchell 8y ago
This seems like a solid example of practicing Deep RL with a few lines of code. With the OpenAI thing you run a program from command-line. Curious if you have any insight about what the algorithm is and how it compares to the Keras one?
- mtrazzi 8y agoBoth the Keras one and the one from spinning up can be called "vanilla policy gradients". The one in Keras is closer to REINFORCE, and the one from spinning up use actor-critic and a multilayer perceptron.