3 ms·Reinforcement Learning, Policy Gradient Is Nothing More Than Random Search2 points by dspoka 9y ago