Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joschu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
joschu
8y ago
When optimizing high-dimensional policies, the gap in sample complexity between PPO (and policy gradient methods in general) and ES / random search is pretty big. If you compare the Atari results from the PPO and ES papers from OpenAI,
2.
▲
by
joschu
13y ago
Love the Isaac Asimov reference in "Multyvac" I use Picloud for my CS PhD research, I hope to be able to continue with Multyvac or the open-source version.