3 ms·
Reinforcement learning using gradient descent, for instance, requires backpropgation and thus much higher memory loads, as each parameter gets updated via the c
by staturecrane 9y ago
Reinforcement learning using gradient descent, for instance, requires backpropgation and thus much higher memory loads, as each parameter gets updated via the chain rule. ES does not require gradient descent and so each individual agent has much lower memory consumption during training than a deep Q-learning agent would.
You may end up using more memory in ES, but only because you can parallelize with as many processes as your system can handle.