4 ms·
In reinforcement learning, you typically need to run many copies of an environment (an Atari emulator for example) to collect experience for your agents. Enviro
by slewis 6y ago
In reinforcement learning, you typically need to run many copies of an environment (an Atari emulator for example) to collect experience for your agents. Environment code usually makes more sense on a CPU, because it’s a lot more branchy and stateful.
Think of two agents that run for awhile and end up in different parts of a virtual world, that require very different code to execute. This might be pretty difficult to parallelize on a GPU.
After collecting a bunch of experience, you update whatever function you’re optimizing. This could be something like a neural network that is the brain of your agents. The update step can happen on a GPU.
But then we need to go collect a lot more experience again, so we’re back in CPU land.
Experience collection often dominates overall compute time in reinforcement learning. Which means you want a lot of CPUs.