3 ms·
Also, because they are focusing heavily on the RL part of the modeling. They obviously have obscene amounts of available compute, but that is not their competit
by zerostar07 8y ago
Also, because they are focusing heavily on the RL part of the modeling. They obviously have obscene amounts of available compute, but that is not their competitive advantage.
- sgillen 8y agowhat exactly do you mean? Are you saying that RL requires less compute? I would say having an obscene amount of compute is definitely a big competitive advantage, especially over a lot of small academic research labs.
- zerostar07 8y ago> We train each GQN model simultaneously on 4 NVidia K80 GPUs for 2 million gradient steps. The values of the hyper-parameters used for optimisation are detailed in Table S1, and we show the effect of model size on final performance in Fig. S4. > The values of all hyper-parameters were selected by performing informal search. We did not perform a systematic grid search owing to the high computational cost.