4 ms·
Fun idea. GPT @ Home :D. Scatter of the inputs would be very cheap as they are tiny LongTensors (sequences of indices), but the Gather of the gradients seems li
by karpathy 6y ago
Fun idea. GPT @ Home :D. Scatter of the inputs would be very cheap as they are tiny LongTensors (sequences of indices), but the Gather of the gradients seems like a bottleneck. These models can be quite large. Maybe each worker only communicates back some sparse or potentially precision-reduced gradients? In the limit, I recall papers that were able to only communicate one bit per dimension. May also be possible to further reduce the number of weights by weight sharing, or e.g. with HyperNetworks.
- londons_explore 6y agoI built this a few years ago: https://github.com/Hello1024/shared-tensor https://github.com/Hello1024/shared-tensor It does updates to weights based on 1 bit precision updates each iteration. It would be fairly trivial to go to less than 1 bit precision too - simply set some threshold (eg 3), and wherever the difference between the weight on the server and the client is greater than 3, transmit a binary "1", else send a binary "0". Then entropy code all the resulting binary. By adjusting the threshold up and down, you trade off the size of the data to send Vs precision.
- 0-_-0 6y agoI read a paper that did exactly what you describe but of course I can't find it now...
- 0-_-0 6y agoI long wanted to see a proof-of-work cryptocurrency that does neural network training instead of burning through hashes. Imagine if 0.5% of the planet's energy consumptions (9 GW) was used for training neural networks instead of mining bitcoin! It would also solve the problem of ASICS being 1000x more efficient than GPUs, so everyone can participate. It would incentivise the development of efficient neural network training hardware. Somebody do this already!
- deleted 6y ago[deleted]