3 ms·
A lot of comment are sneering at various aspects of this press release, and yeah, there's some cringeworthy stuff. But the technical aspects are pretty cool:
by PoignardAzur 2y ago
A lot of comment are sneering at various aspects of this press release, and yeah, there's some cringeworthy stuff.
But the technical aspects are pretty cool:
- Fault-tolerant training where nodes and be added and removed mid-run without interrupting the other nodes.
- Sending quantized gradients during the synchronization phase.
- (In the OpenDiLoCo article) Async synchronization.
They're also mentioning potential trustless systems where everyone can contribute compute, which would make this a truly decentralized open platform. Overall it'll be pretty interesting to see where this goes!
- londons_explore 2y ago> Sending quantized gradients during the synchronization phase. I did this 9 years ago, works pretty well. I don't understand why all ML isn't async and quantized like that now. This project quantizes to 1 bit per weight and it works so well I didn't even make it configurable. https://github.com/Hello1024/shared-tensor https://github.com/Hello1024/shared-tensor
- radarsat1 2y ago> 1 bit per weight does this basically correspond to moving each weight either up or down by a fixed amount? I'm a bit surprised you don't at least need a "stay same" bit, but i suppose it could balance out over multiple iterations. Interesting that it works at all. Although, thinking on it, I could see it maybe even having a nice regularizing effect where every layer would end up have similar weight magnitudes. (like projecting onto the local n-ball as mentioned in a paper posted recently on HN)
- f_devd 2y agoIt has been more formally studied in signSGD[0], and empirically it's comparable to Adam in terms of behavior. [0]: https://arxiv.org/pdf/1802.04434 https://arxiv.org/pdf/1802.04434
- londons_explore 2y agoThis is for keeping the weight vectors in sync between two machines. The weight vectors themselves are regular floats. But the data exchanged between the machines is 1 bit. Basically, you keep track of changes to the weight vector which hasn't yet been propagated to the other machine. You quantize this to 1 bit per weight (ie. a sign bit) and send it, together with a single scale factor X, accumulating the quantization error for the next sync iteration. You choose X to be the RMS or some similar metric of the accumulated error.