3 ms·
And they also have builtin support for 1-bit SGD. Compresses a model "down to 1 bit per weight" [0]. Seems to be a general technique for model compression, but
by iraphael 9y ago
And they also have builtin support for 1-bit SGD. Compresses a model "down to 1 bit per weight" [0]. Seems to be a general technique for model compression, but it's nice for deployment to have it built in. This also doesn't seem to be a new addition to CNTK, just something I didn't know before.
[0] https://www.microsoft.com/en-us/research/publication/1-bit-stochastic-gradient-descent-and-application-to-data-parallel-distributed-training-of-speech-dnns/ https://www.microsoft.com/en-us/research/publication/1-bit-s...
- minimaxir 9y ago1-bit SGD is cutting-edge deep learning tech. (so cutting edge it has a different license than CNTK itself: https://docs.microsoft.com/en-us/cognitive-toolkit/CNTK-1bit-SGD-License https://docs.microsoft.com/en-us/cognitive-toolkit/CNTK-1bit...)
- Permit 9y agoIs this common? Do libraries such as Tensorflow occasionally license portions of themselves only for non-commercial use? I ask because when CNTK was first shared it was also under a non-commercial license. It seemed to subsequently drop off the radar for many people (eg. It wasn't mentioned in Stanford's ConvNet course or in Udacity's Machine Learning course).
- justnikos 9y agoThere's actually two 1-bit things going on with CNTK 2.0. One is the 1-bit SGD that has long been in CNTK and has been criticized for its weird license. I am not a lawyer, but my understanding is it says something like you cannot use this unless you call it from inside CNTK. The terms are not that bad and you don't have to use 1-bit SGD if you don't like them. CNTK 2.0 has another 1-bit thing going on as well which is binary convolution. This uses the Halide compiler to generate code that is 10x faster than optimized 32-bit convolution. This still seems to be at a proof of concept stage. Disclaimer: I work at Microsoft.