4 ms·
CNTK has significantly higher performance with one or more machines; great multi-gpu scalability. Can train harder on bigger datasets given your resources.
by ajwald 10y ago
CNTK has significantly higher performance with one or more machines; great multi-gpu scalability. Can train harder on bigger datasets given your resources.
- corysama 10y agoCan someone clarify this? In my head "one or more machines" means "always". Does CNTK generally have higher perf even on a single machine? Or is ajwald trying to say it is better at scaling to multiple machines.
- cmarschner 10y agoCNTK user & contributor here. CNTK overall has very low framework overhead and has tensors with dynamic axes as first-level citizens. This means that sequences can be expressed without needing to do padding, sorting of the input data, or any other workarounds, and can be packed automatically by the toolkit in an optimal way. In particular, while it is laying out the rectangular structure it uses to traverse multiple RNNs of a minibatch in parallel, it fits shorter sequences into the holes and can reset the RNN state for these sequences while it is traversing this structure. This makes CNTK especially suitable for expressing RNN models (for CNNs many of the calls are just forwarded to CuDNN, so the difference might be much lower). As for distribution, a) it has an extremely simple way to run data parallelism (for CNTK 1 it was just using MPI and starting the worker with a few extra options. I think CNTK 2 will add this in a week or so to the Python bindings), b) it has 1-bit SGD and more recently BlockMomentum, which are just dead simple methods to use for distributing the gradients, and they just work. All of these are open source (though 1-Bit SGD and BlockMomentum are patented).
- dcl 10y agoAn algorithm for updating parameters is patented?
- nl 10y agoThere's a saying: lies, damn lies and benchmarks. In this case AFAIK mi one has replicated these claims. MS hasn't got it in the standard CNN inference benchmark yet though: https://github.com/soumith/convnet-benchmarks https://github.com/soumith/convnet-benchmarks. This is the benchmark that showed how slow TensorFlow was on non-Google libraries until Google fixed it.