2 ms·
There are some benchmarks of a neural network toolkit built on top of this: https://github.com/soumith/convnet-benchmarks https://github.com/soumith/convnet-ben
by benanne 11y ago
There are some benchmarks of a neural network toolkit built on top of this: https://github.com/soumith/convnet-benchmarks https://github.com/soumith/convnet-benchmarks
compare NVIDIA's own cuDNN R2 versus NervanaSys-16 and NervanaSys-32. Pretty impressive!
I've also tried out his GEMM implementation on a GTX 980. Seems like it can be up to twice as fast as the one from cuBLAS for some matrix sizes.