3 ms·
The performance advantage was still in the ballpark of 4-8x faster for training MNIST on the CPU, which while smaller than most networks people are training on
by celrod 4y ago
The performance advantage was still in the ballpark of 4-8x faster for training MNIST on the CPU, which while smaller than most networks people are training on their GPUs, still has more than 40 thousand parameters.
For someone with a statistical background, this is a lot of parameters. John von Neumann could wiggle a lot of elephant trucks.
A lot of practical/useful models fill the range from the tiny ones we may use in UDEs and SciML to this MNIST convnet.