3 ms·
*11.5% as much compute.
by Willson50 6y ago
*11.5% as much compute.
- The_rationalist 6y agoBut what was the baseline hardware for a reasonable training time?
- nl 6y agoPage 7 has a table of one training step on TPUv3 and V100 GPUs. I don't completely understand this: NFNet is slower than its competitors on this benchmark, but they claim higher efficiency. This isn't obvious to me.
- trott 6y ago> I don't completely understand this: NFNet is slower than its competitors on this benchmark, but they claim higher efficiency. Take a look at F1 and B7: They have the same accuracy, but F1 is smaller and much faster.