21 ms·
As of now, Apple's version is uncompetitive for training. It is decent for inference, which is good but doesn't cover the training usecase which is the main use
by ctchocula 4y ago
As of now, Apple's version is uncompetitive for training. It is decent for inference, which is good but doesn't cover the training usecase which is the main usecase for TPUv4 and Nvidia GPUs like A100. A100 provides 312 TFlops, but M1 GPU only 5.1 TFlops, which is 2 orders of magnitude slower.
This blog [1] does make the argument that M1 Max has about 8x lower performance than a consumer grade Nvidia GPU 2080, but also uses 8x less watts, so it's possible Apple can make a product with similar performance/watt. However, I would say that they are farther away than TPUv4 for now, because the product doesn't exist.