4 ms·
https://en.wikipedia.org/wiki/FLOPS https://en.wikipedia.org/wiki/FLOPS Intel Rocket Lake FP64: 32 Flops per core per cycle at 8 cores and 5GHz Nvidia Ampere:
by uniqueuid 4y ago
https://en.wikipedia.org/wiki/FLOPS https://en.wikipedia.org/wiki/FLOPS
Intel Rocket Lake FP64: 32 Flops per core per cycle at 8 cores and 5GHz
Nvidia Ampere: 1/32 Flops per core per cycle at 2048 cores and ~2 GHz
Even though there are more architectural differences, there IS a difference in orders of magnitude ... for the things that you can do on a GPU (SIMD).
- bhouston 4y agoThis is because FP64 is emulated on the NVIDIA Ampere as it is missing dedicated FP64 hardware. This is not a fair comparison. FP32 is 64x faster on NVIDIA Ampere than FP64 per: https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-gpu-architecture-whitepaper-v2.pdf https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-gpu-... Ampere has 60 TFLOPs of Performance on its top GPU, the Hopper per: https://en.wikipedia.org/wiki/Ampere_(microarchitecture) https://en.wikipedia.org/wiki/Ampere_(microarchitecture). Which is roughly 50% to 100% faster than Rocket Lake depending and that doesn't account for the parallelism/memory benefits that comes with the GPU architecture.
- hajile 4y agoLet's analyze 3 real cores. A100 has 6912 cores (excluding tensor as they don't do 32-bit float math). Each core does 2 (32-bit) floats per cycle and runs at 1410MHz peak frequency. It is 826mm^2 and has an official 300w TDP. This yields 19491.84 GFLOPS. Threadripper (Zen 3) has 64 cores. Each core does 32 (32-bit) floats per cycle and runs at 4300MHz. It has 8 CPU chiplets (80.7mm^2 each) and one IO die (416mm^2) for a total area of 1061.6mm^2. This chip yields 8806.4 GFLOPS. Alder Lake has 8 P-cores and 8 E-cores. If you have an unlocked chip with AVX-512, that yields 64 (32-bit) floats per cycle running at up to 5.3GHz along with 32 (32-bit) floats per cycle running at up to 4GHz. Die area is around 215.25mm^2 (though much of that die area is taken up by the GPU) and it has a TDP of 150w (peak 241w). This chip yields 2713.6 GFLOPS for the P-cores and 1024 GFLOPS for the E-cores for a total of 3737.6 GFLOPS. The first thing of note is that orders of magnitude is a gross overstatement as there isn't even a single order of magnitude of difference between the slowest and fastest systems here. Dividing GFLOPS by die area gives 23.60 GFLOPS/mm^2 for the A100, 8.30 GFLOPS/mm^2 for the Threadripper, and 17.36 GFLOPS?mm&2 for the Alder Lake. Considering the GPU size of the Alder Lake, I suspect that GFLOPS per die area actually favors Alder Lake over the A100 (though this is also somewhat offset by the presence of tensor cores and other non-related stuff in A100). Also noteworthy is that power is a much weirder metric due to how turbos affect all the things. I suspect this is where the A100 has a major advantage not to mention real GFLOPS dropping off steeply in long-running workloads (though this will affect all of these systems to greater or lesser degrees). On the flip side, I should add that branchy code dramatically slows down actual performance on a GPU while only slowing down a CPU when a misprediction happens (in theory this should be only 1-5% of cases). A more interesting question would be Alpha EV9 with it's proposed 1024-bit SIMD unit running at around 2GHz. That chip would have delivered 256 GFLOPS of performance per core around 2005. Scaling to a modest 32-core chip at 4GHz gives a very impressive 16384 GFLOPS while retaining all those CPU advantages too. This leads to the observation that adding large SIMD/vector units to a CPU isn't really hard. A small, in-order core with very wide vector units running at a slower speed to conserve energy seems like an interesting idea and several startups are working in this direction (not to mention this being the basis of Intel's Larabee and later Knight's Corner designs).
- bhouston 4y ago> This leads to the observation that adding large SIMD/vector units to a CPU isn't really hard. A small, in-order core with very wide vector units running at a slower speed to conserve energy seems like an interesting idea and several startups are working in this direction (not to mention this being the basis of Intel's Larabee and later Knight's Corner designs). As well as the Sony/IBM Cell processor used in the PS3. https://en.wikipedia.org/wiki/Cell_(microprocessor) https://en.wikipedia.org/wiki/Cell_(microprocessor) Wide SIMD or just in-order massive FP capabilities is not new. But it is hard to program for. In many ways special purpose ASICs fill this role if you just need raw massive FP capabilities. Or eventually memory embedded computing can even take this further.
- hajile 4y agoThe Cell was a very flawed implementation. Load hit store issues where you could wait around almost a hundred cycles (and stall it because no OoO). Getting rid of SMT and adding complete OoO support would have been a much better result even if the die area was a bit bigger as a result. The SPE had NO branch prediction and were pretty close to VLIW in design. It relied on the provably false trope of a "sufficiently advanced compiler" that can somehow solve the halting problem and of course leading to a 50-50 chance that you plowed ahead on the wrong branch and now must wait almost 20 cycles for everything to reset. As it only had 128 registers to unroll loops into, this problem is basically guaranteed to happen a lot. They didn't provide any cache per core. Instead, programmers have to micromanage the 256k of RAM as if it were cache. Accurately predicting every possible memory pattern is impossible (it would require proving all execution paths which the halting problem says is impossible). Even modern GPUs include cache and choose the appropriate fetching patterns on the fly. I'll also note that the SPE used a different ISA from the CPU (PPE) which adds yet another layer of headaches. We'll never know how powerful the cell could have been because it had so many footguns that devs were seemingly incapable of avoiding all of them. These problems were called out by places like Anandtech or Real World Tech a year and a half before the PS3 launched and was then reiterated by the people using it. Sony allegedly took to threatening teams that wanted to release their finished xbox 360 games before the PS3 version and even took to flying around a couple teams of cell developers to try to get games out the door quicker. A system using the same ISA for both CPU and GPU isn't such a common idea in practice. There are supposedly a couple companies trying to use RISC-V to do this. I guess we'll have to wait and see what they can come up with. I certainly don't see them repeating these major mistakes.
- qeternity 4y ago> here IS a difference in orders of magnitude No, that is a single order of magnitude.