5 ms·
If you factor out chip size and clock frequency, they do actually stagnate. They add new features, but legacy rendering performance is improved by single-digit
by hydroreadsstuff 6y ago
If you factor out chip size and clock frequency, they do actually stagnate. They add new features, but legacy rendering performance is improved by single-digit percentage points, I believe.
If you look at A100 compared to V100 for e.g. FP32 FMA performance (not tensor). 14.1TFLOPS -> 19.5 = +38%, for 2x transistors (16->7nm), +35% SMs and 250W->400W is not that great. Note that NVIDIA uses boost clock for all A100 numbers and seems not have published any base clock so far. So their is a chance that actual sustained A100 performance is lower.
Turing GPUs have rather large dies. TU102 754nm vs GP102 471nm. So comparing them as is, isn't quite fair.
On the CPU front Intel used to use rather small dies for consumers (and even use die shrinks to just cram more chips onto a waver -> more $$$), but now that AMD forces their hand, they are giving in. But of course a lot this area goes into extra cores, not single threaded performance (diminishing returns there).
- derefr 6y ago> legacy rendering performance is improved by single-digit percentage points, I believe. I feel like one would expect this, given that the rendering engines of any given generation, are architected under specific assumptions about the optimal relative "shape" of their graphics pipeline. The clearest example being old game consoles. You could write a SNES emulator for the SuperFX coprocessor, that used the host's GPU to render the polygons, but the rendering would not go any faster than it does on the SNES, because the draw commands are just being trickled out as the physics-engine work necessary to update their positions gets done in fits and starts per scan-line. That trickling-out was necessary on a SNES, because there wasn't time to run all that logic during VBLANK; but in the modern era, we have the opposite problem — that the logic could all be completed during VBLANK (with tons of room to spare), but instead is being "dragged out" across 512 HBLANKs, such that the GPU only gets the full picture of the completed scene at the last moment. (Despite the recent source-code leak, rewriting StarFox to make it render at 60FPS or more will not be a trivial process.) The same thing is true, to a varying extent, of all legacy renderers. They're written for graphical pipelines that just don't match the one we have now. The one we have now is "wider" — more parallel — in so many places, but if it's just being used to recapitulate a long, serial set of fixed-function legacy pipeline stages, that width doesn't help it any.
- pizza234 6y agoWell, from the user perspective, chip size and clock are crucial factors, as GPU workloads tend to be parallelizable. Besides, chip size and clock don't come for free, so I think they shouldn't be discarded regardless. AFAIK, with each generation, Nvidia has increased raw performance by 20-30%, which is significant. On the other hand, the wall of physical limits is getting closer (which may restrict the chip size), but until then, GPUs are faring very well. GPUs functionalities also have a very different nature (as you point out), but this can play well for the user. Realism depends on many functionalities, which have plenty of headroom for improvement (in the sense of hardware support); see ray tracing, which supposedly, is going to be significantly faster on Ampere.