2 ms·
The tensor cores are basically more efficient because they don't implement a 'read from register file; execute one instruction; write to register file'-style ar
by frogblast 6y ago
The tensor cores are basically more efficient because they don't implement a 'read from register file; execute one instruction; write to register file'-style architecture:
They are basically a big array of ALUs hard-wired in a grid to do a matrix multiply: The connectivity between ALUs is baked into silicon: no general purpose register file to hold the intermediate results, just wires directly from the output of one ALU into the input of the next (well, and registers for pipelining).
No overhead for register files, instruction decoding, instruction scheduling, load/store units, cache, etc. (load/store is handled by the usual SM instruction set).
This is basically the same concept you find in any "AI acceleration" hardware (Google's TPUs, various neural-thingys you find in Phones, etc).
As for using in games, this hardware is generally idle, except for a few cases (typically because Nvidia did the work to light them up):
- upscaling from lower resolution rendered products to higher resolution (Nvidia markets this under the name "Deep Learning Super Sampling"), which is basically taking a derivation Temporal Anti-aliasing and mapping it onto the tensor cores to exploit that hardware
- Doing CNN-like denoising and reconstruction of a noisy ray traced image.