4 ms·
As a chip designer, I'm dubious of this claim: "The advantage of performing matrix multiplication using only addition for arithmetic is that it then becomes fe
by Drunk_Engineer 3y ago
As a chip designer, I'm dubious of this claim:
"The advantage of performing matrix multiplication using only addition for arithmetic is that it then becomes feasible to build special- purpose chips with no multiplier circuits. Such chips will take up less space per on-chip processor."
Perhaps that is true in a toy design, but in real-world chips the multiplier uses only a very tiny fraction of chip real-estate. And even if the matrix-multiply can be eliminated, there are other uses for multiply operations.
I once attended a chip design conference where NVidia discussed its latest GPU. In one of the slides showing block layout, the designers pointed out how barely any silicon was being used for actual floating operations -- the vast majority was for pipelining and moving bits around.
- brigade 3y agoThat’s because CPUs and GPUs are useful for more than just matrix multiplication. TPUs aren’t; they assume highly regular data movement in and out of the ALUs.
- hedgehog 3y agoIt of course depends on workload but I know more than a little about this specific problem space and reducing the space+energy cost of the multipliers is useful. If the idea proposed worked well enough it might be a useful block in a camera ISP chip, audio interface for wake word and speech preprocessing, and similar applications where the models are small and energy is precious.
- creato 3y agoI don’t know, sorting requires a lot of moving data around, which is expensive for energy too. Maybe this will be used for vectors of small fixed/bounded size, but then you don’t get to amortize the cost of the logs as much either.
- imtringued 3y agoI once saw a professor show me their latest taped out CPU (multi project wafer of course) and it had one tiny core that is supposed to control the wake-up of the larger processor. The sleep controller has no memory and is a thin slice that is barely visible. Meanwhile the primary processor is much larger but it too is completely dwarfed by the memory it is being surrounded by. The idea of throwing out parts of the ALU is ridiculous. The only situation where this would make sense is in some kind of processing in memory situation where your logic process does not permit large CPUs and you expect to have hundreds of cores per chip with multiple chips on a single DIMM.
- daniel-cussen 3y agoNot that tiny, n not such little energy, if Intel AVX-512 requires throttling the chip down to 60% of full speed (n similar slowdowns are necessary on Xeon Phi, 1.4 GHz to 1.2 GHz for full) SIMD when using "heavy" operations like multiplication (and also count-leading-zeroes aka CLZ) so it's a lot of energy for sure.