3 ms·
Can anyone comment on where efficiency gains come from these days at the arch level? I.e. not process-node improvements. Are there a few big things, many small
by _hark 1y ago
Can anyone comment on where efficiency gains come from these days at the arch level? I.e. not process-node improvements.
Are there a few big things, many small things...? I'm curious what fruit are left hanging for fast SIMD matrix multiplication.
- yeahwhatever10 1y agoSpecialization. Ie specialized for inference.
- vessenes 1y agoOne big area the last two years has been algorithmic improvements feeding hardware improvements. Supercomputer folks use f64 for everything, or did. Most training was done at f32 four years ago. As algo teams have shown fp8 can be used for training and inference, hardware has updated to accommodate, yielding big gains. NB: Hobbyist, take all with a grain of salt
- jmalicki 1y agoUnlike a lot of supercomputer algorithms, where fp error accumulates as you go, gradient descent based algorithms don't need as much precision since any fp errors will still show up at the next loss function calculation to be corrected, which allows you to make do with much lower precision.
- cubefox 1y agoMuch lower indeed. Even Boolean functions (e.g. AND) are differentiable (though not exactly in the Newton/Leibniz sense) which can be used for backpropagation. They allow for an optimizer similar to stochastic gradient descent. There is a paper on it: https://arxiv.org/abs/2405.16339 https://arxiv.org/abs/2405.16339 It seems to me that floating point math (matrix multiplication) will over time mostly disappear from ML chips, as Boolean operations are much faster both in training an inference. But currently they are still optimized for FP rather than Boolean operations.
- muxamilian 1y agoIn-memory computing (analog or digital). Still doing SIMD matrix multiplication but using more efficient hardware: https://arxiv.org/html/2401.14428v1 https://arxiv.org/html/2401.14428v1 https://www.nature.com/articles/s41565-020-0655-z https://www.nature.com/articles/s41565-020-0655-z
- gautamcgoel 1y agoThis is very interesting, but not what the Ironside TPU is doing. The blog post says that the TPU uses conventional HBM RAM.
- nsteel 1y agoThere's been some talk/rumour of next-gen HBMs having some compute capability on the base die. But again, not what they're doing here, this is regular HBM3/HBM3e. https://semiengineering.com/speeding-down-memory-lane-with-custom-hbm/ https://semiengineering.com/speeding-down-memory-lane-with-c...