4 ms·
"If AI inference remains as desirable as Kimball expects, the evolution is likely to follow the same trajectory as the CPU. The CPU didn’t improve along a singl
by aschla 17d ago
"If AI inference remains as desirable as Kimball expects, the evolution is likely to follow the same trajectory as the CPU. The CPU didn’t improve along a single axis but instead across simultaneously. Once transistor scaling slowed, chip and system architecture innovations of all kinds proliferated. The list of individual innovations that led to today’s ubiquitous, powerful personal compute could fill dozens of books. A few decades from now, the history of AI inference innovation will show similar depth."
Of the areas mentioned in the article, which are the most likely to have the most prominent innovative impact, and what will they entail?
- jjtheblunt 17d ago> The list of individual innovations that led to today’s ubiquitous, powerful personal compute could fill dozens of books. https://a.co/d/0dMk8urP https://a.co/d/0dMk8urP which is the Hennesey and Patterson computer architecture book would serve the role of the "dozens of books" hyperbole rather well.
- rhdunn 17d agoOne of the biggest limitations right now is memory capacity (storing large models/contexts in memory) and bandwidth (transferring the relevant data/weights to the silicon that is performing the operations on that data). This would cover things like: 1. having more memory on the card/chip and/or faster access to that memory; 2. integrated memory and compute units optimized for matrix and vector multiply add operations; 3. optimized load circuitry to e.g. read memory in the stride and span (next row, next column) access patterns common to matrices or ensure that no/few parts of the chip are stalled waiting on data or operations to complete. Another aspect is quantizations. These are similar to SIMD vector operations in that you are performing an operation on a block of n-bit data values at the same time, so can have optimized circuitry. For 2 or 3 valued quantizations you can reduce various addition and multiplication operations to logic operations, avoiding circuitry for things like the half-adder, full-adder, and carry-lookahead. Then there's adding specific circuitry for common operations such as ReLU like is done in hardware acceleration of image, video, etc. processing. There's a trade off here as optimized hardware would perform better at the specific operations but if those are too specific then they can't be used by different/newer model architectures. (Though it does make sense to try and optimize common operations/logic where possible.) It would be interesting to see if these designs can/will benefit training as well, as that would bring down the time/cost/energy of training large models as well as making it easier for local fine-tuning.
- ip26 17d agoA large fraction of the innovation in CPUs is driven by working around the memory wall. I anticipate AI inference will follow the same trend, and innovations that work around the autoregressive nature will be enormously impactful. Speculative decoding is an example. An accurate draft model can reduce the number of times you stream through memory by a factor of 4x.
- nixon_why69 17d agoSpeculative decoding is software, though. How do you work around the memory wall when you're going to have to stream all weights, no matter what? Latency-hiding tricks don't matter when you're bandwidth constrained.
- searealist 17d ago2x is more typical, and only for dense models, not the sparse models all frontier labs use.
- jononor 16d agoSome kind of compute-in-memory architecture is a good candidate, I think. There are many alternatives here, researched for many years prior to the LLM craze. However economies of scale dominate in chip industries, and this tends to favor more conventional or incremental approaches (to piggyback on existing scale). Alternatively someone needs to have a way of bootstrapping the insane scales needed to be competitive with a better-but-different approach. So it could be that boring and straightforward stuff like two-chip prefill+decode takes most.