4 ms·
Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fatho
by topspin 2mo ago
Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now.
One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.
- mdp2021 2mo ago> decade after decade That can only work when there is physical capacity for improvement though. > underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So > * That will yield what it has always yielded[:] orders of magnitude jumps in capacity and performance* That will yield new and renewed hardware technologies. (Already the distinction between SRAM and DRAM was overly specialistic before this boom - now it's on our mind as we know we need to "expand", "make cheap", "integrate" or find alternatives.)
- topspin 2mo ago> That can only work when there is physical capacity for improvement though. There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered. Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn't been a use case for really dense, high performance ROM. Now there is. ROM used to be a big deal in computing and media (cartridges, optical disks, etc.,) but that tapered off long ago; volatile and R/W storage was sufficient and convenient for the time, and the inference model use case, where dense, high speed ROM can have extremely high value, didn't exist. Now there is a use case, and industry is thinking about something they haven't cared about in a long time. Current fabrication nodes, stacked in the third dimension à la NAND flash, could produce staggeringly dense, fast and low power ROM. That's why AMD snatched up Taalas: they're thinking about an aspect of the future that has been (reasonably) neglected.
- mdp2021 2mo agoSure, but actually, our current need is not really for "ROM": it is for "CiM", compute-in-memory - we want to minimize the data movement bottlenecks. That some implementations could be read-only is actually a disadvantage. Clearly there are possibilities, some of them proven (proof-of-concept, in-production etc.) - but taking for granted "Moore's law" like spaces for them may not be founded on what we know at this stage.
- deleted 2mo ago[deleted]
- topspin 2mo ago> our current need is not really for "ROM" A fast, low power ROM is the key ingredient to near term local inference with large models at low power. If I could offer you a $500 ROM that provided the model data for frontier inference on power similar to a desktop GPU, you would buy it, and consider it a bargain, even when it came time to pay another $500 for the upgrade.
- mdp2021 2mo ago> A fast, low power ROM is the key ingredient Surely it is clear to you that Read-Only /Memory/ does not /compute/, and our need is to compute through the data in the memory... That is CiM - a technology not that similar to ROM... Because a plain ROM does not solve problems in this area... In other words, > If I could offer you a $500 ROM that provided the model data Then I would have a physical token containing what I already had as a file, and the problem of running that file into something efficient would remain... Because the ROM does not "run" its contents...
- topspin 2mo ago> Surely it is clear to you that Read-Only /Memory/ does not /compute/ > Because the ROM does not "run" its contents... Conventional GDDR/HBM don't compute either, yet inference is implemented using these. Compute isn't the inference bottleneck. Inference requires high bandwidth, high capacity memory. The compute resources necessary are fungible, comparatively cheap and already available, at least for a small number of concurrent loads, such as in most local inference use cases. > Then I would have a physical token containing what I already had as a file I suspect you are not grasping what I mean by ROM. Dense, high performance ROM would not be the hardware equivalent of a "file", with performance bottlenecked by low bandwidth, high latency storage media, serialized for RW coherence reasons. It would have extremely high bandwidth, on par with GDDR, low latency due to a dedicated high performance bus, high concurrency due to a lack of any RW coherence obligations, and operate at low power (no gate leakage, no dynamic refresh,) and low cost compared to equivalent GDDR/HBM capacity. Essentially what high performance ROM would provide is high capacity, low power HBM, albeit read-only. At that point all you need is sufficient TOPS to run the inference algorithm. The compute part is already available, affordable and readily scales up and down as per performance/cost/power budgets.