6 ms·
This paper is light on background so I’ll offer some additional context: As early as the 90s it was observed that CPU speed (FLOPs) was improving faster than m
by refibrillator 2y ago
This paper is light on background so I’ll offer some additional context:
As early as the 90s it was observed that CPU speed (FLOPs) was improving faster than memory bandwidth. In 1995 William Wulf and Sally Mckee predicted this divergence would lead to a “memory wall”, where most computations would be bottlenecked by data access rather than arithmetic operations.
Over the past 20 years peak server hardware FLOPS has been scaling at 3x every 2 years, outpacing the growth of DRAM and interconnect bandwidth, which have only scaled at 1.6 and 1.4 times every 2 years, respectively.
Thus for training and inference of LLMs, the performance bottleneck is increasingly shifting toward memory bandwidth. Particularly for autoregressive Transformer decoder models, it can be the dominant bottleneck.
This is driving the need for new tech like Compute-in-memory (CIM), also known as processing-in-memory (PIM). Hardware in which operations are performed directly on the data in memory, rather than transferring data to CPU registers first. Thereby improving latency and power consumption, and possibly sidestepping the great “memory wall”.
Notably to compare ASIC and FPGA hardware across varying semiconductor process sizes, the paper uses a fitted polynomial to extrapolate to a common denominator of 16nm:
> Based on the article by Aaron Stillmaker and B.Baas titled ”Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7nm,” we extrapolated the performance and the energy efficiency on a 16nm technology to make a fair comparison
But extrapolation for CIM/PIM is not done because they claim:
> As the in-memory accelerators the performance is not based only on the process technology, the extrapolation is performed only on the FPGA and ASIC accelerators where the process technology affects significantly the performance of the systems.
Which strikes me as an odd claim at face value, but perhaps others here could offer further insight on that decision.
Links below for further reading.
https://arxiv.org/abs/2403.14123 https://arxiv.org/abs/2403.14123
https://en.m.wikipedia.org/wiki/In-memory_processing https://en.m.wikipedia.org/wiki/In-memory_processing
http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration.TechScale/ http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration.Tech...
- bilsbie 2y agoThanks for the background. Whatever happened to memristors and the promise of memory living alongside cpu?
- deleted 2y ago[deleted]
- dewarrn1 2y agoThat's funny, I had thought that memristors were a solved problem based on this talk from a while back (2010!): https://www.youtube.com/watch?v=bKGhvKyjgLY https://www.youtube.com/watch?v=bKGhvKyjgLY, but I gather HP never really commercialized the technology. More recently, there does seem to be interest in and research on the topic for the reasons you and the GP post noted (e.g., https://www.nature.com/articles/s41586-023-05759-5 https://www.nature.com/articles/s41586-023-05759-5).
- tonetegeatinst 2y agoI believe asianometry did a YouTube video on memristors....might be worth watching.
- Lerc 2y agoOr even an architecture akin to an atonishingly large number of RP2050's. It does seem like it would work well for certain types of nnet architectures. I've always been partial to the idea of two parallel surfaces with optical links, Make a connection machine style hypercube where the bit of the ID of every processor indicates its location in the hypercube. Place all of the even parity CPUs on one surface and the odd parity CPUs on the other surface, every CPU would have line of sight on its neighbour in the hypercube (as well as the diametrically opposed CPU with all the ID bits flipped)
- phh 2y ago> Or even an architecture akin to an atonishingly large number of RP2050's. Groq and Cerebras are probably that kind of architecture
- iml7 2y agoIt came and went in the form of optane.
- chatmasta 2y agoI'm a layman on this topic, so I'm definitely about to say something wrong. But I recall an intriguing idea about a sort of "reversion to analog," whereby we use the full range of voltage crossing a resistor. Instead of cutting it in half to produce binary (high voltage is 1, low voltage is 0), we could treat the voltage as a scalar weight within a network of resistors. Has anyone else heard of this idea or have any insight on it?
- nickpsecurity 2y agoThey mostly failed in the market. I have a list of them here: https://news.ycombinator.com/item?id=41069685 https://news.ycombinator.com/item?id=41069685 I like the one that’s in RAM sticks with an affordable price. I could imagine cramming a bunch of them into a 1U board with high-speed interconnects. Or just PCI cards full of them.
- _zoltan_ 2y agoWhile this might have been true for a while before 2018, since then 400GbE ethernet became the fastest adapted interconnect, and today 1.6Tbit interconnects exist. PCI-e V4 came and went so fast that it lived maybe 2 years. NVMeOF has been scaling with fabric performance and it's been great. 400GB/s interconnect on the H100 DGX today.
- fsndz 2y agoTrue. Samsung's Dr. Jung Bae Lee was also talking about that recently. "rapid growth of AI models is being constrained by a growing disparity between compute performance and memory bandwidth. While next-generation models like GPT-5 are expected to reach an unprecedented scale of 3-5 trillion parameters, the technical bottleneck of memory bandwidth is becoming a critical obstacle to realizing their full potential." https://www.lycee.ai/blog/2024-09-04-samsung-memory-bottleneck-gpt5 https://www.lycee.ai/blog/2024-09-04-samsung-memory-bottlene...