5 ms·
In-mem (generally) means no (re)loading of data from a storage device.
by limit499karma 2y ago
In-mem (generally) means no (re)loading of data from a storage device.
- yjftsjthsd-h 2y agoSure, but I don't think that makes sense here; when I run an LLM on CPU, I load to memory and run it, when I run on GPU I load the model into the GPU's memory and run it, and I don't have anything like that much money to burn but I imagine if I used an FPGA then I would load the model into its memory and then run it from there. So the fact that they're saying "in-memory" in contrast to ex. GPU makes me think that they're talking about something different here.
- mmoskal 2y agoIt's a different kind of memory chip that also does some computation. See https://en.m.wikipedia.org/wiki/In-memory_processing https://en.m.wikipedia.org/wiki/In-memory_processing
- adrian_b 2y agoWhile this has been proposed repeatedly for many decades, I doubt that it will ever become useful. Combining memory with computation seems good in theory, but it is difficult to do in practice. The fabrication technologies for DRAM and for computational devices are very different. If you implement computational units on a DRAM chip, they will have a much worse performance than those implemented with a dedicated fabrication process, so for instance their performance per watt and per occupied area will be worse, leading to higher costs than for using separate memories and computational devices. The higher cost might be acceptable in certain cases if a much higher performance is obtained. However it is unavoidable that unlike with a CPU/GPU/FPGA, where you can easily reprogram the device to implement a completely different algorithm, a device with in-memory computation would be much less flexible, so it either will implement extremely simple operations, like adding to memory or multiplying the memory, which would not increase much the performance due to communication overheads, or it would implement some more complex operations, which might implement some ML/AI algorithm that is popular for the moment, but which would be hard to use to implement better algorithms when such algorithms are discovered.
- vlovich123 2y agoI suspect that the attempts to remove the DRAM controller and embedding it into the chips directly will succeed in meaningfully reducing the power per retrieval and increase the bandwidth by big enough that it’ll postpone these more esoteric architectures even though its pretty clear that bulk data processing like LLMs (and maybe even graphics) is better suited to this architecture since it’s cheaper to fan out the code than it is to shuffle all these bits back and forth.
- p1esk 2y agoIn-memory doesn’t mean in-DRAM. https://arxiv.org/pdf/2406.08413 https://arxiv.org/pdf/2406.08413
- adrian_b 2y agoSRAM does not have enough capacity to be useful for in-memory computation. The existing CPUs, GPUs and FPGAs are full of SRAM that is intimately mixed with the computational parts of the chips and you could not find any structure improving on that. All the talk about in-memory computing is strictly about DRAM, because only DRAM could increase the amount of memory from the up to hundreds of MB of memory that is currently contained inside the biggest CPUs or GPUs to the hundreds of GB that might be needed by the biggest ML/AI applications. All the other memory technologies mentioned in the paper linked by you are many years or even decades away from being usable as simple memory devices. In order to be used for in-memory computing, one must first solve the problem of making them work as memories. For now, it is not even clear if this simpler problem can be solved.
- p1esk 2y agoLet’s see: Mythic uses flash, d-Matrix uses SRAM. Encharge is the only one who uses capacitor based crossbars, but those are custom built from scratch and very different from any existing DRAM technology. Which companies are using DRAM for in-memory computing?
- fulafel 2y agoNot in the context of discussing hardware architectures. (Context in the abstract is "First, we present the accelerators based on FPGAs, then we present the accelerators targeting GPUs and finally accelerators ported on ASICs and In-memory architectures" and the section title in the paper body is "V. In-Memory Hardware Accelerators")