6 ms·
Samsung's Processing-in-Memory (PIM)
- plywoodShadow 1mo agoWhat about energy consumption? Wouldn't active cooling be needed for RAM as well as for CPU and GPU?
- topspin 1mo agoThe better question is: what is the net gain for the overall system? If PIM reduces the net thermal load and power consumption of the system for the same workload, then it’s a win regardless of where the heat sinks end up. The customers Samsung has in mind for this today are not limited to commodity designs. They’re using novel designs with each new hardware generation, so moving heat sinks around is not a deal-breaker.
- sehw 1mo ago[dead]
- reliabilityguy 1mo agoInteresting that Samsung still pursues PIM. IIRC they had a paper in ISCA21 or 22 where they showed HBM2 module with PIM, which back then impressed me quite a lot. That being said, I am not sure what’s the killer application for this technology, and without such application adoption is unlikely.
- xyzzy123 1mo agoAs I understand it, the killer app is llms. You could run MACs directly in RAM, offloading a lot of work from CPU and cutting down on insane (external) memory bandwidth required. Imagine (this is a fantasy pitch but potentially achievable for some use cases) wanting to run a larger llm and all you have to do is buy more RAM so it fits.
- embedding-shape 1mo ago> Imagine (this is a fantasy pitch but potentially achievable for some use cases) wanting to run a larger llm and all you have to do is buy more RAM so it fits. Isn't this how it works today already? Granted you wanted to run it on RAM rather than VRAM.
- petu 1mo agoYes, but running out of RAM is impractical due to low memory bandwidth. According to the article/Samsung RAM dies inside can support way higher bandwidth than they expose, they're limited by external interface / bus width: > Together, they can utilize the chip’s internal bandwidth across all 16 banks, which comes out to 614 GB/s. For comparison, regular DRAM accesses can hit two banks in parallel and max out at 76.8 GB/s. And that's just for single 64-bit IC. So way faster and more power efficient.
- dannyw 1mo agoYou can scale with more memory channels. Workstation/server platforms go up to 12 or 16 channels if I remember correctly. Consumer platforms have been stuck at dual channel for decades; most of it I attribute to intentional product segmentation. I'm hoping that LLMs might change eventually for an upcoming consumer platforms; going to 4 channel would be really nice.
- zerd 1mo agoHow do you scale past 16 channels though? 16 channels give you around 614 GB/s, while PIM can do that per chip, so it can achieve 58TB/s.
- amelius 1mo agoYou: "AI, please write me $COOL_APP." AI: "Sorry, all the hardware is made for running AI."
- eru 1mo ago
- netfortius 1mo ago> That being said, I am not sure what's the killer application for this technology... Build it, and they will come ;)
- deleted 1mo ago[deleted]
- imtringued 1mo agoThe idea is this: You have an eight socket server with 96 memory slots, you add 96x PIM memories into the server (optimistic), load all the LLM parameters or KV cache in RAM and exclusively let it perform GEMV and let it rip. 614 GB/s x 96 = 58,944 GB/s. Alternatively, the memory is used for embedded inference tasks. You can now upgrade from the limited single or two digit MB SRAM accelerators to reasonably fast single digit gigabyte models. Without MoE you could reach 100 tokens per second with an 8B fp8 model on a single channel. With MoE you might break 500 tokens per second.
- hoppp 1mo agoAnd now the 96 memory slots need individual cooling.
- throwaway173738 1mo agoWhich might be easier since the surface is larger
- reliabilityguy 1mo ago> load all the LLM parameters or KV cache in RAM and exclusively let it perform GEMV and let it rip. Won’t you have a bunch of extra reads/writes via the CPU because these DIMMs won’t be able to compute matrix multiplications?
- LogTrim 1mo ago[flagged]
- johnnyApplePRNG 1mo ago[flagged]
- consp 1mo agoSo you basically dispose of cache for the memory region used? I wonder what the offsets of the cache misses is going to be in practice (the article addresses it but there is no solution/impact given by samsung).
- manmal 1mo agoI guess you also get very high bandwidth that way? I‘m not sure that would come for free though.
- WithinReason 1mo agoIt comes at a cost of a fragmented memory space, which is fine for some applications, like LLMs
- PunchyHamster 1mo agoif your working set fits in cache PIM is irrelevant
- yvdriess 1mo agoCaches are important for cpu core performance even when the the working set doesn't fit.
- skissane 1mo ago> So you basically dispose of cache for the memory region used? I wonder what the offsets of the cache misses is going to be in practice (the article addresses it but there is no solution/impact given by samsung). This is a temporary issue. JEDEC's LPDDR6-PIM is going to add defined commands for Processing-in-Memory operations. Once there are standardised commands, it will be possible for the CPU vendors to make the CPU cache aware of what is happening. Of course, that doesn't solve it for this generation of the technology. But I think this generation is more of a demo for early adopters to gain experience with it. It will likely take a few years for all these issues to be solved, but there is no principled reason why they can't be.
- userbinator 1mo agoIn-memory computation was already possible with regular DRAM: https://news.ycombinator.com/item?id=22712811 https://news.ycombinator.com/item?id=22712811 Add a new set of CPU instructions like “rep macb” ...and it's been long enough now, that I can say there was an effort to implement this on standard x86 memory controllers and have the existing string instructions do so, back in the days of SDR SDRAM, but the tradeoffs weren't (yet) in favour.
- PunchyHamster 1mo agothat's not at all comparable, you're still paying memory latency and not getting any extra bandwidth
- saejox 1mo agoIf we could buy a 64gb stick and run a 32b model with 30tps on it. This could sell
- Torkel 1mo ago[dead]
- bhouston 1mo agoI wrote up a theoretical post here about LMM performance of a MacBook Pro with PIM memory: https://ben3d.ca/blog/m5-max-samsung-lpddr5-pim-650-tokens-per-second https://ben3d.ca/blog/m5-max-samsung-lpddr5-pim-650-tokens-p...
- asaddhamani 1mo agoInteresting read. Definitely hoping this works out so we can have cheap LLM machines at home
- deleted 1mo ago[deleted]
- pragma_x 1mo agoWhat I find amusing about moving compute to a RAM bank is it _almost_ resembles where we were with ISA-based extended RAM back in the 1980's. Some cards featured a CPU that took over the whole system and/or functioned like an upgrade. Others were a "computer on a card" that provided other features. I think this goes to show how cyclic tech can be. So, something like Samsung's invention here might have gained traction, as overcoming the slow PC ISA bus would have been a huge accelerator, kind of like where we are now.
- vardump 1mo agoMaybe we can connect multiple Samsung RAM sticks to a grid and call it RAMsputer.
- mr_toad 1mo ago“ Each PIM block only has fast access to its locally attached DRAM bank. All other input data has to be brought in through the DRAM chip’s comparatively constrained external interface. PIM blocks can’t directly exchange data with each other, so the host has to move data using regular DRAM reads and writes if one PIM block needs to use results generated by another.” So how big are these banks? If you can’t fit the weights of a layer into one bank then presumably you lose a lot of the speed gains.
- londons_explore 1mo agoNot necessarily. If your weights have to go across 2 banks, you just have to split and transfer the input and output vectors, which are much smaller.
- deleted 1mo ago[deleted]
- imtringued 1mo agoFor GEMV you lose nothing. The reason is quite simple. You can split the matrix along both dimensions so you just tile it into 64x64 or whatever fits into the bank and just fill it up. The biggest problem is load balancing the tiles across all banks for maximum parallelism. For GEMM I believe there is no point in doing PIM, you are better off with a GPU or NPU.
- londons_explore 1mo agoWhilst processing in memory is clearly the future, I am unconvinced by this implementation. Matrix multiplication involves getting every entry of the input and output matrices to be at the same multiplier at the same time. (Ie. N^2). To do that, a lot of data movement needs to happen. Movement is the main thing - the multiplication and addition is a sideshow as far as energy and silicon space is concerned. You need a 'around the chip' ring shift register to pass every element of one matrix past every element of the other.
- zozbot234 1mo ago"Movement is the main thing" is precisely why pursuing compute-in-RAM makes some sort of sense to begin with. But DRAM fabrication processes are quite specialized and do not perform well with pure compute logic. The overall profile of this thing will arguably be similar to a rather weak NPU, though with much better memory bandwidth - one key limitation, as with NPUs, will be the bespoke programming model and lack of support for the latest compressed/quantized number formats, which heavily limits the usefulness of being able to access memory directly. GPUs, even weak iGPUs, can dequantize/pad parameters on the fly which adds a lot of flexibility - and expose standard, well understood compute capabilities via CUDA, Metal or Vulkan. This is not quite comparable unfortunately.
- rivetfasten 1mo agoI was under the impression they're doing a chiplet thing to get heterogenous processes for that reason.
- ACCount37 1mo agoIf what we want to do with this is make cheap QKV sweeps, then "a weak NPU with a lot of mem bandwidth" seems good enough? Exactly the tool for that job, and nothing else. Also spares us the trouble of dealing with weights. By the time we're in QKV realm, the weights have already weighted.
- danmaz74 1mo agoGiven how important matrix multiplication with a huge number of fixed parameters is becoming, there is an enormous incentive to design much more efficient architectures where this very simple compute is colocated with memory. Inference cost would come down a lot.
- ginko 1mo agoFeels like the most realistic/short-term way to make use of this would be to set up some barebones RTOS to run from CPU cache with the PIM memory being used for compute only and use the device as a network attached accelerator.
- krater23 1mo agoSelf changing RAM and a complex way to interact with it in software. A new security nightmare is emerging.
- intrasight 1mo agoIt doesn't suffer from the inherent security issues of the von Neumann architecture. Memory will only be data.
- OptionX 1mo agoSo instead of putting more cache on the cpu you just put the cpu on the cache.
- dotancohen 1mo agoSort of, but the cache architecture is replaced with memory architecture. But you could look at it either way.
- harshaw 1mo agoThis is somewhat orthogonal to the article, but the whole bubble on AI data centers seems to presume that the need for compute is so massive that it far exceeds the expected optimizations we would expect with at scale inference (PIM, ASICs, etc). I would expect that there is a set of optimizations like this one (or variations) that would someone negate the buildout. But it's not really discussed.
- Tenoke 1mo agoThere's been a ton of optimizations already, it hasn't remotely reduced demand even temporarily. More efficiency just makes the compute have even higher ROI per $ and watt spent.
- roryirvine 1mo agoWith sufficient optimisation, there ought to be a tipping point beyond which local inference is good enough. And, sure, datacentre compute will still be needed for training but one of the biggest current uses will begin to taper off. The question really is how soon we reach that tipping point, and whether it's before or after the current bubble runs out of steam for some other reason.
- Tenoke 1mo ago>there ought to be a tipping point beyond which local inference is good enough There's no such ought really. Even at current levels you'd need like a 100x gain from here to approach current top proprietary models (probably a lot more for say Mythos or Mythos 2), and it's not like they are stoppng to improve. This is before we even account that you'd just be running 1 agent then, and not a swarm like you'd be able to in the cloud or that you can do only so much compression before you are losing out
- eureka7 1mo ago> (probably a lot more for say Mythos or Mythos 2) Not everyone needs that large of a model, though.
- hham 1mo agomemory plus memory bandwidth > GPU, this is the gist of it.
- samuelknight 1mo agoI saw them present a similar concept at Hot Chips in 2020 or 2021. It's still a cool idea, however people should remember that there are like 20 of these exotic accelerators designs pitched at trade shows every year that go nowhere.
- p1esk 1mo agothat go nowhere But some do end up in the industry: Mythic AI, Encharge AI, d-Matrix.
- tesnorindian 1mo agoIn memory compute has been in talks since LLMs took up. I remember few flocks were trying to get RISC V cores in the memory like these papers https://arxiv.org/abs/2602.01827 https://arxiv.org/abs/2602.01827
- fulafel 1mo agoIt has been a prominent "next major shift" idea in computer architecture since the 90s, to deal with the memory wall. Eg David Patterson advocating it in the 1990 and 1997 articles. or CRAM [1]. [1] https://www.eecg.toronto.edu/~stumm/Theses/Elliott-PhD98.pdf https://www.eecg.toronto.edu/~stumm/Theses/Elliott-PhD98.pdf
- bob1029 1mo agoThe tradeoff with putting the compute in the memory is that you have to know exactly where the dependent information will be at all times. Most problems do not fit this pattern very well. AI, gaming and crypto being the most obvious exceptions. It is incredibly constraining to develop applications using specialized hardware like this. You might as well spin out an ASIC for whatever it is you are doing. All 3 applications noted above eventually got their own flavors. I think the Von Neumann bottleneck is mostly a feature. The fact that communication of information across distances is expensive should not be immediately assumed to mean that it is universally flawed to do this. You are paying for something when you use all those joules. I'd argue we are usually wasting our energy with regard to information communication (e.g., lighting up a network interface & copper because we couldn't be bothered to use SQLite), but other times this stuff is fundamentally required for practical solutions to exist.
- Eridrus 1mo agoIn a world where AI is writing all of the code, the difficulty of the task may no longer be a blocker.
- trollbridge 1mo agoIndeed. SIMD type code used to be too hard for me to write very often, but now it’s easy to crank out tons of it.
- jeffbee 1mo agoThere are fundamental issues here and I think the article only touched on a few. On the software side this completely blows up the whole virtual memory concept. We will need different operating systems.
- zeusk 1mo agowhy would it? the parent OS can already handle physically contiguous allocations so these should be no different (with the exception that a separate interface can be used to do compute over these buffers/pages).
- howdyhowdy 1mo agoI thought DMR and Venice were supporting memory encryption by default. With the keys living on the CPU side and no standards for key sharing, I wonder how this will gain traction.
- rivetfasten 1mo agoThe discussion about cpu caching challenges etc seems like it's missing the point. Wouldn't it be more likely to DMA the results to a GPU anyway?
- HarHarVeryFunny 1mo agoI remember taking VLSI design as part of my Comp. Sci. degree at Bristol, UK c.1980, using the Conway & Mead book, and "Commingling of Processing and Memory" was mentioned even back then. Obviously you (eventually) need your data where the compute is, especially in a non-von-Neumann architecture where moving data around isn't an option even if you were OK with the performance drop. It seems kinda obvious that eventually AI will be implemented as low power dataflow custom chips integrating memory/state & compute, but who knows!
- BiraIgnacio 1mo agoVery cool paradigm, I wonder if it's easier to have processing moved/added to memory or have more memory in the CPU. In either case, as you said, it would be having both data and compute closer than they are right now.
- throwaway173738 1mo agoMight as well just go whole hog and change the entire computer architecture, then. A lot of the arguments against this change boil down to computers and software code don’t work well with this today.
- nottorp 1mo agoIs this about the fake craters in moon photos?
- artyomsv 1mo ago[dead]
- sergq 1mo ago[flagged]
- latchkey 1mo ago2021 https://patents.google.com/patent/US11600340B2/en https://patents.google.com/patent/US11600340B2/en https://x.com/xennygrimmato_/status/2025376089607209218 https://x.com/xennygrimmato_/status/2025376089607209218
- glitchbot 1mo agoThat game is still going? I have a char from 06!
- sciencesama 1mo ago10 years ago hpe labs had something similarly including the operating system ! and the cto got ousted and a new ceo camea nd the whole new direction is always networking !
- rubin55 1mo agoIt sounds to me like something HP was betting on a while ago: Memristors[1]. Afaik, it never went anywhere, but I thought the ideas were pretty remarkable at the time. [1]: https://en.wikipedia.org/wiki/Memristor https://en.wikipedia.org/wiki/Memristor
- emsign 1mo agoIt's funny how the article just assumes PIM be for computing weights of a model without even mentioning it in the introduction.