3 ms·
To be fair, nobody has upgradable memory in any system that has enough memory bandwidth and compute power to run LLMs with decent performance. It might be inter
by wtallis 26d ago
To be fair, nobody has upgradable memory in any system that has enough memory bandwidth and compute power to run LLMs with decent performance. It might be interesting to compare against some decade-old x86 server or workstation stuffed full of LRDIMMs to reach 1.5–2TB of RAM, but the bandwidth would be only slightly faster than a desktop today with high-end DDR5: nowhere close to GPU bandwidth. So performance would still suck.
Designing for extreme expandability comes with pretty steep tradeoffs.
- LTL_FTC 26d agoTake a look at AMD’s 12-channel memory servers. The newer Epycs are up to 16-channels now, 1.6TB/s. Pretty great for inference.
- wtallis 26d agoSure, if you want to make a comparison where the price tags aren't the same order of magnitude, then a recent server is obviously going to be powerful. But since the baseline of this comparison is a laptop and several Thunderbolt SSDs, the kind of servers or workstations with 1.5–2TB of RAM that you can reasonably compare against would have to be the really old ones, barely new enough to support that much total RAM. And despite the theoretically high memory bandwidth of recent EPYC CPUs, approximately nobody who can afford one is doing LLM inference on them.