4 ms·
A server with 512GB of high-bandwidth GPU addressable RAM in a server is probably a six figure expenditure. If memory is your constrain, this is absolutely the
by InTheArena 2y ago
A server with 512GB of high-bandwidth GPU addressable RAM in a server is probably a six figure expenditure. If memory is your constrain, this is absolutely the server for you.
(sorry, should have specified that the NPU and GPU cores need to access that ram and have reasonable performance). I specified it above, but people didn't read that :-)
- Numerlor 2y agoA basic brand new server can easily do 512gb. Not as fast as soldered memory but it should be maybe mid to high 5 figures
- la_oveja 2y ago5 figures? can be done in 6k https://x.com/carrigmat/status/1884244369907278106 https://x.com/carrigmat/status/1884244369907278106
- InTheArena 2y agoThat's CPU only memory, not high bandwidth, and not addressable by the GPU.
- jeffbee 2y agoThere isn't anything particularly high-bandwidth about Apple's DDR5 implementation, either. They just have a lot of channels, which is why I compared it to a 24-channel EPYC system. I agree that their integrated GPU architecture hits a unique design point that you don't get from nvidia, who prefer to ship smaller amounts of very different kinds of memory. Apple's architecture may be more suited to some workloads but it hasn't exactly grabbed the machine learning market.
- buildbot 2y agoM3 Ultra has 819GB/s, and a single epyc cpu with 12 channels has 460GB/s. As far as I know, llama.cpp and friends don’t scale across multiple sockets so you can’t use a dual socket Turin system to match the M3 Ultra. Also, 32GB DDR5 RDIMMS are ~200, so that’s 5K for 24 right there. Then you need 2x CPUs, at ~1K for the cheapest, and you need 2, and then a motherboard that’s another 1K. So for 8K (more, given you need a case, power supply, and cooling!), you get a system with about half the memory bandwidth, much higher power consumption, and very large.
- adrian_b 2y agoPartial correction, an Epyc CPU with 12 channels has 576 GB/s, i.e. DDR5-6000 x 768 bits. That is 70% of the Apple memory bandwidth, but with possibly much more memory (768 GB in your example). You do not need 2 CPUs. If however you use 2 CPUs, then the memory bandwidth doubles, to 1152 GB/s, exceeding Apple by 40% in memory bandwidth. The cost of the memory would be about the same, by using 16 GB modules, but the MB would be more expensive and the second CPU would add to the price.
- buildbot 2y agoAh, I didn’t realize they’d upped the memory bandwidth to DDR5-6000 (vs 4800), thanks for the correction! The memory bandwidth does not double, I believe. See this random issue for a graph that has single/dual socket measurements, there is essentially no difference: https://github.com/abetlen/llama-cpp-python/issues/1098 https://github.com/abetlen/llama-cpp-python/issues/1098 Perhaps this is incorrect now, but I also know with 2x 4090s you don’t get higher tokens per second than 1x 4090 with llama.cpp, just more memory capacity. (All if this only applies to llama.cpp, I have no experience with other software and how memory bandwidth may scale across sockets)
- adrian_b 2y agoThe memory bandwidth does double, but in order to exploit it the program must be written and executed with care in the memory placement, taking into account NUMA, so that the cores should access mostly memory attached to the closest memory controller and not memory attached to the other socket. With a badly organized program, the performance can be limited not by the memory bandwidth, which is always exactly double for a dual-socket system, but by the transfers on the inter-socket links. Moreover, your link is about older Intel Xeon Sapphire Rapids CPUs, with inferior memory interfaces and with more quirks in memory optimization.
- KeplerBoy 2y agoaddressable is a weird choice of words here. CUDA has had managed memory for a long time now. You absolutely can address the entire host memory from your GPU. It will fetch it, if it's needed. Not fast, but addressable.
- p_ing 2y agoWindows has been doing this since what... the AGP era? Though this is a function of the ISA rather than the OS.
- Numerlor 2y agoAh seems like I remembered the CPU price for a higher tier CPU which can cost the 6k on their own. Thinking about it you can get a decent 256gb on consumer platforms now too, but the speed will be a bit crap and would need to make sure the platform ully supports ECC UDIMMs
- behnamoh 2y agoexcept that you cannot run multiple language models on Apple Silicon in parallel
- jeffbee 2y agoThat doesn't sound right. The marginal cost of +768GB of DDR5 ECC memory in an EPYC system is < $5k.
- InTheArena 2y agoGPU accessible RAM.
- numpad0 2y agomoot point if tok/s benchmark results are the same or worse.
- DrBenCarson 2y agoNot moot if you care about producing those tokens with the largest available models
- kjreact 2y agoAre the benchmarks worse? Running LLMs in system memory is rather painful. I am having a hard time finding benchmarks for running large models using system memory. Can you point me to some benchmarks you’re referring to?
- adrian_b 2y agoIn a dual-socket EPYC system, the memory bandwidth is higher than in this Apple system by 40% (i.e. 1152 GB/s), and the memory capacity can be many times higher. Like another poster said, 768 GB of ECC RDIMM DDR5-6000 costs around $5000. Any program whose performance is limited by memory bandwidth, as it can be frequently the case for inference, will run significantly faster in such an EPYC server than in the Apple system, even when running on the CPU. Even for computationally-limited programs, the difference between server CPUs and consumer GPUs is not great. One Epyc CPU may have about the same number of FP32 execution units as an RTX 4070, while running at a higher clock frequency (but it lacks the tensor units of an NVIDIA GPU, which can greatly accelerate the execution, where applicable).
- deleted 2y ago[deleted]
- energy123 2y agoWhat is the memory bandwidth to the CPU cores? Is it competitive with 8-channel DDR5 servers for non-GPU compute?