4 ms·
You don't have to go that far down the page to see it is paging to system RAM: Requirements: 380GB CPU Memory 1-8 ARC A770 500GB Disk
by Cheer2171 2y ago
You don't have to go that far down the page to see it is paging to system RAM:
Requirements:
380GB CPU Memory
1-8 ARC A770
500GB Disk
- superkuh 2y agoYep. That's why the headline is incorrect. 380GB of the model on CPU system RAM and 32GB on some ARC GPUs. The ratio, 380/32, is obvious. Most of the processing is being done on the CPU. The GPU are little bit icing in this context. Fast, sure, but having to wait for the CPU layers (that's how layer splits work with llama.cpp). I think changing the end of headline to "Xeon w/380GB RAM" would stop it from being incorrect and misleading.
- Cheer2171 2y ago"with" does not mean "entirely on" Edit: but what you added in your edit is right, it would be more accurate to append the system ram requirement
- ryao 2y agoWhat if it does not need to read from system RAM for every token by reusing experts whenever they just happen to be in VRAM from being used for the previous token? If the selected experts do not change often, this is doable on paper.
- hmottestad 2y agoThat’s probably the main performance benefit of using the GPU. If you’re changing the active expert for every single token then it wouldn’t be any faster than just running it on the CPU. Once you can reuse the active expert for two tokens you’re already going to be a lot faster than just the CPU. More GPUs let you keep more experts active at a time.
- hexaga 2y agoExpert distribution should be approximately random token-by-token, so not likely.
- superkuh 2y agoThat's not how llama.cpp works. It's a layer split. The GPUs handle a few layers and the CPU handles the rest. The GPU layers no matter how fast they complete still have to wait on the CPU layers.
- colorant 2y agoThe ipex-llm implementation extends llama.cpp and includes additonal CPU-GPU hybrid optimizations for sparse MoE