5 ms·
Lanes will probably be an issue, so a threadripper pro or an epyc cpu, add half a grand at least for the motherboard and it’s starting to look grim.
by PartiallyTyped 3y ago
Lanes will probably be an issue, so a threadripper pro or an epyc cpu, add half a grand at least for the motherboard and it’s starting to look grim.
- thfuran 3y agoAnd that's before you even get your first power bill.
- PartiallyTyped 3y agohey, at least you will cut down on the heating costs!
- easygenes 3y agoFor LLM applications, the performance loss when power limiting 3090 to 200w is fairly low and you get peak perf/w.
- yumraj 3y agoSo even with power limiting, with 4 3090s, you’re looking at 800w from GPUs alone. So about 1000w give or take. Yes? M2 Ultra [0] seems to be max 295w [0] https://support.apple.com/en-us/HT213100 https://support.apple.com/en-us/HT213100
- easygenes 3y agoYeah, but watt for watt the 3090s will output more tokens, as a single 3090 has more memory bandwidth than an M2 Ultra. That's the main performance constraint for LLMs. Dramatically oversimplifying of course. There will be niches where one will be the right choice over the other. In a continuous serving context you'd mostly only want to run models which can fully fit in the VRAM of a single 3090, otherwise the crosstalk penalty will apply. 24GB VRAM is enough to run CodeLlama 34B q3_k_m GGUF with 10000 tokens of context though.