4 ms·
Your going to need a lot more than a few, 800G VRAM needed
by foxhop 2y ago
Your going to need a lot more than a few, 800G VRAM needed
- TechDebtDevin 2y agoOof.
- glitchc 2y agoChrist!!
- lolinder 2y agoQuantized to 4 bits you'll only need ~200GB! 5 4090s should cover it.
- angoragoats 2y agoYou'll probably need 9 or more. 4090s have 24GB each.
- woodson 2y agoI wonder if AutoAWQ works out of the box, given no architectural changes (?). That would be most straightforward together with vLLM for serving.
- pat2man 2y agoTwo 128gb Mac studios networked via thunderbolt 4?
- Teknomancer 2y agoThis is actually a promising endeavor. Id love to see someone try that.
- angoragoats 2y agoThere's already at least one project that attempts this: https://github.com/exo-explore/exo https://github.com/exo-explore/exo
- downvotetruth 2y agoIf an implementation had NVidia's Heterogeneous Memory Management implemented, then 192 GB RAM DDR5 + GPU VRAM would seem to be close.
- AaronFriel 2y agoIf previous quantization results hold up, fp8 will have nearly identical performance while using 405GiB for weights, but the KV cache size will still be significant. Too bad, too, I don't think my PC will fit 20 4090s (480GiB).
- knicholes 2y agoI've got a motherboard that will support 8!
- Zambyte 2y ago40,320 4090s?? What witchcraft is this?! :D
- sebastiennight 2y agoAll the more impressive when you realize that Groq's infrastructure (based on LPUs) was built using only 6!
- beeboobaa3 2y agohow is this even useful? no one can run it.
- jermaustin1 2y agoYou don't use the 405B parameter model at home. I have a lot of luck with 8B and 13B models on a single 3090. You can quantize them down (is that the term) which lowers precision and memory use, but still very usable... most of the time. If you are running a commercial service that uses AI, you buy a few dozen A100s, spend a half million, and you are good for a while. If you are running a commercial inferencing service, you spend tens of millions or get a cloud sponsor.
- beeboobaa3 2y agoI can't expect all my users to have 3090s and if we're talking about spending millions there are better things to invest in than a stack of GPUs that will be obsolete in a year or three.
- jermaustin1 2y agoNo, but if you are thinking about edge compute for LLMs, you quantize. Models are getting more efficient, and there are plenty of SLMs and smaller LLMs (like phi-2 or phi-3) that are plenty capable even on a tiny arm device like the current range of RPi "clones". I have done experiments with 7B Llama3 Q8 models on a M3 MBP. They run faster than I can read, and only occasionally fall off the rails. 3B Phi-3 mini is almost instantaneous in simple responses on my MBP. When I want longer context windows, I use a hosted service somewhere else, but if I only need 8000 tokens (99% of the time that is MORE than I need), any of my computers from the last 3 years are working just fine for it.
- loudmax 2y agoIf you want to run the 405B model without spending thousands of dollars on dedicated hardware, you rent compute from a datacenter. Meta lists AWS, Google and Microsoft among others as cloud partners. But also check out the 8B and 70B Llama-3.1 models which show improved benchmarks over the Llama-3 models released in April.
- whalesalad 2y agofollow the trail of tears to my credit card