3 ms·
Imagine how many 4090s you could buy and run in a cluster for $15,000 though
by naillo 3y ago
Imagine how many 4090s you could buy and run in a cluster for $15,000 though
- J_Shelby_J 3y agoI hate to burst your bubble, but more than two 4090s is going to put you at BoM + labor costs around $15k. Especially if you have to upgrade your electrical and hvac.
- elorant 3y agoThey seem to have six, either 4090s or 3090s judging by the total amount of VRAM. How many more do you think you could get with $15k considering all the other hardware costs? I doubt you could make it more efficient at this price point.
- ColonelPhantom 3y agoThe Tinybox is planned to use AMD, not NVIDIA. The GPU they're using to build Tinybox is the 7900 XTX.
- heyitsguay 3y agoIt might be tough to make more efficient, but $15k seems exactly about the price of "stick 6 4090s in a decent box and throw in a couple grand for my troubles", versus any revolutionary hardware configuration. The way it advertises running fp16 Llama 70B feels a bit contrived too, given the prevelance of quantizing to 8 bit at minimum.
- elorant 3y agoIn my opinion the best hardware to run big models is to go and get a mac studio ultra. You have 192GB of unified RAM which can run pretty much every available model without losing performance. And it would cost you half that price.
- qwytw 3y ago> without losing performance But isn't M2 ULTRA over 20x slower than this thing? ~30 TFlops vs 738.
- gsuuon 3y agoAccording to this [1] article (current top of hn) memory bandwidth is typically the limiting factor, so as long as your batch size isn't huge you probably aren't losing too much in performance. https://finbarr.ca/how-is-llama-cpp-possible/ https://finbarr.ca/how-is-llama-cpp-possible/
- elorant 3y agoBy losing performance I meant you don’t need to quantize the model a lot since it fits in RAM. My bad for not clarifying it.