3 ms·
I have just bought 4 servers with 10 GPU (Titan X Pascal) and 384 GB of RAM for the price of a macbook. I found a solar powered CoLo facility to put them in. Th
by dougSF70 3y ago
I have just bought 4 servers with 10 GPU (Titan X Pascal) and 384 GB of RAM for the price of a macbook. I found a solar powered CoLo facility to put them in. The economics of owning vs renting from AWS are worth it. My server set up is the same price as running one model on AWS.
- crabbone 3y agoIf memory serves, NVidia's drivers can only be used for Tesla / Ampere family, if you are using them in a datacenter. Titan / GeForce are not allowed. So, there's that (but maybe I'm missing something). Now, "running" may also mean different things. For example, you may want to do things s.a. performance diagnostics (in order to understand if your code uses resources efficiently), and then you'd need stuff like NVML, DCGM and co. Consumer-grade hardware might not be supported, or might not in principle support diagnostics collection / instrumentation. Or, if you have multiple GPU-dependent workloads that cannot saturate your resources -- you might think of MIG as being a way to address that... and, again, consumer-grade GPUs won't help you here... I'm not saying you shouldn't try self-hosting. I'm actually all for it. But, you also need to be mindful of the pros and cons. NVidia must have some reason to pitch the datacenter family of GPUs to, well, datacenters. They aren't just blowing up prices. Also, to give some sense of comparison: Titan X ~= 3.5K CUDA cores. V100 ~= 5K CUDA cores. A100 ~= 7K CUDA cores. H100 ~= 18.5K CUDA cores. It probably doesn't translate directly into H100 being six times as fast as Titan X, but, I hear that these GPU workloads might be lengthy...
- dougSF70 3y agoI bought these from a Chinese internet firm. They dismantled their data center and I am buying it.
- dougSF70 3y agoAlso, Datacenter GPUs only have passive cooling, presumably to allow for more Cuda cores. I get that what I have is older tech but for the cost of one H100 GPU I can have 60 of these servers (600 GPUs) plus some change to pay for CoLo fees.
- crabbone 3y agoI'm not trying to discourage, and I don't have concrete numbers on hand, but there are other factors, beside the price of h/w. Like, obviously, electricity use, data locality / moving it around (with more smaller units you'd have to move it more, also, not sure if consumer-grade GPUs support NVLink, and even if they do, then at what bandwidth?) For some of these, you could obviously pay with your time. Sometimes that time is very valuable, and sometimes you have a lot to spare. Also, the amount of VRAM (3x)... Sometimes having too little of it means having to re-write the program, or it could dramatically impact the speed. Similarly for bandwidth (8x). But, again, if the kind of workload you have isn't constrained by either, then you could probably win by running on more smaller / older GPUs.
- iJohnDoe 3y agoCan you please share where to buy this from? I would like to purchase one. Thanks!
- _boffin_ 3y agoUhh… can you share details… this seems like price point of things when they’ve fallen off the back of a truck.