4 ms·
CEO of an AI startup & former AWS employee here. The cloud sucks for training AI models. It's just insanely overpriced in a way that no "Total Cost of Ownershi
by mattmireles 6y ago
CEO of an AI startup & former AWS employee here.
The cloud sucks for training AI models. It's just insanely overpriced in a way that no "Total Cost of Ownership" analysis is going make look good.
Every decent AI startup––including OpenAI––has made significant investments in on-premise GPU clusters for training models. You can buy consumer-grade NVIDIA hardware for a fraction of the price that AWS pays for data center-grade GPUs.
For us in particular, the payback on a $36k on-prem GPU cluster is about 3-4 months. Everything after that point saves us ~$10k / month. It's not even close.
When I was AWS, I tried to point this fact out to the leadership––to no avail. It simply seemed like a problem they didn't care about.
My only question is why isn't there a p2p virtualization layer that lets people with this on-prem GPU hardware rent out their spare capacity?
- blueblisters 6y agoAre TPUs too application specific to replace GPUs? It seems cloud TPUs could be competitive with GPUs in terms of $ per number of target epochs, provided you can do data parallelism for your workloads. Also, IBM offers bare metal pricing which is somewhat cheaper than virtualized instances attached to GPUs (and faster too). I think GPU virtualization is not quite there yet because Nvidia does not give access to core GPU functionality needed for efficient virtualization - you're stuck with using their closed-source libraries and drivers.