5 ms·
(Author here) I agree with a lot of your points, especially the fact that TCO is practically impossible to calculate, is highly subjective, and should ideally
by sabalaba 8y ago
(Author here)
I agree with a lot of your points, especially the fact that TCO is practically impossible to calculate, is highly subjective, and should ideally include opportunity costs which, almost by definition vary from person to person.
I've definitely thought through some of your points. First I want to say that if your load is spiky you have no business buying hardware. No matter how you slice it, you're simply not going to be able to get hundreds of GPUs of throughput with an 8 GPU machine so it's not possible to get the same "product". So for those doing short bursts of hyperparameter search, stick with cloud.
But, like you said, for those with high utilization and known utilization patterns, it makes a lot of sense to go on-prem. Let's just say pretty much anybody who is buying a reserved instance from AWS should consider buying hardware instead.
>- Cost associated with finding an admin who understands how this thing works
Pricing based on what co location facilities will provide you without much search and is pretty generous. $10k per year per server is silly high and should cover that search cost.
> - Try before you buy
You can try essentially this exact machine on AWS. That instance and hardware are almost exactly the same. Now of course this doesn't apply to many instance types.
> - Time it takes for the box to be built, shipped and sent to the data center
5 business days is our mean lead time for that unit. Most workstations are 2 days.
> - Time (and cost) it takes to install software, drivers, etc
It comes with that stuff pre-installed as mentioned in the article. See Lambda Stack for driver / framework / CUDA woes: https://lambdalabs.com/lambda-stack-deep-learning-software https://lambdalabs.com/lambda-stack-deep-learning-software
> - 3-year? What is the useful life of this? When does it seemingly become obsolete?
I'll admit it's a long time but 3 years is about right for GPUs from my experience. I personally purchased new GPUs for my workstations in 2012 (Fermi), 2015 (Maxwell), 2017 (Pascal).
> - AWS will continually upgrade their hardware and you keep paying the same
Yes, but only if you never commit to a reserved instance, in which case costs are 3x higher. If you buy a reserved instance you don't get upgraded hardware.
> - Spending $90k instead of $184k in year 1 with the option to turn it off if you want (no longer need). This could be very valuable for a startup who wants elastic spending patterns.
Yea, most of these are people who are already using GPUs really often for internal training infra and are finding it expensive.
> - Returns, breakage, warranty in case of a hardware failure
Our hardware comes with a 3 year parts warranty. Of course, this doesn't pay for your time lost but it's pretty rare to see parts fail and the $10k / year more than covers in-colo swap outs.
I agree with all of your points. I don't think that we really disagree with much here. TCO is hard to calculate :).
- choppaface 8y agoK80s are 2014 and people still use them at scale. So 3 years isn't unreasonable, right? A major drawback of having this sort of hardware on-prem (and this unit in particular) is that it doesn't include local storage for a 100TB+ scale dataset. There's not just the admin cost of hard drives and RAID, but there's the infra cost of synchronizing data to the machine. An average, reasonable, product-oriented Engineering Manager could easily choose cloud over a $100k savings if it means less headaches. Especially since GCloud will give big discounts if they're trying to outbid Amazon. On the other hand, if Lambda offered some sort of NAS solution plus caching software (Alluxio! :), that would make the offering a lot more competitive. Oh and maybe throw in a network engineer to set up peering and such.. ;) What might make the comparison more compelling is to take anonymized workloads from your customers, examine how often the workload results in idle (wasted) time, and to factor that into the figures. If a user's workload is elastic (many are, especially in R&D), then the user would be drawn towards ML Engine, Sagemaker, Floydhub, Snark, etc... The downside to these services is that they almost always involve expensive copying of data from NAS to worker machines. But if user utilization is high, and the dataset is static, the on-prem machine is a prime investment. While the savings is a big and noteworthy number, iteration speed is of principle concern to readers here and likely a good segment of your customers. I love the Lambda offerings, but there are a lot of deep learning hackers who can't make the most of bare metal. The 1080TI server is probably the killer solution at 1/10th the cost. But NVidia legal blah blah blah :(
- evilmoo 8y ago> 5 business days is our mean lead time for that unit. Most workstations are 2 days. So this is an advert? Gotcha. I think if you called out at the top of the article then it would have been a little more honest.
- mbesto 8y agoThanks for the response! Yes you're right, it is difficult to calculate. I hope people reading my comment took it into the context of a larger view on the process to calculate TCO, as opposed to just directly scrutinizing your products/article.