17 ms·
V100 Server On-Prem vs. AWS P3 Instance Cost Comparison
- purplezooey 8y agodamn, "includes hiring a part time system administrator".
- fisherjeff 8y agoMust be extremely part-time, for $10k/year total cost.
- Scoundreller 8y agoThey're hiring a share of colo staff. It's a comparison to AWS, so everything that you install/do/operate on that server is extra in both cases.
- Spivak 8y agoI mean $10k/yr/server doesn't sound too unreasonable. Just sounds really bad when you're only looking at one server. $100k/yr total compensation (i.e. 60k-ish salary) for someone babysit 10 servers isn't super unreasonable.
- yjftsjthsd-h 8y agoHeck, I'd take that job! ...seriously, any chance that's a real thing? Sounds better than what I do now.
- vidarh 8y agoYes, sort of. You can find people with small-ish number of servers that will pay stupid money to have someone on-call when you count on a per server basis. In practice it's a nice side-gig, but you'll tend to need several of them, as people do understand they're paying a premium to have you accessible, and do expect to pay (substantially) less per server if they have more of them. In practice this will tend to include out of ours availability and/or devops type work, not just low level sysadmin stuff or physical maintenance, as a lot of that can be farmed out to "remote hands" at the colo providers on hourly rates with 24/7 availability and will certainly cost a tiny fraction of that $10k/year.
- cferr 8y agoI make just a couple of k's more than that to babysit a dozen. It's a thing.
- vidarh 8y agoIf a single server requires anywhere close to that in sysadmin time per year, it is broken or the sysadmin in question is incompetent, or we're talking a supercomputer of much greater complexity than this thing. To me it seems like a very conservative estimate, or allowing for said sysadmin to provide a lot of valueadd services (e.g. devops type services) that you'll typically need for a cloud setup as well.
- bithavoc 8y agoBuy servers if you have stable workloads, otherwide rent virtual machines in the cloud.
- wenc 8y agoThe article also assumes 100% utilization on the cloud. I wonder if continual GPU-based training of DNN models is perhaps a fairly circumscribed use case? Most DNN model training workloads are lumpy and transient.
- ramraj07 8y agoPractically though the moment you use an instance more than 50% of the time AWS incentivizes buying the annual plan.
- NotAnEconomist 8y agoIf I'm reading the chart right, GPU servers on AWS are cheaper if you utilize them less than 65-75% of the time, and turn them off when not in use.
- bithavoc 8y agoI think so, if you can afford turning off a server during the night then it makes sense to take advantage of some sort of hourly billing which is what most clouds offer by default.
- deleted 8y ago[deleted]
- stephenbez 8y agoEC2 has per-second billing: https://aws.amazon.com/blogs/aws/new-per-second-billing-for-ec2-instances-and-ebs-volumes/ https://aws.amazon.com/blogs/aws/new-per-second-billing-for-...
- rocky1138 8y ago
- ilaksh 8y agoThis is just the most extreme example. AWS is just really expensive. If you want a VPS take a look at Digital Ocean or Linode.
- wenc 8y agoI think the expense is mostly because these are GPU instances, which are not yet commoditized. Unlike VMs, multi-tenancy on GPUs is just a little bit harder.
- ilaksh 8y agoRegular AWS EC2 is still a lot more expensive than the alternatives that I mentioned.
- Namidairo 8y agoThe 15 tables in the nvidia grid documentation kind of shows how much of a mess it is. Different resolutions and maximum tenants depending on card and license type, and you can't allocate resources homogeneously.
- vidarh 8y agoActually the savings on this server looks to me to be low compared to what I'd usually expect. You should use AWS for convenience, not cost. They're expensive for cloud services, and even the cheapest cloud providers are expensive compared to renting dedicated for all but the most transient workloads, and of dedicated hosting providers I only know Hetzner to get close to the costs I could get for renting colo space or doing truly on-prem hosting. Even then the only reason Hetzner is competitive is because I'm in London where space/power is expensive, and they're in Germany, where it is cheap (e.g. they rent out colo space as well, and prices are at 1/3 to 1/4 of what I've paid in London).
- coleca 8y agoHetzner rents out 1080Ti GPUs which are not available in most regions or from most cloud providers, hence the lower cost. This article refers to the much more expensive Tesla V100 GPUs. From what I understand the NVidia license for the 1080Tis prevents cloud providers from offering them for uses other than blockchain. Since AWS can't control what you actually do with it, they simply don't offer 1080Tis. Source: https://www.theregister.co.uk/2018/01/03/nvidia_server_gpus/ https://www.theregister.co.uk/2018/01/03/nvidia_server_gpus/
- hughesjo 8y agoIt's not fair comparison unless we are comparing all in costs that include ops.
- riku_iki 8y agoThey include ops cost for on-prem server
- deleted 8y ago[deleted]
- dc_gregory 8y agoI'm inexperienced in the hardware front, would that machine likely not break down under a large load for 3 years straight? Nothing is set aside for hardware failure etc.
- taspeotis 8y agoDepends what the warranty is, I think three years is standard and you can pay for five or seven. (You pay commensurately, however.)
- dc_gregory 8y agoAh, excellent point, a warranty is something I completely overlooked. Probably some value "lost" doing it yourself due to downtime if something goes wrong, but likely negligible.
- baroffoos 8y agoOur on prem server has been running for about 7 years on a high load. Its a little unpredictable but I would expect a server to last longer than 3 years
- iheartpotatoes 8y agoBut isn't that $69,000 peanuts compared to the cost to hire sysadmins who are on call 24/7 to swap out: RAM, fans, drives, power supplies, provision new images, etc.?? So you saved on cloud costs, but now you have the admin burden: 1-3 people for $200k each. No?
- badestrand 8y agoFrom the article: "Our TCO [(Total Cost of Ownership)] includes energy, hiring a part-time system administrator, and co-location costs"
- rossdavidh 8y agoIf you have only a single server, and you need 24/7 support, then probably it is true you don't want to hire a full-time sysadmin for one server. But, for the people who are doing this kind of thing, they probably have more than one, so the cost of sysadmin is spread across more than one server.
- iheartpotatoes 8y agoI thought ads on HN were discouraged?
- jim_bailie 8y agoMaybe not! I wonder how much if anything was paid for this placement. And I also wonder if I should be preparing an "article" about my company. We could sure use some extra exposure. Anyway, despite my belly-aching, it was an interesting read.
- warent 8y agoYC has nothing to gain by selling ads on this forum. The business is already extremely wealthy and successful. Any ad revenue from this would be absolutely peanuts. It just doesn't add up.
- jim_bailie 8y agoLook, it was an informative piece and I'm glad I read it. I even book-marked it for future reference. But let's be candid, It was a promotional piece as well and if I "owned" this forum I would certainly have something to gain by charging a modest fee for such placements. It may sound like I'm angry about this or have a negative feeling, but I don't; not at all.
- warent 8y agoThere's no argument whether you would have something to gain. We're talking about YC, not you.
- ringaroll 8y agoNice.
- raincom 8y agoThis makes sense only if your prospective clients want "lift and shift" into the cloud. But lots of people are using AWS for their services like S3, RDS, Cloud Front, Route 53, etc.
- ec109685 8y agoYou can still use s3, even if you aren’t totally in aws.
- sudhirj 8y agoNow if they'd only throw in S3 for pseudo-infinite data storage, reliable SQS for work queue management, 10/25/100 Gigabit networking between the instances, redundant power supplies and cooling and racks in carefully selected stable locations for free, I'd buy a dozen!
- KaiserPro 8y agoUsing both S3 and SQS outside of AWS is perfectly possible, and the Colo provides decent bandwidth, cooling and power, that's literally the point of them
- sudhirj 8y agoI never get the point of these things. This is like saying a Raspberry Pi hooked up to your home router is cheaper, why bother buying Gmail for your company. Or like saying that if you spend 20 hours every day in an Uber buying a car would be cheaper over 3 year term. Nobody is disputing that buying similarly specced hardware on the market is likely to have a cheaper TCO, the point of cloud providers is that also have a bunch of stuff attached to the servers, which turns to be pretty important.
- hughesjo 8y agoMost HN devs downvoting here never worked on real scalable systems to even understand what you are talking about, tbf
- dang 8y agoCould you please stop posting unsubstantive (and in this case, supercilious) comments to Hacker News? If you know more, it would be really good for you to share some of what you know, so that we can all learn something. If you don't want to take the time to do that, that's fine, but in that case please don't post. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- vidarh 8y agoBecause the vast majority of systems never need that scale. In other words: while it's an interesting niche situation, to most people how to handle small static workloads is much closer to what is actually relevant to them.
- Thorrez 8y agoSave $69k sounds a lot more significant than save 38%. If you're going to be using the server less than 73% of the time, AWS sounds better.
- deleted 8y ago[deleted]
- BGZq7 8y agoKeep in mind, the AWS cost is already for a reserved instance. On-demand would cost more.
- mbell 8y agoI'd be curious what the TCO is when factoring in storage. i.e. What is replacing S3 for data storage in the colo setup?
- deleted 8y ago[deleted]
- secabeen 8y agoThat system has slots for 16 2.5" drives in the back. I'd guess they can buy whatever commodity drive/SSD they want, and store the data there. Even throwing some cycles and memory at ZFS, the cost is small compared to the rest of the box. You'll need additional off-site backup, but that's starting to get out of the scope of the article.
- mbell 8y agoI don't think tossing 16 SSDs into a single enclosure is a fair comparison with S3. I also think that storage is absolutely in scope for an article like this.
- vidarh 8y agoTrue, for most workloads 16 SSDs in a single enclosure will be far faster and more efficient, but will require offsite backups. If you need redundant blob storage, pretty much every colo providers has solutions, and most of them are going to be cheaper than S3. Worst case you can use S3, and then need to factor in the bandwidth cost difference.
- secabeen 8y agoYep. Amazon even offers the Storage Gateway appliance to facilitate these workflows.
- ec109685 8y agoYou can still use s3, even if you aren’t in aws.
- mbesto 8y ago> Our TCO includes energy, hiring a part-time system administrator, and co-location costs. Which is a myopic view of TCO. This ignores so many things about purchasing on-prem hardware (for good or bad): Admin: - Cost associated with finding an admin who understands how this thing works Speed / convenience: - Try before you buy - Time it takes for the box to be built, shipped and sent to the data center - Time (and cost) it takes to install software, drivers, etc Maintenance / Capitalization / Finances: - 3-year? What is the useful life of this? When does it seemingly become obsolete? - AWS will continually upgrade their hardware and you keep paying the same - Hardware can be capitalized, which means you can push it to the balance sheet (for tax or valuation purposes) - Spending $90k instead of $184k in year 1 with the option to turn it off if you want (no longer need). This could be very valuable for a startup who wants elastic spending patterns. Hidden costs: - Returns, breakage, warranty in case of a hardware failure I understand why there is a market for this product, but it's not always an apples/orange comparison. Generally speaking, if you know what your workload is going to be, (I'd be hard pressed if a lot of orgs really know the answer to this) then on-prem hardware is not a terrible choice, but it has to be analyzed appropriately.
- microcolonel 8y ago> - Time (and cost) it takes to install software, drivers, etc I'd imagine this is part of having a system administrator.
- vidarh 8y ago> - Cost associated with finding an admin who understands how this thing works Managing AWS resources is not "free" either, or the market for people to handle devops work on AWS related to ongoing operations would be non-existent. My experience is when I've done consulting, clients on AWS used to end up paying more per instance than on-prem clients ended up paying per physical server. > - Time (and cost) it takes to install software, drivers, etc This is largely an initial cost of a day or two of setup when you start setting up an on-prem setup + for the first server of a totally new model. If you're buying one, then sure, you need to factor in a bit of time. If you're buying more, then if you use a competent admin the second one should be a matter of inserting boot media or (preferrably) configuring PXE booting, and picking an IP.
- canadev 8y agoInteresting note about Lambda Labs, all of the press links on https://lambdalabs.com/?ref=blog https://lambdalabs.com/?ref=blog are about a ~"privacy violating Google Glass app" that recognizes faces and geotags photos of them. I don't see why they choose to promote that link now.
- zten 8y agoWho's running model training 24/7 to justify reserving this instance or co-locating your own hardware? (Apologies in advance for not being very imaginative) Their ImageNet timing fits within the bounds of a Spot Duration workload, so in the most optimistic scenario, you can subtract 70% from the price - assuming spot availability for this instance type. (Of course, there are many more model training exercises that don't even remotely fit inside 6 hours.)
- bubblethink 8y agop3dn.24xlarge's pricing makes no sense at all. It feels like aws did it to pull off some PR/marketing stunt without any real users in mind. I've tried getting spot instances for it, but aws just errors out. So they don't even have enough of them to allow spot instances. And it's a gpu machine. So the usual arguments of scaling up on demand or adapting to load don't really apply. You either have this usecase or you don't. And if you do, just buy the hardware.
- hughesjo 8y agoYou are not the target audience since you don't have the use case nor the budget
- rb808 8y agoI bought a second hand Xeon E3-1246 v3 (8 VCPU), 16GB memory for $250 on ebay. That's less than it costs to rent an a1.xlarge for 6 months. Hardware is so cheap now, esp with SSDs and memory getting cheaper. Don't automatically rent!
- vidarh 8y agoOr if you're going to rent, consider dedicated hosting providers too. Providers like Hetzner often work out far cheaper (especially if you're doing anything requiring a lot of outbound bandwidth). AWS is great for convenience when you can afford it, but it is a really expensive solution, even when factoring in the extra things you have to deal with to rent, lease or buy dedicated servers.
- riku_iki 8y agoSadly no major dedicated provider offers machines with 8 GPUs..
- elchief 8y agoIn my mind, the reason to use something like AWS is to a) get your servers in minutes instead of weeks and b) easily right-size your service Once your service is somewhat stable in terms of size and you can afford longer lead times, then you should return to on-prem to save money
- joefourier 8y agoWhat's this false dichotomy between AWS and on-prem? Dedicated servers at Hetzner, OVH, Datapacket, etc. are much cheaper than AWS and can also be ready in minutes.
- elchief 8y agoI should have said on-prem or dedicated (or whatever the opposite of expensive AWS is)
- chx 8y agoThis false dichotomy between colo and AWS is just making me exhausted 'cos I have been repeating this for so long: just rent a dedicated server. There surely are some cases where colo is the best choice but as the years, now a decade pass since I have tried to spread this, it makes less and less sense every year -- and it never did much in the first place. Maybe if you have several racks worth of equipment? I am not familiar with that size.
- deleted 8y ago[deleted]
- mippie_moe 8y agoLambda Labs engineer here. Here's what we're trying to argue: Many benefits of running infrastructure in the cloud are lost for offline batch processing jobs. Training machine learning models doesn't require low response times, high availability, geographic proximity to clients, etc. Yet, with cloud, you're paying for all this extra infrastructure. The main benefits of cloud for tasks with machine learning type workloads are cost saving (if utilization is low) and no electrical set-up. On the other hand, cloud is extremely expensive for groups that require high base levels of GPU compute. The article is arguing that such groups can save a huge amount of money by moving infrastructure on-prem.
- chx 8y agoI feel I am talking to a brick wall. This often happens and this is what exhausts me. You are still arguing between cloud and colo / on-prem (even ignoring that on-prem and colo is very different but w/e) and what I am saying is that there is a third option neither the article nor your reply even acknowledges.
- icelancer 8y agoI have been hard-pressed to find Threadripper 2990wx dedicated servers with 1080ti/2080ti GPUs in them for rental, much less at a reasonable cost. Colo/on-prem was our only option for our mega transcoding server.
- deleted 8y ago[deleted]
- ti_ranger 8y ago1)How long does it take to get a Lambda Hyperplane operational, from the point I place an order? A p3dn.24xlarge is a few minutes. My experience in deploying new hardware-based solutions is that it typically takes between 6 and 12 weeks. 2)How does the TCO compare when I only need to train for 2 hours a day? 3)If I was previously training on AWS p2.16xl, I could upgrade to p3.16xl with basically 0 incremental cost (for the same workload). Does LambdaLabs offer free (0 capex) upgrades? If so, how long would it take to upgrade?
- ec109685 8y ago1) 5 days (according to the author) 2) Obviusly the cloud is better in that case 3) You can’t always upgrade on aws eigher if you have paid for a reserved instance: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ri-modifying.html https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ri-modif...
- mschuster91 8y ago> My experience in deploying new hardware-based solutions is that it typically takes between 6 and 12 weeks. Most of this is due to arcane ordering processes both on client and supplier side.
- deepnotderp 8y agoIt may be useful to note that most deep learning workloads for training are pretty latency insensitive and are pretty flat throughout the day.
- Scaevolus 8y agoIf you're not deploying in a datacenter, you can save even more money by building a workstation with a few 2080 Ti cards, which cost $1200 and give 90% of the speed of the $3000 Titan V: https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v100-vs-titan-v-vs-1080-ti-benchmark/ https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v...
- altmind 8y agoExactly my thoughts - instead of one datacenter class gpu, you can use multuiple desktop components for better performance and more performance per buck. I've seen some sweet supermicro 2U chassis that can fit and run cool four 2080Ti. I remember, though, that NVIDIA disallows in EULA the use of nvidia driver blob in consumer gpu in datacenter environment. Time will show how legally enforcible is this restriction.
- freeone3000 8y agoWe have one of these. Aside from the mechanical differences (it's actually a 2.5U server due to how the power cables connect to consumer cards), you're losing out on NVLink. You're going to be restricted to single-card training, which means either batching or other restrictions to ensure it fits in the memory of a single card. This setup is able to train four one-card models in parallel; it is not able to train a four-card model.
- riku_iki 8y agoNVLink supports 2080 Ti
- freeone3000 8y agoThere's a physical connector also called NVLink, yes, but the cards don't present as unified at the driver level; nvidia-smi shows "link: off" even using the bridge. It's effectively SLI, which can reduce memory bandwidth load and cross-card transfer costs, but doesn't unify the cards the way Volta (actual) NVLink does.
- m0zg 8y agoIf it's really on-prem (i.e. not in the "datacenter" as per NVIDIA EULA), you could spend a lot less than $100K+ for a lot more throughput by purchasing consumer-grade cards and HEDT gaming hardware. Sure you'll have 4 GPUs per box and not 8, and sure, each GPU will have 11GB and not 32, but the whole machine (_with_ the GPUs) will cost just a tad more than a single V100. So if you don't really need 32GB of VRAM per GPU (and you most likely don't), it'd be insane to pay literally 5x as much as you have to.
- manigandham 8y agoCloud computing was never about price, it was about the ability to provision and operate infrastructure instantly through an API. If you can take advantage of that flexibility to build reactive capacity then you can save money, but that wasn't the initial driving point.
- latch 8y agoyou MIGHT save money. The price difference can be substantial enough that you might simply be better off having it "over provisioned" off hours. Also, not having "reactive capacity" might simply mean that your 95th percentile goes from 100ms response time to 300ms. Which again, might be a more cost effective approach. I think the initial development and ongoing cost of maintaining that reactive capacity is also substantial enough to be considered. Most of the apps that truly need the elasticity should probably going hybrid anyways. Baseline on dedicated hardware, spikes on spot instances.
- icelancer 8y ago> If you can take advantage of that flexibility to build reactive capacity then you can save money Almost everyone thinks they can; very few small businesses have the sysadmin/devops to pull it off.
- boulos 8y agoDisclosure: I work on Google Cloud. First, thanks for writing this up. Too many people just take a “buy the box, divide by number of hours in 3 years approach”. Your comparison to a 3-year RI at AWS versus the hardware is thus more fair than most. You’re still missing a lot of the opportunity cost (both capital and human), scaling (each of these is probably 3 kW, and most electrical systems couldn’t handle say 20 of those), and so on. That said, I don’t agree that 3 years is a reasonable depreciation period for GPUs for deep learning (the focus of this analysis). If you had purchased a box full of P100s before the V100 came out, you’d have regretted it. Not just in straight price/performance, but also operator time: a 2x speedup on training also yields faster time-to-market and/or more productive deep learning engineers (expensive!). People still use K80s and P100s for their relative price/performance on FP64 and FP32 generic math (V100s come at a high premium for ML and NVIDIA knows it), but for most deep learning you’d be making a big mistake. Even FP32 things with either more memory per part or higher memory bandwidth mean that you’d rather not have a 36-month replacement plan. If you really do want to do that, I’d recommend you buy them the day they come out (AWS launched V100s in October 2017, so we’re already 16 months in) to minimize the refresh regret. tl;dr: unless the V100 is the perfect sweet spot in ML land for the next three years or so, a 3-year RI or a physical box will decline in utility.
- deleted 8y ago[deleted]
- freediver 8y agoThe actual cost of running a cloud instance is inflated. The cheapest way to run them is using spot/interruptible instances which for most deep learning jobs will suffice. If anything there will be some upfront cost to set it up in a way that it automatically manages interruptions, storage etc. Also by not limiting yourself to AWS you can have many other options. With this setup you can get 2x4x V100 on Azure for a total of $42k/year (assuming running 24/7). Even if one spent $40k to write code for spot instance management this is by far the cheapest solution for GPU compute both short term and long term. source for calculation: https://cloudoptimizer.io https://cloudoptimizer.io