6 ms·
I just spent $50K on coloc hardware. I'm taking a $10K/mo Azure spend down to a $1K/mo hosting cost. But the real kicker is that I get x5 the cores, x20 RAM,
by joshuaellinger 7y ago
I just spent $50K on coloc hardware. I'm taking a $10K/mo Azure spend down to a $1K/mo hosting cost.
But the real kicker is that I get x5 the cores, x20 RAM, x10 storage, and a couple of GPUs. I'm running last-generation Infiniband (56gb/sec) and modern U.2 SSDs (say 500MB/sec per device).
I figure it is going to take me about $10K in labor to move and then $1K/mo to maintain and pay for services that are bundled in the cloud. And because I have all this dedicated hardware, I don't have to mess around with docker/k8s/etc.
It's not really a big data problem but it shows the ROI on owning your own hardware. If you need 100 servers for one day per month, the cloud is amazing. But I do a bunch of resampling, simple models, and interactive BI type stuff, so co-loc wins easily.
- wpietri 7y agoI'm sure your right for your case. But I'd add one caveat for those less experienced: if you own the hardware, you need to be prepared to go to the colo when something breaks. The various clouds are a much nicer experience when hardware fails. At the very least people should have enough spare capacity that a hardware failure means going sometime in the next couple of weeks, rather than getting up at 3 am and fixing things under pressure.
- foobiekr 7y agoOperations teams deal with both. You design your system with enough spare capacity that you can live somewhat degraded for a time - you must if only due to the lead time. Software failures are far far far more common than hardware failures so once you combine these, the occasional midnight trip to the colo is both rare and oddly satisfying for hero types.
- latch 7y agoOr take the middle road and just get rent the hardware (aka, dedicated hosting). You pay more than colo but still way less than cloud, get the same level of hardware support as a cloud provider but the same performance as colo.
- kavalg 7y agoYep, for example Hetzner offers bare metal servers as well as cloud instances at a very reasonable price. (Not affiliated in any way. Just a happy customer.)
- TheSpiceIsLife 7y agoI would have assumed the colo provider would offer Remote Hands, so you’d only need to send replacement hardware. That’s how the DC I used to work in operated.
- wpietri 7y agoIf you have enough spare capacity and the problem is pretty mundane, sure, that can work. But if not, then it's off to the colo while the rest of the company freaks out.
- joshuaellinger 7y agoPrior to the cloud, I ran at a coloc facility for 15 years. I break stuff much more often than having it actually fail. So... make yourself robust against human error first and you'll probably cover the hardware side as a side-effect. I am more likely to hose a machine during an OS upgrade and not have time to recover than I am to have an SSD fail. But spare capacity is a good idea, especially if you have real-time traffic.
- dmak 7y agoHow did you estimate your hardware needs?
- adtac 7y agoI plan to do this in the near future once my GCP credits are used up (18 months of credits left). My plan is to temporarily shift to dedicated hardware through a service like Hetzner to evaluate what kind of hardware I need. I can simply redirect a fraction of the traffic and extrapolate. Since this is elastic there will be no upfront costs, but I can play around with different sizes. Once I'm happy with my estimate, buy real hardware and move the rest over. At least that's the plan. I don't think you can do much more than an educated guess and I think this will be as close as I can get. Not AI related btw.
- joshuaellinger 7y agoGee... if only there was a service where you could spin up machines on demand. (joke) I kinda worked backwards from the cost. I ran the business for a year on Azure but each 'sample' of the resample took about 2 mins so it precluded any near real-time analysis. I ported the kernel to a GPU locally using python/numba and it ran in about 10 seconds and that was enough to seal-the-deal. From there, I spec-ed out a GPU server and then machines that matched each role in my environment. I decided I was willing to spend $50K and just started loading up the machines.
- burnte 7y agoI did similar at my current and last job. Rather than spend $24k/month, I spent $50k, bought a shitton of hardware, built a virtualization cluster at Corp, and upgraded our connections. Accounting thought i was a wizard.
- walshemj 7y agoEspecially as they can amortize that cost in the annual accounts - there might even be RnD tax credits they can use
- dboreham 7y agoWe never went cloud, except for ancillary things like build machines, nagios etc that run on tiny VMs. Whenever I looked at the economics I could buy a server of the class we needed for roughly 2x the monthly rent for the equivalent from Amazon.
- eyegor 7y agoYes, it's quite obvious when you actually have compute needs. At my current employer, we spent about 100k to build a small single purpose hpc. One year later, I calculated the azure costs (help bargain for more servers) would have been around 1.5m. This is almost 24/7 use though, and add another ~150k in electricity.
- foobiekr 7y agoFor my own company we built out at two regionally distinct colo facilities. That worked really well and operations was efficient and costs were moderate, clearly tied to CAPEX increments which were predictable. Recent projects have been on AWS. For a project that is roughly on the scale of our colo in terms of instances, though with aggregate lower performance, we are buying one of our colos every year. It’s insane. Network costs are particularly egregious in AWS. But there is absolutely no way we’d be permitted to build colo facilities for many reasons and there are many reasons why even if we could get permission to do so we would choose not to due the resulting death by a thousand cuts orchestrated by the team who happens to have inserted themselves as the owner for DC/colo like things.
- angry_octet 7y agoYes, cloud costs are the cost of having poor internal management, such that inefficiency and incompetence reigns unchallenged. The enormous cost differential is unfortunately borne by the unit doing the work, rather than the one preventing it from being done efficiently.
- foobiekr 7y agoThat is a very accurate summary.
- bsenftner 7y agoI ran a ML based 3D reconstruction service for 7 years - given face photos of a person, reconstruct a realistic 3D likeness. I licensed a finished 3D reconstruction algorithm, purchased $50K worth of servers plus a federal reserve bank quality hardware firewall, and put it all in a Los Angeles downtown co-lo (the former Enron data center, actually.) I paid $600 a month to run that, as opposed to the equal compute capability being $96K per month if run at Amazon. It kills me to see people being raped by the cloud, but everyone just lines up like good little boys...
- angry_octet 7y ago(I don't think people like the gratuitous imaging generated by that last sentance. Much too real.)
- mrosett 7y agoAgreed. Totally unnecessary
- andrew311 7y agoWhat colo company did you use?
- joshuaellinger 7y agoIn Austin, DataFoundry is by far the best. It was overkill for me and went with something off the beaten path but they have an amazing facility. I wound up at a facility run by a fiber vendor because they'd sell me a fixed 250mbps pipe for the same price that a data center would sell me 20mbps pipe that bursts to 1gbps. It only works for me because of the nature of my business -- most people would be better off somewhere else. Choosing a co-loc facility is complicated. My recommendation is to tour and get quotes from 3-5 vendors in your area before choosing anyone. Ideally, take someone who has done it before.
- sabalaba 7y agoYea we’re seeing this all over the place at Lambda (https://lambdalabs.com https://lambdalabs.com). Most people running consistent GPU training or inference jobs are building on-prem clusters or even groups of workstations. It just doesn’t make financial sense to use the big the cloud service providers for those with consistent workloads. I always hear stories where folks have saved hundreds of thousands in infrastructure costs with owning + co-lo.
- chintler 7y agoI agree to this, and I think lambdalabs is quite precisely positioned for on-prem training. As an aside, thank you for your one-line installer script for tf/keras. Earlier, my team used to spend days figuring out the CUDA/tf/keras/CUDNN etc dependency charts, and you've brought that down to ~0.
- htrp 7y ago+1 for the one line install
- marcus_holmes 7y agothe point of Cloud is that it solves the problem of variable demand. I used to run on-prem back in the 2000's, and we were constantly dealing with demand fluctuation crises. Spinning up new physical servers to deal with new demand, or being massively over-specced when demand dropped, was a real pain. I'm starting a new thing this week, and using the Cloud for it because I have no idea what our demand will be. I can start small, scale up with our customer growth, and never have to worry about ordering new servers a month in advance so I have enough capacity when (or if) I need it. At some point in the future, when our needs are clear and relatively stable, it might make sense to migrate to on-prem and save those costs.
- deleted 7y ago[deleted]
- throwaway9d0291 7y agoI half-agree. The Cloud specifically solves the problem of _highly_ variable demand. If your peak demand is 100x your baseline and only happens for ~1h each day, cloud is almost certainly a good choice. If it happens for ~12h a day or it's only 5x your baseline, the cost of the cloud is such that you're likely to save with dedicated hardware, even though much of your hardware sits around doing nothing part of the time. > never have to worry about ordering new servers a month in advance so I have enough capacity when (or if) I need it. There is a middle-ground that's very much worth considering: renting dedicated servers. It's not quite as cost-effective as colocation and owning your hardware when you have at least a cabinet worth of stuff but it does offload the management of the hardware and provisioning to somebody else. They can also usually be provisioned in a matter of minutes. In some cases (e.g. Packet.net) these machines can even be treated essentially like cloud instances, with hourly pricing. There's also yet another middle ground: using dedicated to handle the known and predictable baseline traffic and using the cloud to handle the unexpected bursts.
- supermatt 7y agoThanks for the pointer to packet.net. This looks like it will scratch a current itch.
- Merrill 7y agoThis whole topic recapitulates all the arguments for business units acquiring and operating their own servers versus continuing to suffer the internal bill-backs from the corporate data center. Some of the same caveats apply with respect to software updates, configuration control, security, availability, business continuity, disaster recovery, and what happens if the local admin is hit by a bus.
- StreamBright 7y agoExactly. These examples are mostly apples to oranges comparisons. I have worked over 20 years in OPS and it is really hard to do cheaper than AWS ____in the long run___. If you are unlucky and bought a batch of SSDs that are faulty exactly 1 month after warranty expires or you have downtime because of other low-level reasons that AWS shields you from, your co-loc cost can quickly go up. I don't even want to go into networking hoops, that is a whole different problem to deal with global network vendors. If you can be sure you never run into these, or your business is resiliant to these sort of problems, or you have a dedicated highly skilled team (like dropbox) than co-loc might be a good idea. Otherwise it is pretty damn hard.
- StreamBright 7y agoNetwork redundancy, electricity redundancy, bandwidth included? Otherwise, it is a bit of apples to oranges. What about firewalls? I mean you could ignore all that and say you only need raw computing power. On the k8s note, nobody is forcing you to use k8s on the top of Azure. Now do the calculation for ongoing operations for 5 years, taking into consideration normal hardware failure and maintenance cost. You need to swap out old hardware to get a new CPU, etc. I have tried to use co-loc vs cloud for ~100 nodes and cloud won, by 30%.