16 ms·
It's wild to me that sites managing storage clusters at petabyte scale are doing it on AWS. I would think that by then you could save millions more by migratin
by deeblering4 5y ago
It's wild to me that sites managing storage clusters at petabyte scale are doing it on AWS. I would think that by then you could save millions more by migrating to your own colocated hardware.
- mwcampbell 5y agoI would hope that for a high-throughput DB cluster like this, they're using instance-local storage rather than EBS. If that's the case, then they're probably already taking advantage of EC2 reserved instances to save a lot compared to the on-demand prices that we usually see.
- jhk727 5y agoWe are, though we started out using EBS. As you mentioned, NVMe instance storage performs much better for our workload. We work around the lack of durability through strong automation of point in time restore/swapping in of new nodes in case of hardware failures. And yes, reservations make a massive difference economically.
- Stevvo 5y agoFar lower risk and capital investment than collocation. I've never had to store petabytes of data, but I would imagine the considerations are not too different to smaller scales.
- aaronblohowiak 5y agoThe amount of staff you have to have on-hand and amount of pre-planning (and up-front capital commitment) can all make that very unattractive long after the basic per-GB price would seem to make it attractive.
- b112 5y agoNo. You already have 24x7 staff at this scale. Hardware requires thought and skill, but then so does software. It isn't voodoo.
- fwip 5y agoNot always true. Some data is intrinsically bigger than others. If you have a petabyte of chatlogs, sure, you have 24x7 obligations to millions of people. If you have a petabyte of astronomy data, you have like 3 research scientists using it.
- koolba 5y agoIt is a different skill though. Going from zero to one for physical infrastructure is a significant leap in both cost and operational process. You need to manage inventory, provide 24/7 physical access, and set up supply chains to ensure you have ongoing availability.
- Spivak 5y agoY'all are making "buy an asset tag printer", "have a rep from Dell/HP" and "use the data center's remote hands if you need it" sound crazy complicated.
- z3t4 5y agoYou don't have to do everything in house, you can for example buy servers with a on site support agreement. Then you just have to buy new servers at regular intervals, you don't need to have a guy that can fix a server with a soldering pen. Same for internet connection, you can buy transit, no need to become your own ISP. So you don't need people who deal with peering agreements, etc. For electricity you can make a support deal with a local electrician company. You don't need a guy who can build and maintain a custom power supply unit. It does help however to have someone with basic sysadmin and network skills. But if you don't have that, you will sooner or later screw up your AWS infrastructure too.
- teknopurge 5y agoOr save even more using decentralized options like Storj.io or Filecoin. IMO, market pendulum is swinging(in-motion) right now: cloud ---(cost drivers)---> dedicated HW/colo with hybrid or custom cloud---(operational drivers)---> web3 decentralization [^--- we are here ---^]
- tmikaeld 5y agoYou must be kidding? Those solutions are not even close to running any database of this scale, did you even read the article?
- teknopurge 5y agoyep. Those solutions aren't there yet, and arguably too early, but the market direction is clear. In five years I would not be surprised to see some form of transactional data store(high TPS, blob store) on a decentralized layer. Crazy to talk about different FS compression schemes when trying to optimize business/app logic higher up the stack. should be abstracted away by now. (yea, I know it's not, but should be)
- _nickwhite 5y agoAWS is more often than not a better solution than colo when factoring in the on-site engineers, techs, and operational complexity costs a company will pay to monitor and respond to hardware-related events. One could build out a datacenter management team with on-call engineers, or, they could pay AWS to handle all that, and focus on innovation and products that make their company unique and (hopefully) profitable. AWS makes a lot of sense for companies that wish to inoculate themselves from the hardware layer, and it would probably take a company many magnitudes larger than Heap to realize any real benefits from self-hosting at a colo. This isn't even considering the fact that uptime matters, and you'll need more than 1 colo to really do it right. I say this as someone who built, manages and operates datacenters and colo spaces.
- runlevel1 5y agoI'm not saying it never happens, but I've never seen moving to AWS (or Azure, GCP, etc.) save in people costs at any tech company with a large resource footprint. It just shifted where the time is spent and who had to spend it. The public cloud and managed services work great for the most common use cases, but go outside those and you start having to engineer around limitations. If you have a sizable footprint in any given dimension you're trading one complexity for another.
- bombcar 5y agoThe "who had to spend it" is huge - companies love paying providers and hate paying people/depreciating costs.
- markus_zhang 5y agoCuriously our company moved from AWS to on-premise a couple of years ago. Something about CAPEX -> OPEX was mentioned back then.
- bombcar 5y agoThat's the second part of the loop, when the cost of AWS is high enough that you can show immediate dollar savings by bringing it in-house.
- sorenjan 5y ago"Nobody Ever Got Fired for Buying IBM" Maybe it's worth a couple of million to not have to deal with the risk, and just keeping status quo.
- dilyevsky 5y agoExactly, most management would prefer to just set investors money on fire and keep risk profile low if their business model can support it
- jasode 5y ago>petabyte scale are doing it on AWS. I would think that by then you could save millions more by migrating to your own colocated hardware. Usually, those types of judgements are based on thinking of AWS as a "dumb datacenter" such as a bunch of harddrives or just bare cpu. AWS is more cost-effective if you use high-level AWS services instead of just storing files in the cloud. In this case, it looks like Heap is also using AWS Redshift and probably a bunch of other services in the AWS portfolio. A similar comment I made previously: https://news.ycombinator.com/item?id=28288352 https://news.ycombinator.com/item?id=28288352 So for self-hosting hardware, Heap would not only build up the petabytes of diskspace, they also have to replicate Redshift functionality and the entire AWS services portfolio they're using. If you use enough AWS _services_, it becomes cheaper than self-hosting because you don't have to reinvent the wheel.
- deeblering4 5y agoMy main take aways from this are that cloud vendor lock-ins are real, and they can be hard to break free from. Perhaps that's more of a cautionary tale for new projects than a justification for the expense though.
- r3trohack3r 5y agoWhat you're calling "vendor lock-ins" I'm calling "providing sufficient value to justify cost." It's not that migrating out isn't possible, it's that Amazon is providing "Engineering/SiteOps Departments as a Service" at a price that's hard to compete with in house.
- hackerfromthefu 5y agoWhat's the newspeak for the high egress fees?
- r3trohack3r 5y agoI'm not sure that's what I was arguing against. From the GP's post - I was trying to say I read this as "the collection of services AWS provides would require several in-house engineering teams to compete with" not "vendor lock-in". A single service is relatively easy to replicate if it is core to your business, but an entire on-demand datacenter w/ abstractions like Time Series databases, pub/sub services, etc. isn't as trivial to do yourself. Many of these services require teams of engineers to manage at scale, and engineers that _understand_ them well. The on-demand service catalog of the cloud provides significant value. It's more than "pub/sub" or "blob storage" as a service. It's an entire engineering organization and data center as a service w/ pre-built architectures for you to start using today.
- WJW 5y agoThe discounts you can get from doing anything at big enough scale will push your costs back to colocation prices. Don't assume that anyone with a cloud bill over 200k is playing anywhere near the price you read on the pricing page.
- Robotbeat 5y ago“Will”? I doubt it. Definitely not with that level of certainty.
- tomnipotent 5y ago> certainty Considering the number of people here commenting about costs but have never managed a P&L, I don't think certainty is high on the list.
- FpUser 5y agoNope. I had chance to compare what one org had for 600k of real money after all the discounts. Not even remotely close to what one can get for rented dedicated servers.
- thehappypm 5y agoIt’s not always that dollars you spend are evil. Sometimes the more expensive or slower solution is better because it means you can have fewer people working on it, or have lower-skilled people managing it. And new people are easier to hire and make productive because it’s a shared skill. And you get updates for free. And outages are likely to be shorter. And.. the list goes on.
- jhk727 5y agoAuthor here - as others have noted, there's a lot of benefits to a company of our size operating infrastructure on AWS vs. managing physical hardware. A couple of the highlights for managing our primary database cluster include: - Automation - this was noted by another commenter, but with AWS we can fully automate the instance replacement procedure using autoscaling groups. On hardware failure, the relevant database is removed from its autoscaling group and we automatically start restoring a fresh instance from latest backup. This would be much more difficult if we were to manage our own hardware. - Flexibility - we have the ability to easily change instance classes via selling/buying reservations. Some of our biggest wins historically have come from AWS releasing new instance families - we've been able to swap out the hardware for our entire cluster over a week or so, for negligible cost (often saving money in the process due to the cost per unit of hardware decreasing on new instance classes). While we could leverage the same developments in a self-managed environment, it would be more difficult, and likely more expensive due to how capital-intensive self-hosted is. Additionally - there's a ton of value in the integration of the AWS ecosystem. We use many AWS managed services, including heavy use of RDS, Kinesis, S3, and others. For a company with a relatively small engineering team managing a large infrastructure footprint, it hasn't made financial sense yet to invest in moving to self-hosted infrastructure.
- shrubble 5y agoIt seems shocking to me, that you haven't yet migrated, however you know your costs/benefit ratio more than I do. Have you ever examined a split model, where some parts of the load are run on your own or rented dedicated servers, and some runs on AWS? Separately to your comment about LVM... the LVM snapshot requires that a separate part of the volumes be set aside to hold the snapshot data. If the snapshot volume fills up with changes being made to volume that holds your data before the snapshot completes, then the snapshot will fail. This does not occur with ZFS as you have noticed.
- nine_k 5y agoWith the price of the egress being what AWS charges, such a split may still not make economical sense if it crossed a data-intensive boundary. Also, with many customers also using AWS, much of the traffic may not even leave a datacenter, improving speed, reliability, and maybe even cost.
- dayjah 5y agoI feel it's fair that they're on AWS right now. Generally the arc of MVP->IPO involves using the cloud to find product market fit, and as that fit improves your revenues should also. Moving from the cloud to a colo would then be driven by capital investment to bring down COGs; to either improve PPS or get to cash-flow positive. Heap using AWS just means they've not yet reached a point on that trajectory where the capital investment moves the needle enough to warrant it. That could be for any number of reasons.
- namdnay 5y agoKeep in mind that nobody at large scale is paying the sticker price for AWS (or Google or Azure)
- ksec 5y ago>storage clusters at petabyte scale are doing it on AWS. I had to double check just in case, but petabyte is only 1000 terabyte. It may be big in terms of database, but rather small in absolute terms. You could fit a single Petabyte in a 1U server. I doubt they pay listed price. And AWS is now mostly a Enterprise and Sales game. So once you ran other cost involved in managing, I would think you need to be multiple Rack scale before the cost break down better for your own hardware. And that is excluding other benefits of sitting inside AWS ecosystem. The only thing I think AWS isn't so good at is the low cost, sub $1000 per month spending scenario. Where you are paying a lot more just for staying inside the ecosystem for things you may not be using. Those tends to flavour Linode or DO.
- toast0 5y ago> You could fit a single Petabyte in a 1U server. That seems a bit over the top. I see 18 TB drives available, but let's posit 20 TB drives, so you need 50 of them. I don't think you can fit 50 3.5" drives in a 1U space, even if there's no motherboard or power supply. 50+ drive storage chassis are generally 4U. I did see some 16 drive 1U servers though, so I'm pretty sure you could fit that much storage into 3U even though I also didn't see any 3U storage chassis.
- namibj 5y ago2.5" SSDs are AFAIK the medium for this.
- ksec 5y agoThis has been a thing since 2017 [1] using EDSFF. It is technically possible for 2 PB per 1U, but price and market demand doesn't seems to value these kind of density. So they are mostly used as half petabyte for now. [1]https://www.supermicro.com/newsroom/pressreleases/2017/press170914_32_JBOF.cfm https://www.supermicro.com/newsroom/pressreleases/2017/press...
- deleted 5y ago[deleted]
- enginaar 5y agoBank of America saves $2 billion per year https://www.google.com/search?client=safari&rls=en&q=bank+of+america+saves+2+billion+on+cloud&ie=UTF-8&oe=UTF-8 https://www.google.com/search?client=safari&rls=en&q=bank+of...
- tgsovlerkhgsel 5y agoI'm not surprised. Scale makes IT cheaper, and AWS has scale. That means the actual total cost (not what is paid to Amazon, but what it actually costs to run) for something running on AWS will almost always be lower than a custom data center, due to the one-off reinvent-the-wheel work you'd have to do to run your own. The only remaining question is, who keeps the savings. AWS can make a nice profit by selling their services at a higher price than it costs them to provide their part, but still cheaper than running your own data center. That means AWS will always be able to provide you the service cheaper than if you build your own. Whether they're also willing to do that is another question, but it seems logical. At list prices, it's probably not worth running in AWS, but I highly doubt someone doing petabyte scale is paying list prices. Amazon has every motivation to provide a hefty discount to make the "own datacenter" approach unattractive.