15 ms·
Author here - as others have noted, there's a lot of benefits to a company of our size operating infrastructure on AWS vs. managing physical hardware. A couple
by jhk727 5y ago
Author here - as others have noted, there's a lot of benefits to a company of our size operating infrastructure on AWS vs. managing physical hardware. A couple of the highlights for managing our primary database cluster include:
- Automation - this was noted by another commenter, but with AWS we can fully automate the instance replacement procedure using autoscaling groups. On hardware failure, the relevant database is removed from its autoscaling group and we automatically start restoring a fresh instance from latest backup. This would be much more difficult if we were to manage our own hardware.
- Flexibility - we have the ability to easily change instance classes via selling/buying reservations. Some of our biggest wins historically have come from AWS releasing new instance families - we've been able to swap out the hardware for our entire cluster over a week or so, for negligible cost (often saving money in the process due to the cost per unit of hardware decreasing on new instance classes). While we could leverage the same developments in a self-managed environment, it would be more difficult, and likely more expensive due to how capital-intensive self-hosted is.
Additionally - there's a ton of value in the integration of the AWS ecosystem. We use many AWS managed services, including heavy use of RDS, Kinesis, S3, and others. For a company with a relatively small engineering team managing a large infrastructure footprint, it hasn't made financial sense yet to invest in moving to self-hosted infrastructure.
- shrubble 5y agoIt seems shocking to me, that you haven't yet migrated, however you know your costs/benefit ratio more than I do. Have you ever examined a split model, where some parts of the load are run on your own or rented dedicated servers, and some runs on AWS? Separately to your comment about LVM... the LVM snapshot requires that a separate part of the volumes be set aside to hold the snapshot data. If the snapshot volume fills up with changes being made to volume that holds your data before the snapshot completes, then the snapshot will fail. This does not occur with ZFS as you have noticed.
- nine_k 5y agoWith the price of the egress being what AWS charges, such a split may still not make economical sense if it crossed a data-intensive boundary. Also, with many customers also using AWS, much of the traffic may not even leave a datacenter, improving speed, reliability, and maybe even cost.
- mhio 5y ago> This does not occur with ZFS as you have noticed I'm not sure exactly what "This" refers to? Just wanted to note that a ZFS snapshot can be destroyed when the parent pool runs out of space too, but you don't need to allocate a volume.
- shrubble 5y agoFor LVM you usually have to pre-allocate the space. So perhaps you think that you will need 8GB to hold the changes during the snapshot operation, and it works great for 6 months until more data is added and 1 customer does a lot of small updates during their maintenance window, which overlaps with your backup schedule... and the snapshot operation fails. No data is lost in this case, but the back up doesn't finish.
- newsclues 5y agoAre you in the gap oxide computer is trying to fill?
- deleted 5y ago[deleted]
- jes 5y agoThank you for writing this article. I’m curious: The engineers that brought these significant cost savings to your company, did they receive a share of the money saved?
- AstroDogCatcher 5y agoThank you - haven't laughed like that in a while.
- Aachen 5y agoI'm not sure how appropriate it is to take a serious comment where the author has a genuinely unpopular opinion and say you laughed really hard at it.
- lmm 5y agoIt's clearly not a serious comment, and I say this as someone who thinks there should be more workers' cooperatives in the tech industry.
- jes 5y agoI was and still am entirely serious. I don’t know you and doubt that you and I have ever said three words to each other. What’s your basis for your claim to clearly know my intent? The article author works for a company that has 54 open positions. Thousands of capable developers will or have already read his article. Why can he not talk about how great it is to be a developer at his company? Seems like a missed opportunity to me. Also, if you are an engineer that can save your company millions of dollars in real life, that’s worth something, isn’t it? You won’t get paid for it unless you have the courage to ask hard questions. Your experience will obviously be different than mine. But please don’t pretend that you know me, or what I’m thinking when I offer a comment. It would be much better to show up with curiosity rather than condescension. I wish you well.
- Aachen 5y ago
- tinco 5y agoThe things you mention are ostensibly true, yet still they don't make sense to me. It might make sense when you're a startup that's growing, but when your SSD costs are so large you can save millions on them, then the numbers just don't add up. In my experience, doing things in the cloud is about as expensive per 12-18 months as buying the hardware up front is. That's super interesting for a fast growing startup that could go bust any minute and wants to spend every second of their time on growing, expanding and marketing. But when you're spending so much on AWS you can save millions just by reducing filesystem overhead by 20%, it should have stopped making sense a while ago. $2 million should get you a team of 10 sysadmins and devops engineers. Sure automation would be more difficult, but you'd have the manpower to achieve it. Isn't that what running a business is about? Flexibility, when you're growing quickly it's nice that you can provision new hardware instantly, but AWS is so expensive you could continuously over provision your hardware by 50% and still always be ahead of the AWS price curve. And as I said, you could fully swap out your hardware every 18 months and be at the same price basically. You could even hire a merchant to offload your old hardware and recuperate 50% of those costs. And I'm not saying to throw AWS overboard altogether, that you have your core business outside of AWS's datacenter doesn't preclude you from buying into RDS, Kinesis, S3. Is AWS just cutting you more financial slack than we're getting as a tiny company? Or am I underestimating the costs of getting that sysadmin team on board?
- zerd 5y agoThere's also opportunity cost. What if you instead of hiring a team of 10 people (probably need more, managers, distributed geographically etc), you hire more developers, marketing, sales and increase revenue. So instead of saving $2M, you make $3M more revenue. Making up numbers of course, but companies can estimate this. If that's the case it makes sense to keep paying AWS, until the ratio changes, and then they can cut costs.
- gwright 5y ago> In my experience, doing things in the cloud is about as expensive per 12-18 months as buying the hardware up front is. It seems like you are glossing over the other costs: staff to implement and manage, development time and maintenance for automation to re-implement everything that AWS includes, data center costs (not clear if you were thinking of hardware ownership only or data-center also). I'm not saying you didn't think about those things just saying that they can't be ignored in these types of comparisons.
- qatanah 5y agocurious to know what instance type are you using for AWS? Are they i3 instances?
- throwdbaaway 5y agoOne thing to consider is that when running on-prem, you will be able to procure disks that advertise 512B sector size, and actually performing well when doing so. With that, you can use 8KB record size on ZFS, and still obtain the same 5.5x compress ratio, assuming you are currently using 4KB sector size with 64KB record size on ZFS. With 8KB record size, there would be no read/write amplification caused by postgres. More importantly, ZFS fragmentation would be much less of an issue. Of course, disks procurement can be a nightmare process, which is why AWS is printing money.