14 ms·
Pinterest Cut Costs from $54 to $20 Per Hour With Automatic Shutdowns
- dclusin 14y ago$54 per hour for what? Is this simply the cost of running the instances? It seems to me like pinterest is throughput bound and the vast majority of their costs would be accrued via transit and storage charges.
- aioprisan 14y agoI agree with your questions. I'd be interested what percent that signifies versus the total per hour bill for all services, not just CPU time, but bandwidth as well
- jackowayed 14y agoFor now they really only saved $300k/yr, which is less than the cost of two engineers. Though I guess cutting a nontrivial expense that is likely to grow by nearly a factor of three is pretty great and probably worth investing in. That said, this adds complexity to their systems with the only benefit being cost savings. Given that we can assume that no code is perfect, it's likely that at some point the auto-downscaling will cause an outage or period of slow responses, which could easily lead to lost usage and trust that costs them as much as they're saving on ops.
- pdog 14y agoIs cost savings really the only benefit? It seems to me that this is a better engineering outcome.
- jspthrowaway2 14y agoThat depends on for who. For Amazon, perhaps.
- achompas 14y agoBetter for Pinterest too. It means they need to develop services that can be launched and terminated quickly, and that their services are much more resilient during times of high load.
- sliverstorm 14y agoIMO designing systems that can be powered on and off regularly (and quickly) is a good thing in itself; it encourages "proper" setup. I've found servers that have been on for months or years tend to need manual intervention after a reboot. When the machine could reboot every day, you can't have that. In other words, it's just good design.
- calpaterson 14y agoNitpicking: but it's not good "in itself", it just has other benefits.
- jacques_chester 14y ago> IMO designing systems that can be powered on and off regularly (and quickly) is a good thing in itself Despite the fact that, in theory, a mainframe should never go down, most dinosaur pens will power cycle them regularly just to see what happens if you come up from a cold start. A mate of mine worked in a dino pen where they did this on Saturday evenings. He told amusing stories.
- jedberg 14y ago> That said, this adds complexity to their systems with the only benefit being cost savings. That's not the only benefit. They also get much better reliability. The systems scale down with load, but they also scale up. As their load increases their system scales up along with it, giving greater reliability during increased load. Since they have to architect for that, it makes scaling up in general easier.
- azakai 14y ago> For now they really only saved $300k/yr, which is less than the cost of two engineers. It took them 2 weeks to implement. That means it's saved a lot more than it took to implement (2 weeks times a couple of engineers) already. To use your metric, if they saved close to $300K/year, that means they could afford to add 2 engineers to their staff, which is very significant.
- jbigelow76 14y agoThat's pretty interesting. I've been frustrated in the past by both AWS and GoGrid (and I'm sure every other cloud provider) that keep incurring VPS instance costs even when the instance is shut off. I understand that even if I'm not using the VPS the resources need to be kept in reserve (in theory), but the solution of destroying and reprovisioning instances sucks pretty bad, way to time consuming if you are dealing with only a handful and operationalizing it is not not worth it. I'd love to move to a provider that let me provision an extra instance or two for either failover or testing/staging but not be charged for it if I wasn't running traffic to it. EDIT: I stand corrected, I might have been thinking of Rackspace's cloud (can't remember what it's called now) instead of AWS. But I know for a fact I am right on GoGrid (and pretty sure Azure) because I have a long email chain arguing about charges for provisioned instances in off states.
- james33 14y agoUnless AWS has changed something since I last used it (which admittedly has been at least a year), they don't charge when an instance is off except for storage.
- josegonzalez 14y agoWell, not exactly true. If you pay a the Heavy Utilization instances, you are charged regardless of whether you even have the instance allocated.
- zwily 14y agoThat's also not exactly true. If you purchase a reserved instance, you pay the up front price, and then the reduced hourly price for whenever the instance is running.
- zwily 14y agoTo anyone reading this later... I am wrong here. Heavy Utilization reservations are charged whether or not you have instances running. Medium and Light are only charged when running.
- jules 14y agoGiven the premium that Amazon charges, how does this compare to dedicated servers? And why are dedicated servers in the US so much more expensive than in Europe?
- aioprisan 14y agopower is cheaper in Europe than in the US
- PanMan 14y agoNope, it's the reverse, so that can't be the case. Some quick googling: http://en.wikipedia.org/wiki/Electricity_pricing http://en.wikipedia.org/wiki/Electricity_pricing US is between 8-17 cents. EU is 20-30+. See also http://blogs.platts.com/2012/11/20/electric_prices/ http://blogs.platts.com/2012/11/20/electric_prices/
- ddon 14y agohttp://www.energy.eu/ http://www.energy.eu/
- gizzlon 14y agoThat wikipedia article is not enough to back your claim. As you can see, it lists a lot caveats. In my limited knowledge, electricity pricing is quite complicated, and those numbers are probably not even close to what business and industry are actually paying. Edit: The blog-post is more convincing, but again, are those numbers really comparable? Maybe they are, I don't know enough about the subject, I just find it a little too simple to just compare numbers from different websites without deep knowledge of the topic. Here's another link, the prices differ (?): http://epp.eurostat.ec.europa.eu/statistics_explained/index.php/Energy_price_statistics http://epp.eurostat.ec.europa.eu/statistics_explained/index....
- druiid 14y agoTheir engineer answered earlier in the thread. It was more along the lines of they had a small team and needed to scale up quickly (which I think is probably most of these EC2 stories you see, really). He also said that now their engineering team has more breathing room and considering dedicated/colo would be in the cards.
- RyanGWU82 14y agoI'm Ryan Park -- I'm the Pinterest engineer quoted in the article. I'm happy to answer any questions about our setup. Just to clarify, the auto-scaling was specifically for a pool of web application servers. At the time I gathered the numbers, there were 80 servers in that pool. In the last few months we've been moving toward a service-oriented architecture, and we've been able to use the same code to auto-scale the internal services. Of course it's not possible to auto-scale stateful servers like databases, but it's still saving us a considerable amount of money. We implemented the auto-scaling in early 2012, so it's been in use for almost a year now. It only took about 2 weeks of engineering to build the system. It does need occasional maintenance, but it's still worth the effort given how much money it saves us.
- ntoshev 14y agoHow do you decide whether to run a spot or on-demand instance to handle dynamic load? What happens if spot instance cost suddenly spikes? Also, how do you use regions/availability zones?
- RyanGWU82 14y agoRight now we run everything in the US-East region, and we have all our services balanced across 4 availability zones. If there's a problem in a single AZ, it will affect every layer of our system, but only about 25% of the hosts in that layer. Some of our services are automatically resilient and can handle that easily. Others aren't so great, but we're working on more automatic failover. When we need more servers for an auto-scaled service, we open spot requests and start on-demand instances at the same time. For most services, we want to run about 50% on-demand and 50% spot. We have a watchdog process that continually checks what's running. It launches more instances whenever there aren't enough, and terminates instances when there are too many. So if the spot price spikes and a bunch of our spot instances are shut down, the watchdog will launch replacement instances on-demand. It will also request more spot instances once the price has dropped back to normal. In reality we don't often run into spot capacity issues -- maybe once a month, and it's almost never apparent to our users. I spoke about this in detail at AWS re:Invent last month, and the full talk is available online here: http://www.youtube.com/watch?v=73-G2zQ9sHU http://www.youtube.com/watch?v=73-G2zQ9sHU
- WestCoastJustin 14y agoHere is the AWS re: Invent STP 204: Pinterest Talks Rapid, Cost Effective Scaling on Amazon Web Services -- https://www.youtube.com/watch?v=73-G2zQ9sHU https://www.youtube.com/watch?v=73-G2zQ9sHU
- jkat 14y agoIn our investigation, we've found AWS to be around 5x more expensive. Being able to save 63% during off hours (or, let's just say, reducing AWS' bill an average of 30%) doesn't really seem to make much of a difference - and that's with paying a large amount upfront.
- cperciva 14y agoIf spot prices spike and spot instances are shut down, on-demand replacement instances are launched. Spot instances will be relaunched when the price goes back down. I heard this at AWS re:invent, and thought I must have misunderstood. I'm still confused as to how this can possibly be a good strategy. The pool of EC2 on-demand instances has a finite size -- it isn't magical -- and does hit its capacity limit from time to time. When there's high demand for EC2 instances -- say, when there's an outage in another AZ -- you're likely to see both spot prices going up and a lack of capacity in the on-demand pool. As a result, this strategy seems designed to only ask for on-demand instances at the times when they're least likely to be available.
- jedberg 14y agoSpot instances come from both leftover on-demand instances as well as unused reserved instances. So it's quite possible to run out of on-demand and still have a low spot price.
- cperciva 14y agoIt's possible, sure... but only if the people who find that they can't launch on-demand instances don't think to try spot instances instead.
- jedberg 14y agoMany people aren't set up to handle spot instances. You need to be much more resilient to single instance failures than when using on-demand or reserved instances.
- usaar333 14y agoI deal with this a lot with my work at PiCloud and you very rarely see a lack of capacity in the on-demand pool (across every availability zone). You actually see high spot prices due to what at first glance seems "irrational"; incredibly high bids on instance types, sometimes 3-4x over on-demand prices. I suspect such high bids are placed by customers of the spot instances who absolutely do not want their workload terminated early by Amazon and are willing to take the risk of paying more to run it to completion. If you are a webapp though, like Pinterest, you don't have this desire. Hence, it makes sense to dynamically switch.
- teeja 14y agoI'm wondering if you did (or anyone has done) any cost effectiveness on using SSD's rather than HD's. (Or maybe that's an order-of-magnitude too low to consider?)
- thezilch 14y agoNetflix has done such benchmarks [0]; TLDR: _their workload_ cut costs in half with hi1 instances. [0] http://techblog.netflix.com/2012/07/benchmarking-high-performance-io-with.html http://techblog.netflix.com/2012/07/benchmarking-high-perfor...
- jcampbell1 14y agoI can't imagine that SSDs would make any difference for web app servers. I am not familiar with app servers workload that uses a material amount of disk I/O. That being said, if I built custom app servers, I'd use SSDs because the cost is small for a system that doesn't need much storage.
- jacques_chester 14y agoReminds me a lot of how some energy-intensive plants can spin up less-efficient / quick-start units during offpeak hours to squeeze a little extra production out. And vice versa: most electrical utilities have slow-starting, efficient-as-possible turbines that never get turned off (baseload -- coal is most common), and a bunch of relatively inefficient but flexible turbines (usually natural gas).
- eru 14y agoNuclear also makes for great baseload. Wind or solar are similar in the sense, that you do not gain by turning them off, but their supply is not stable.
- jacquesm 14y agoNuclear is asymmetrical. It's very fast to shut down but horrendously slow to start back up again. In case of sudden load drops where nuclear plants will be shut down for safety reasons this can cause availability to be affected for weeks afterwards.
- brazzy 14y ago> relatively inefficient but flexible turbines (usually natural gas). Actually, natural gas plants are at least as efficient as coal. They're just more expensive, especially if you turn them on and off a lot, which is pretty bad for the lifetime of a lot of components.
- hga 14y agoDepends on the kind of plant, there's baseline gas plants that heat water to steam and thence to steam turbines, they're akin to coal plants, high capital costs, low operating costs. Then there are straight gas turbines, akin to the ones that power jet airplanes but more like the ones that power most of the US Navy's ships. These have lower capital costs but higher operating costs and are used for peaking.