5 ms·
My job is cost optimizations at a very large corporation. We have been given the order to go all in on AWS. Some things I've found to be particularly annoying:
by brodouevencode 7y ago
My job is cost optimizations at a very large corporation. We have been given the order to go all in on AWS. Some things I've found to be particularly annoying:
* Data transfer will bite you in the ass if you let it. Especially over NAT gateway in very high traffic sites. So you do the right thing and put your application in private subnets, route traffic in over the load balancers and out over the NATGW. Then you get a $20k/mo bill for your microserviced application that has a hundreds of requests per second during peak hours. Pro-tip: the poo-pooed nat instances are actually a cheaper solution, but you're on the hook for maintaining it.
* The CUR can get huge. I mean millions and millions of lines. AWS says you can throw it into S3, query with Athena, etc. etc. But if that data set is huge even _that_ will cost you a lot of money to run reporting, analysis, etc. Especially after you build that dashboard for the refresh happy VP.
* The Cost Explorer is admittedly getting better, but still lacking a lot of necessary detail. You have to pair it up with CloudWatch to get actual cost and usage in a usable way. The value add services like EMR/Elasticsearch service/all the ML stuff do the hideous job of hiding actual usage. You gotta dig hard.
* The third party cost tracking tools (CloudHealth/Metricly/CloudAbility/Cloudyn) are just a wrapper around what you can get out of the CUR. Their value-add is reporting and advisement, and giving recommendations on right sizing and reserved instances and savings plans. Though if your cloud team is sufficiently savvy they can do this themselves.
* No matter how you do your analysis, tagging will make your life so much easier. Can't emphasize this enough.
- ersii 7y ago> * No matter how you do your analysis, tagging will make your life so much easier. Can't emphasize this enough. Do you have any recommended resources for reading or tips on how and what to tag in what way for making ones AWS Life easier?
- ldoughty 7y agoTagging is simple*. The real killer is the things you can't tag, like bandwidth. Be sure to use multiple AWS accounts if you want to split bills (or at least track bills) to sub-groups like per department. Give each of these groups their own account (or set of accounts, preferably, if we're talking a business.. maybe even different accounts for Dev/preprod/prod, if this is a major cost and the project is worth it) For tags, you can make any tag you want and summarize bills by tags... So anything take can be tagged is trackable.. But things like bandwidth are not. It's also hard to enforce tagging when you can't automatically destroy non-complaint objects, so again, separate accounts help here.. if the sub-department wants to know their spend better, THEY are more likely to enforce the rule than A top-down policy from a disconnected IT group... And you can't simply apply a gonna "all things must be tagged" enforced in the AWS level because some items can't be tagged, or the tagging has to happen after creation (for instance, by SDK/cli, you can't create an ec2 instance with tags.. you make the instance, then tag it. The GUI does this behind the scenes so it looks like one step) So again, for major booking boundaries, use different accounts. After that point, it's on the delegated entities to use tags appropriately... And it's often different for each group anyway.
- dragonwriter 7y ago> It's also hard to enforce tagging when you can't automatically destroy non-complaint objects You can automatically destroy non-compliant (with your tagging policy) objects, by querying objects that exist and examining their tags through the API (heck, you could even script the CLI to do this), and, if you use AWS Organizations, you can prevent noncompliant resources with a combination of service control policies (to require tagging) and tag policies (to specify use of tags). > (for instance, by SDK/cli, you can't create an ec2 instance with tags.. you make the instance, then tag it. That's...not true. The runinstances call in the SDK that creates one or more instances from an AMI takes an optional set of tag specifications for tags that can be applied to the instances and/or any of a wide variety of associated resources. (python) https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/ec2.html#EC2.Client.run_instances https://boto3.amazonaws.com/v1/documentation/api/latest/refe... (Java) https://docs.aws.amazon.com/AWSJavaSDK/latest/javadoc/com/amazonaws/services/ec2/model/RunInstancesRequest.html https://docs.aws.amazon.com/AWSJavaSDK/latest/javadoc/com/am...
- tyingq 7y agoIt's not perfect, but you can tag and track some amount of egress charges now.
- brodouevencode 7y ago> And you can't simply apply a gonna "all things must be tagged" enforced in the AWS level because some items can't be tagged, or the tagging has to happen after creation (for instance, by SDK/cli, you can't create an ec2 instance with tags.. you make the instance, then tag it. The GUI does this behind the scenes so it looks like one step) The accepted approach is warn then terminate. Give them an hour and then if nothing's done start the slaughter.
- fphhotchips 7y agoI'm not the person you responded to, but as far as tagging strategy goes, here's a starting point. https://www.finops.org/blog/finops-tagging-automation-strategies/ https://www.finops.org/blog/finops-tagging-automation-strate... The entire FinOps foundation is good for cloud finance management - I believe the author of that post has an O'Reilly book coming out this month on the subject also.
- brodouevencode 7y agoSet the standard early, enforce with whatever means necessary. It's a common practice to use tools like CloudCustodian to terminate instances that do not have identified tags (with extreme prejudice). Also, normalize everything, and implement the standards/normalization using your build toolset (Jenkins/Terraform/CI du jour) to enforce this.
- fphhotchips 7y ago> The third party cost tracking tools (CloudHealth/Metricly/CloudAbility/Cloudyn) are just a wrapper around what you can get out of the CUR. Their value-add is reporting and advisement, and giving recommendations on right sizing and reserved instances and savings plans. Though if your cloud team is sufficiently savvy they can do this themselves. I work for one of the vendors you mentioned; and while you're not wrong - the data sources are all from the vendors - there's a fair bit of work that goes into actually making sense of it to the point you can give it to people actually causing the spend. Also there's work that goes into optimisation, so that we can bear the cost of your second point. Your last point is dead on though. For anyone doing cloud at any scale, tagging is non-optional if you want to do any kind of optimisation, chargeback or the like.
- Rapzid 7y agoNot sure the status on this thing but: https://github.com/ProTip/aws-elk-billing https://github.com/ProTip/aws-elk-billing Parses the detailed billing logs into elastic search. Raw, but a good starting point..
- caro_douglos 7y agoI’ve wondered if the reasoning behind poor ML usage info has to do with the api endpoints being so profitable that notion of even looking at the code and implementing more straight forward metrics might kill the golden goose.
- brodouevencode 7y agoIt's less that and more of the underlying technologies that make up those tools are what surface in the CUR. For instance, you want to know how much you're spending on ElasticBeanstalk. How do you do that? The easiest way is to just get a cost broken down by AMI, then look at the AMIs that are EB (though not the best solution it works most of the time). AI/ML tools do the same. It's hard to break them out (hence why tagging is so important).