13 ms·
AWS Tips I Wish I'd Known Before I Started
- deleted 13y ago[deleted]
- lfuller 13y agoYour body tag is set to "overflow: hidden;". I wasn't able to scroll until I tweaked it manually in the inspector.
- richadams 13y agoOops, sorry about that. Should be fixed now.
- paulgb 13y agoAlso, if you change the first line of http://wblinks.com/css/style.css http://wblinks.com/css/style.css to @import url(http://fonts.googleapis.com/css?family=Droid+Sans:400,700); you should notice an improvement in the boldface font rendering. Great article, btw.
- riffraff 13y agoOT, but how does this work? Why does it improve rendering? Is it downloading both the normal and bold weight rather than only one leaving the browser to do the rest?
- paulgb 13y agoYes, you have it exactly right. The browser will try to make fonts bold on its own, but it's not a match for the bold that the designer intended.
- pmh 13y ago>Is it downloading both the normal and bold weight rather than only one leaving the browser to do the rest? Yes. The original URL only provides the normal weight (400), but you can specify other weights/styles[1]. 700 is the weight for bold and GP's URL requests both 400 and 700. [1]: https://developers.google.com/fonts/docs/getting_started#Syntax https://developers.google.com/fonts/docs/getting_started#Syn...
- richadams 13y agoThanks for the tip! CSS updated.
- alexscript 13y agoSince we're giving a little help here, you have two typos in this sentence. > You can't change bucket names one you've creatd them, so you'd have to copy everything to a new bucket. once and created
- richadams 13y agoThanks, fixed.
- SixSigma 13y agoI can't zoom on Firefox mobile for Android
- krallin 13y agoLots of very useful tips there! There's one that I think could be improved on a little: Uploads should go direct to S3 (don't store on local filesystem and have another process move to S3 for example). You could even use a temporary URL[0,1] and have the user upload directly to S3! [0]: http://stackoverflow.com/questions/10044151/how-to-generate-a-temporary-url-to-upload-file-to-amazon-s3-with-boto-library http://stackoverflow.com/questions/10044151/how-to-generate-... [1]: http://docs.aws.amazon.com/AmazonS3/latest/dev/PresignedUrlUploadObject.html http://docs.aws.amazon.com/AmazonS3/latest/dev/PresignedUrlU...
- jaibot 13y agoI have a desktop client that requests one-time upload URLs from my server via an API. Later they get downloaded and processed somewhere else - never actually touching my web server.
- toomuchtodo 13y agoI've always seen issues pushing objects directly to S3 from a browser using CORS. YMMV.
- ceejayoz 13y agoYou can specify CORS headers for S3, or you can just use a standard form POST.
- toomuchtodo 13y agoYou still need a stub API for generating the signature to sign the upload requests to S3, correct?
- ceejayoz 13y agoNot technically, but generally in practice. You can open things up permissions-wise but run the risk of folks uploading lots of large files. Keeping permissions locked down and doing a signature allows things like file size, location, etc. restrictions.
- drob 13y agoAlong these lines, I recommend installing New Relic server monitoring on all your EC2 instances. The server-level monitoring is free, and it's super simple to install. (The code we use to roll it out via ansible: https://gist.github.com/drob/8790246 https://gist.github.com/drob/8790246) You get 24 hours of historical data and a nice webUI. Totally worth the effort.
- PhilipA 13y agoReally useful article, though I don't agree with not using a CDN instead of S3. There are multiple articles which proves the performance of S3 being quite bad, and not useful for serving assets, comparing to CloudFront.
- tyw 13y agoalso the outbound bandwidth cost of S3 is very high. it would cost us several times what we're paying for s3+cloudfront to serve our content straight from s3.
- tedivm 13y agoEven cloudfront is ridiculously overpriced for a CDN. If you're pushing anything close to real bandwidth you could do a lot better elsewhere.
- Consultant32452 13y agoIf you're pushing anything close to real bandwidth you're probably getting an individually negotiated price that is not public.
- tedivm 13y agoNot with Amazon you aren't, they are amazingly stringent in their pricing on this matter.
- tyw 13y agohttp://aws.amazon.com/cloudfront/pricing/ http://aws.amazon.com/cloudfront/pricing/ the reserved capacity pricing is much better than the on demand pricing. Basically like EC2 on demand vs reservations. We set our reserved capacity at about 70-80% of what we expect to use most of the time. We could probably shave a few tenths of a cent per gig off but we get a good price on everything above what we've reserved so it's worked out. If you use a lot of cloudfront bandwidth without setting up a reservation, yeah... you're gonna pay through the nose.
- mbreese 13y agoI'd also add to the list - make sure that AWS is right for your workload. If you don't have an elastic workload and are keeping all of your servers online 24/7, then you should investigate dedicated hardware from another provider. AWS really only makes sense ($$) when you can take advantage of the ability to spin up and spin down your instances as needed.
- AaronBBrown 13y agoAWS makes a lot of sense in a 24/7 environment, particularly when you are a new startup and don't have enough information (or capital!) to make educated server purchases.
- dangrossman 13y agoThey don't have to make any purchasing decisions. They can rent servers from companies like Softlayer and Rackspace (#3 and #4 behind AWS for YC startups), or spin up much cheaper VPS's (Linode's #2). We're talking $120/month commitments, not buying hardware and driving to a data center to install it. Deploying to a freshly imaged physical server is the same as deploying to EC2, and they can be provisioned for you in an hour or two. Each of those servers gets you many times the performance of an EC2 instance in the same price class, which means much more time to figure out your capacity needs as you grow.
- AaronBBrown 13y agoAs someone who has worked w/ AWS (and Rackspace) for several years with multiple startups... Unless they have dramatically improved their offering in the last couple years, an hour or two from "I need a new server" to delivery is 1) not an accurate timeframe for physical servers from Rackspace and 2) even if it was realistic, that's an eternity when you are trying to iterate quickly. I can have a new server in 30 seconds with AWS and in the course of an hour could have tested my automation tools half a dozen times or more vs reimaging a server over and over again. I'm not saying it's always the right choice or that it's cheaper, but that flexibility combined with some of the pre-canned tools (ELB, RDS, CloudWatch, SQS, SNS) has tremendous value even when you aren't autoscaling.
- j-kidd 13y agoGood article, but I think it touches too little about persistence. The trade-off of EBS vs ephemeral storage, for example, is not mentioned at all. Getting your application server up and running is the easiest part in operation, whether you do it by hand via SSH, or automate and autoscale everything with ansible/chef/puppet/salt/whatever. Persistence is the hard part.
- crescentfresh 13y agoGood point. We're struggling to see the benefits of EBS for Cassandra that has its own replication strategy (ie data is not lost if an instance is lost), voiding the "only store temporary data on ephemeral stores" argument.
- blakesmith 13y agoHow do you handle entire datacenter outages with ephemeral only setup? You can replicate to another datacenter, but if power is lost to both do you just accept that you'll have to restore from a snapshotted backup?
- ismarc 13y agoOur use-case is different (not cassandra or db hosted on ephemeral drives), but what we've found using AWS for about 2 years now is that when an availability zone goes out, it's either linked to or affects EBS. Our setup now is to have base-load data and PG WAL files stored/written to S3, all servers use ephemeral drives, difference in data is loaded at machine creation time, AMI that servers are loaded from is recreated every night. We always deploy to 3 AZs (2 if that's all a region has) with Route 53 latency based DNS lookups that points to an ELB that sits in front of our servers for the region (previously had 1 ELB per AZ as they used DNS lookups to determine where to route someone amongst AZs and some of our sites are the origin for a CDN, so it didn't balance appropriately...this has since been changed) that is in the public+private section of a VPC with all the rest of our infrastructure in the private section of a VPC (VPC across all 3 AZs). We use ELBs internal to the AZ for services that communicate with each other. The entire system is designed to where you can shoot a single server, a single AZ or a single region in the face and the worst you have is degraded performance (say, going to the west coast from the east coast, etc.). Using this type of setup, we had 100% availability for our customers over a period of 2 years (up until a couple of weeks ago where the flash 12 upgrade caused a small amount of our customers to be impacted). This includes the large outage in US East 1 from the electrical storm, as well as several other EBS related outages. Overall costs are cheaper than setting up our own geo-diverse set of datacenters (or racks in datacenters) thanks to heavy use of reserved instances. We keep looking at the costs and as soon as it makes sense, we'll switch over, but will still use several of the features of AWS (peak load growth, RDS, Route 53). The short answer is to design your entire system so that any component can randomly be shot in the face, from a single server to the eastern seaboard falling into the ocean to the United States immediately going the way of Mad Max. Design failure into the system and you get to sleep at night a lot more.
- match 13y ago> Use random strings at the start of your keys. > This seems like a strange idea, but one of the implementation details > of S3 is that Amazon use the object key to determine where a file is physically > placed in S3. So files with the same prefix might end up on the same hard disk > for example. By randomising your key prefixes, you end up with a better distribution > of your object files. (Source: S3 Performance Tips & Tricks) This is great advice, but just a small conceptual correction. The prefix doesn't control where the file contents will be stored it just controls where the index to that file's contents is stored.
- 5ersi 13y agoAww man, my head hurts just looking at this list. Just go with a PaaS, like Heroku or AppEngine, and forget about this sysadmin crap.
- q3k 13y ago> sysadmin crap Without this “sysadmin crap” you would not have your precious PaaS.
- rkalla 13y agoFantastic list with much more depth than I expected. Some surprises that others might be interested in from this article and comments below: [1] Keeping buckets locked down and allowing direct client -> S3 uploads [2] Using ALIAS records for easier redirection to core AWS resources instead of CNAMES. [3] What's an ALIAS? [-] Using IAM Roles [4] Benefits of using a VPC [-] Use '-' instead of '.' in S3 bucket names that will be accessed via HTTPS. [-] Automatic security auditing (damn, entire section was eye-opening) [-] Disable SSH in security groups to force you to get automation right. [1] http://docs.aws.amazon.com/AmazonS3/latest/dev/PresignedUrlUploadObject.html http://docs.aws.amazon.com/AmazonS3/latest/dev/PresignedUrlU... [2] http://docs.aws.amazon.com/Route53/latest/DeveloperGuide/CreatingAliasRRSets.html http://docs.aws.amazon.com/Route53/latest/DeveloperGuide/Cre... [3] http://blog.dnsimple.com/2011/11/introducing-alias-record/ http://blog.dnsimple.com/2011/11/introducing-alias-record/ [4] http://www.youtube.com/watch?v=Zd5hsL-JNY4 http://www.youtube.com/watch?v=Zd5hsL-JNY4
- jamiesonbecker 13y agoI like SSH. But I'm the founder of Userify ;) http://userify.com http://userify.com Also, S3 buckets cannot scale infinitely. This is a huge myth http://aws.typepad.com/aws/2012/03/amazon-s3-performance-tips-tricks-seattle-hiring-event.html http://aws.typepad.com/aws/2012/03/amazon-s3-performance-tip...
- freerobby 13y agoCan you (or somebody else) elaborate on disabling ssh access? Is this a dogma of "automation should do everything" or is there a specific security concern you are worried about? What is the downside of letting your ops people ssh into boxes, or for that matter of their needing to do so?
- pavel_lishin 13y agoBased on the article, it seems it's there to make sure that you're automating everything, instead of logging in to do that one little thing by hand.
- richadams 13y agoThis is correct. The tip about disabling SSH isn't about security, it's just about quickly highlighting areas where you're not automated. When developing an application for example, it's often necessary to SSH in to play with some things. But once you've ready to go to production, you want as much automation as possible. Forcing yourself to not use SSH will quickly show you where you aren't automated.
- freerobby 13y agoThanks. I'm a fan of automation but respectfully disagree with this (see my response above for details).
- richadams 13y agoPerfectly valid. This particular tip certainly seems to have caused some great discussion! It worked for my particular case, but I can definitely see it not working for everyone. I've added a link to this thread to my tip, and expanded on it a little to warn people that it's not for everyone.
- freerobby 13y agoThanks for the reply here and above! Good discussion indeed.
- noelherrick 13y ago> Have tools to view application logs. Yes! Centralized logging is an absolute must: don't depend on the fact that you can log in and look at logs. This will grow so wearisome.
- txttran 13y agoWhat tools do you recommend for centralized logging?
- Fasebook 13y agoWhat's the point of auditing security in the Cloud? Is there any point at which you can know that your making any progress?
- mscarborough 13y agoJust one example -- Amazon will sign a Business Associate's Agreement for HIPAA compliance. That doesn't absolve you of your application security responsibilities, but it does give you piece of mind on the PAAS EC2/S3 side of things.
- tel 13y agoFor further note though, they won't unless you buy dedicated instances. This also disables RDS.
- mslot 13y agoBe very careful with assigning IAM roles to EC2 instances. Many web applications have some kind of implicit proxying, e.g. a function to download an image from a user-defined URL. You might have remembered to block 127.0.0.*, but did you remember 169.254.169.254? Are you aware why 169.254.169.254 is relevant to IAM roles? Did you consider hostnames pointed to to 169.254.169.254? Did you consider that your HTTP client might do a separate DNS look-up? etc. There are other subtleties which make roles hard to work with. The same policies can have different effects for roles and users (e.g., permission to copy from other buckets). IAM Roles can be useful, especially for bootstrapping (e.g. retrieving an encrypted key store at start-up), but only use them if you know what you're doing. Conversely, tips like disabling SSH have negligible security benefit if you're using the default EC2 setup (private key-based login). It's really quite useful to see what's going on in an individual server when you're developing a service. Also, it does matter whether you put a CDN in front of S3. Even when requesting a file from EC2, CloudFront is typically an order of magnitude faster than S3. Even when using the website endpoint, S3 is not designed for web sites and will serve 500s relatively frequently, and does not scale instantly.
- richadams 13y agoGreat point with regards to IAM roles. The applications I've worked on don't download things from user-defined URLs, so this never even occurred to me. Is the purpose of blocking 169.254.169.254 important because it could potentially give users access to the instance metadata service for your instance? I'd be interested to hear more information on securing EC2 with regards IAM roles, you seem to have lots of experience in that area. The disabling SSH tip wasn't really about security (I agree that it has negligible security benefit), it's more about quickly highlighting parts of your infrastructure that aren't automated. It's often tempting to just quickly SSH in and fix this one little thing, and disabling it will force you to automate the fix instead. The CDN info has been mentioned elsewhere too, lots of things I didn't know. I'll be updated the article soon to add all of the points that have been made. Thanks for the tips!
- mslot 13y agoThe way IAM roles work is terrifyingly simple. IAM generates temporary access key identifier and secret access key with the configured permissions and EC2 makes them available to your instance via the instance metadata as JSON at http://169.254.169.254/latest/meta-data/iam/security-credentials/ http://169.254.169.254/latest/meta-data/iam/security-credent... . The AWS SDK periodically retrieves and parses the JSON to get the new credentials. That's it. I'm not entirely sure whether the credentials can be used from a different IP or not, but given a proxying function that does not really matter. I make sure all HTTP requests in my (Java) application go through a DNS resolver that throws an exception if: ip.isLoopbackAddress() || ip.isMulticastAddress() || ip.isAnyLocalAddress() || ip.isLinkLocalAddress() The last clause captures 169.254.169.254. Of course, many libraries use their own HTTP client, so it's easy to make a mistake. I'm trying to bring my usage of IAM roles down to 0 as a matter of policy. Currently, I'm only using an IAM role to retrieve an encrypted Java key store from S3 (key provided via CloudFormation) and encrypted AWS credentials for other functions (keys contained in the key store). I'd be happier to bootstrap using CloudFormation with credentials that are removed from the instance after start-up. Thanks for making updates. There are definitely some great tips in there.
- gesman 13y agoSomeone needs to create such list for Azure as well. And make it Wiki-ized.
- Judson 13y agoOne thing the article mentions is terminating SSL on your ELB. If you want more control over your SSL setup AND want to get remote IP information (e.g. X-Forwarded-For) ELB now supports PROXY protocol. I wrote a little introduction on how to set it up[0]. They haven't promoted it very much, but it is quite useful. [0]: http://jud.me/post/65621015920/hardened-ssl-ciphers-using-aws-elb-and-haproxy http://jud.me/post/65621015920/hardened-ssl-ciphers-using-aw...
- richadams 13y agoGreat post, I had no idea you could do this with ELB. I've added your link to the additional reading list in my post, thanks for sharing!
- michaelmior 13y agoDisabling SSH is an interesting tip. I guess the OP doesn't do any automation via SSH.
- richadams 13y agoJust disabling inbound SSH connections, the servers can still SSH out to other systems to pull in files, configurations, clone git repos, etc. It's just a way to stop yourself from cheating and SSHing in just to fix that one thing, instead of automating it.
- milkshakes 13y agoexcept that some automation frameworks rely on inbound ssh access to the machines. ansible would be an example of such a framework, in its default configuration at least.
- richadams 13y agoAh, I wasn't aware of that, very good point! The goal of the tip is really to stop users SSHing in just to fix that one little thing, so you could still allow your automation frameworks SSH access and just disable it for users.
- michaelmior 13y agoIt can also be useful to SSH into a system to check what's going on with a specific problem. Sometimes weird things happen that you can't always anticipate or automate away.
- jamiesonbecker 13y agoUserify is awesome for this - disable SSH user accounts at any time and then re-enable when you realize you still need SSH to find out why your instance stopped sending logs!! ;)
- michaelmior 13y ago
- Mizza 13y agoThat '.' instead of '-' tip for SSL'd buckets just saved me a large future headache. Good stuff!
- croddin 13y agoI think you reversed them.
- simonlebo 13y agoCan anyone explain how disabling ssh has anything to do with automation? We automate all our deployments through ssh and I was not aware of another way of doing.
- jessaustin 13y agossh is handy if you're creating instances and then setting them up. However, if you're doing that on a regular basis, you might ought to use custom AMIs instead. Then (with proper "user data" management) you can just roll out instances that are already set up how you want.
- ceejayoz 13y agoI believe the idea is that by preventing SSH the temptation to just pop in and tweak something manually isn't possible.
- richadams 13y agoYup, this was the intention. You could still allow your automation processes SSH access, just disable it for your users. The idea is that if a user can't SSH in (at least not without modifying the firewall rules to allow it again), it will force them to try and automate what they were going to do instead. It worked well for me, but it's probably not for everyone.
- ape4 13y agoWow looks like a big pain.
- novaleaf 13y agoi'm a devops noob. what tools should i use to log / monitor all my servers? i don't want to learn some complex stuff like cheff/puppet btw.... anything SIMPLE?
- carlio 13y agoThough I haven't tried it, people tell me that ansible is pretty simple - http://www.ansible.com/home http://www.ansible.com/home For logging, try logstash? http://logstash.net/ http://logstash.net/ Monitoring... well that's a large and complicated topic!
- mblaney 13y agoAs an Australian developer, using an EC2 instance seems to be the cheapest option if you want a server based in this country. Anyone got any other recommendations?
- kibibu 13y agoNinefold aren't bad either
- mootpointer 13y agoAs a Ninefold employee, I'd like to think we're pretty good. We do virtual servers and we have a solid Rails platform as well.
- mblaney 13y agothanks will keep them in mind.
- late2part 13y agoThing I wish I'd known before I started: Don't rely on proprietary AWS solutions when open source solutions work just as well.
- rdl 13y agoI'd probably also say "avoid ELB where possible, especially for instance storage" and "avoid ELB, roll your own."
- Fizzer 13y ago> you pay the much cheaper CloudFront outbound bandwidth costs, instead of the S3 outbound bandwidth costs. What? CloudFront bandwidth costs are, at best, the same as S3 outbound costs, and at worse much more expensive. S3 outbound costs are 12 cents per GB worldwide. [1] CloudFont outbound costs are 12-25 cents per GB, depending on the region. [2] Not only that, but your cost-per-request on CloudFront way more than S3 ($0.004 per 10,000 requests on S3 vs $0.0075-$0.0160 per 10,000 requests on CloudFront) [1] http://aws.amazon.com/s3/pricing/ http://aws.amazon.com/s3/pricing/ [2] http://aws.amazon.com/cloudfront/pricing/ http://aws.amazon.com/cloudfront/pricing/
- richadams 13y agoDoh, I feel stupid now. I only looked at bandwidth costs, not the request prices. That's what I get for editing my post late at night based on reading, instead of based on personal experience. For low bandwidth, you're absolutely right, the costs are at best the same. For high bandwidth however (once you get above 10TB), CloudFront works out cheaper (by about $0.010/GB, depending on region). But that wasn't taking into account the request cost, which as you point out, is more expensive on CloudFront, which can negate the savings from above depending on your usage pattern. I'll update my post accordingly, thanks for pointing this error out!
- _hyn3 13y agoYou do have to pay for S3 to CloudFront traffic, so really you're paying twice. (Although the S3 to CF traffic might be cheaper than S3 to Internet, according to the Origin Server section on the Cloudfront pricing page.) http://aws.amazon.com/cloudfront/pricing/ http://aws.amazon.com/cloudfront/pricing/ Also, S3 buckets cannot scale infinitely. They have to have their key names managed appropriately to do it. http://aws.typepad.com/aws/2012/03/amazon-s3-performance-tips-tricks-seattle-hiring-event.html http://aws.typepad.com/aws/2012/03/amazon-s3-performance-tip... Finally :) I like SSH. But I'm the founder of Userify! http://userify.com http://userify.com
- Estragon 13y agoHow hard is it to roll your own version of AWS's security groups? I want to set up a Storm cluster, but the methods I have come up with for firewalling it while preserving elasticity all seem a bit fragile.
- jamiesonbecker 13y agoCheck out Dome9. Amazing tool and I think they work with both AWS and elsewhere.
- jamiesonbecker 13y agoWith regards to managing ssh, keys, etc... userify. Disclaimer: founder.
- kolev 13y agoOne painful to learn issue with AWS is the limits of services, which some of them are not so obvious. Everything has a hard limit and unless you have the support plan, it can take you days and weeks to get those lifted. They are all handled by the respective departments and lifted (or rejected) one by one. Many times we've encountered a Security Group limit right before a production push or other similar things. Last, but not least, RDS and CloudFront are extremely painful to launch. I have many incidents where RDS was taking nearly 2 hours to launch - blank multi-AZ instance! CloudFront distributions take 30 minutes to complete. I hate those two taking so long as my CloudFormation templates pretty much take an excess of an hour due to the blocking RDS and CloudFront. Last, but not least - VPC is nice, I love it, but it takes time to get what's the difference between Network ACL and Security groups and especially - why the neck do you need to run NATs?! Why isn't this part of the service?! They provide some outdated "high" availability scripts, which are, in fact, buggy, and support only 2 AZs. Also, a CloudFront "flush" takes over 20 minutes - even for empty distributions! Also, you can't do a hot switch from on distribution to another as it also take 30 minutes to change a CNAME and you cannot have two distributions having the same CNAME (it's a weird edge case scenario, but anyway).
- kolev 13y agoJust recalled another big annoyance! CloudFormation allows you to store JSON files in the user data, which is a bit similar to CloudInit, but... it turns your numbers into strings! So, imagine you need to put some JSON config file in there and the software expect an integer and craps out if there's a string value instead. I won't even bring how limited and behind the API CloudFormation is... Even their AWS CLI is behind and doesn't support major services like CloudFront. They even removed the nice landing page of the CLI took, which made it very obvious which services are NOT supported - I guess they just got embarrassed by having so many unsupported ones!