13 ms·
Terraform best practices for reliability at any scale
- Terretta 3y agoThis should be mandatory reading for anyone doing IaC, using TF and AWS or not, less for how you do it, more for what and why. // shout out to AWS CAB alums
- pezh0re 3y agoThis is a great read, but I always seem to run into cases where I need to define something like a security group and then reference it when deploying ec2 instances. I'd love to decouple to reduce my plan time, but I haven't figured a way out as of yet. To be fair, I haven't used terraform -chdir yet.
- c0Re69 3y agoTry Terragrunt https://terragrunt.gruntwork.io/docs/ https://terragrunt.gruntwork.io/docs/
- JohnMakin 3y agoyou can pull it in via a data source, but then of course this creates a coupling between multiple modules/state files.
- OJFord 3y ago-chdir is useful, but nothing to do with this (it's literally just cd before running command). As sibling said, use data source to read remote state outputs.
- 36chamber 3y agonot sure what you mean here. Do you use the same security group for all of your instances? i am usually creating a security group per group of related ec2 instances.
- mdhb 3y agoSorry if this comes across weird or snotty it’s not supposed to. But I’m coming at this from a GCP lens and got half way through the article about how the recommended unit of isolation in the AWS environment is entirely different AWS accounts and I’m kind of hung up on that. Is that really a thing people tend to do often? Doesn’t it get super unwieldy? How does billing work? What about identity? I have so many questions. EDIT: Despite the fact that the root resource in both GCP and AWS is an organization, when I heard “account” I mistook that to be AWS terminology for an organization.
- dragonwriter 3y ago> But I’m coming at this from a GCP lens and got half way through the article about how the recommended unit of isolation in the AWS environment is entirely different AWS accounts and I’m kind of hung up on that. Is that really a thing people tend to do often? Doesn’t it get super unwieldy? There are AWS systems above the account level for managing it (Organizations), so its not quite as bad as it might naively seem, but, yes, its more unwieldy than GCP’s projects.
- mdhb 3y agoOh thank god, that’s much better than I naively thought. Thanks for the heads up.
- stock_toaster 3y agoYou can have sub-accounts that roll up billing to a main account. Still messy, but probably cleaner (security policy wise) and possibly safer (fewer production impacting accidental config changes?) than having a single giant account with _lots_ of things mixed together.
- mdaniel 3y agoAs the resident pedant, one cannot have "sub-accounts" in AWS. One can 100% have Accounts that are a member of an AWS Organization, which itself has a Management Account that does as you describe as the "main account", but there is no parent-child relationship between Accounts, only OUs and Accounts or OUs and other OUs (which the Organization, itself, counts as an OU)
- badblock 3y agoSome of this seems like old advice, instead of having directories per environment you should be using workspaces to keep your environments consistent so you don't forget to add your new service to prod.
- rcrowley 3y ago(Hi, I’m one of the authors of the article at the root of this thread.) I’ve gone back and forth on workspaces versus more root modules. On balance, I like having more root modules because I can orient myself just by my working directory instead of both my working directory and workspace. Plus, I feel better about stuffing more dimensions of separation into a directory tree than into workspace names. YMMV.
- dharmab 3y agoWhat do you think about multiple backends? It seems to be working well for me to have a single root module but with a separate backend configuration per environment.
- rcrowley 3y agoThat does work well for environments because typically you’d run exactly the same code, maybe with different cluster sizes or instance types, in each environment. But it doesn’t work well for isolating two services where the code is significantly or even entirely different.
- RulerOf 3y agoMultiple backends are unwieldy if you're using terraform at the command line, but they beat workspaces handily for discoverability. They're a fine option if you're applying through CI though, as the drudgery of utilizing them is handled effortlessly by the robots.
- 36chamber 3y agoDo you always store modules in the same repo as the terraform itself? Why not put them in seperate repos that can be tagged and versioned and then referenced like below? source = "git::https://bitbucket.org/foocompany/module_name.git?ref=v1.2 https://bitbucket.org/foocompany/module_name.git?ref=v1.2"
- swozey 3y agoThis was a good read but really if you already follow the common best practices of IAC/terraform/aws multi-account I don't think you're going to learn much. The comments in here kind of made me think I was going to hop in and take away some huge wins I hadn't considered. But I have been working with Terraform and AWS for a very long time. If you're unfamiliar with AWS multi-account best practices this is a good read. https://aws.amazon.com/organizations/getting-started/best-practices/ https://aws.amazon.com/organizations/getting-started/best-pr...
- bcjordan 3y agoI remember periodically coming across services/platforms that purport to make setting up secure AWS accounts / infra configuration easier and default secure — anyone know what I may be thinking of?
- swozey 3y agoActually the article here is one of those options - https://substrate.tools/ https://substrate.tools/ I don't know how integrating this into an environment where you already have tons of AWS accounts would go but it's interesting. Thankfully I only have to make new accounts when we greenfield a service and that's maybe a yearly thing.
- xyzzy123 3y agoHere's my #1 tip, most important: Try to keep your stateful resources / services in different "stacks" than your stateful things. Absolutely 100% completely obvious, maybe too obvious? Because none of these guides ever mention it. If you have state in a stack it becomes 10x more expensive and difficult to "replace your way out of trouble" (aka destroy and recreate as last resort). You want as much as possible in stateless, disposable stacks. DONT put that customer exposed bucket or DB in the same state/stack as app server resources. I don't care about your folder structure, I care about what % of the infra I can reliably burn to the ground and replace using pipelines and no manual actions.
- sverhagen 3y agoIs a "stack" here a (root) folder on which you'd do a "terraform apply"? I've never know what to call those, surely they aren't "modules". And, so, you're saying: try to have a separate deployment (stack then?) that contains the state, so you can wipe away everything else if you want to, without having to manage the state?
- xyzzy123 3y agoIt's not exactly about the folder, the IaC from a single folder / project can be instantiated in multiple places. Each time you do that, it has a unique state file, so I usually hear it referred to as a "state". In cfn you can similarly deploy the same thing lots of times and each instantiation is called a "stack", so stack/state tend to get used inter-changeably. And yes, that's a succinct rephrasing. When you first use iac it maybe seems logical to put your db and app server in the same "thing" (stack or state file) but now that thing is "pet like" and you have to take care of it forever. You can't safely have a "destroy" action in your pipeline as a last resort. If you put the stateful stuff in a separate stack you can freely modify the things in the stateless one with much less worry.
- swozey 3y agoCan you elaborate on this? I've never heard of this IAC structure and I'm trying to figure out what the benefit/cons are. Maybe it's just Friday and I'm checked out already. If you run a terraform apply and only update microservices but you also have your dbs/stateful things in the same stack/app, you're only updating the microservices so how would this affect the db/stateful at all? On the opposite end - I feel like there would be scenarios where I needed to update the stateful AND stateless services with the same terraform apply. Maybe I'm adding a new cluster and adding a db region/replica/securitygroup and that new cluster needs to point at the new db region. In your scenario I would have updated microservices trying to reach out to a db in a region that doesn't exist yet because I have to terraform apply two different stacks. How would you deal with a depends_on? Maybe I'm misunderstanding this.
- cube2222 3y agoThe article recommends to split up your state files for various advantages, but also expands into how to manage it later in a custom way. I agree with the splitting, but based on many home-grown automation systems I've seen around this I'd really recommend you to use one of the specialized CI/CD systems that are built around automating these kinds of workflows. Once you reach the "many state files" phase, you'll save a lot of engineering time this way. They'll take care of, among others, running the right state files, in the right order, with the right parameters. But they'll also take care of many other things you need to run Terraform at scale and with big amounts of engineers (happy to expand but don't want to kitchen-sink this comment). Disclaimer: Take this with a sensible grain of salt, as I work at Spacelift[0] - one of the TACOS (and of course the one I'll shamelessly link and recommend!). But really, don't use tools like Jenkins for this as you scale, it'll likely hurt you in the long run. [0]: https://spacelift.io https://spacelift.io
- swozey 3y agoI'm sure that you have no control over this but I really wish Spacelift would increase the cost of its cloud tier and lower the cost of Ent. I'm in the anti-goldilocks zone. Ent seems priced for large teams when I practically fit into the cloud offering sans missing a few required features. Great product though from what I've experienced.
- sausagefeet 3y agoDisclaimer: Co-founder of Terrateam. For Terrateam[0], we have probably 70% of the enterprise offering but at around 1/10th the price. If there are any features that are deal breaker, feel free to reach out to me and we'll see what we can do. That being said, Spacelift is a much more luxurious piece of software than us. We are very utilitarian, but we have to rationalize that low price-point somehow. [0] https://terrateam.io https://terrateam.io
- cube2222 3y agoSorry to hear that! Pricing is hard. If you haven't yet, please try talking to our sales team. There's usually a way to make all sides happy with some custom agreements - after all, we'd love for you to be able to use our product as much as you need.
- gerl1ng 3y agoThe solution at the end almost looks like the manual setup of terragrunt which we are using to manage lots of base infra in many different accounts. What would be interesting here would be to see how they actually reference the outputs from one layer onto the next layer. That is something that is not even solved nicely in terragrunt and one of the major annoyances for me there. Using dependencies and the mock_output option is creating lots of noise in the plan outputs as the dependencies are only completely resolved when terragrund applies all the modules. But it seems I also missed a few additions to terraform - so probably there are better ways to take outputs from one terraform run into another one.
- time0ut 3y agoI’ve been using Terragrunt [0] for the past three years to manage loosely coupled stacks of Terraform configurations. It allows you to compose separate configurations almost as easily as you compose modules within a configuration. Its got its own learning curve, but its a solid tool to have in the tool box. Gruntwork is a really cool company that makes other tools in this space like Terratest [1]. Every module I write comes with Terratest powered integration tests. Nothing more satisfying than pushing a change, watching the pipeline run the test, and then automatically release a new version that I know works (or at least what I tested works). [0] https://terragrunt.gruntwork.io/ https://terragrunt.gruntwork.io/ [1] https://terratest.gruntwork.io/ https://terratest.gruntwork.io/
- mike_d 3y agoThey seem very insistent on keeping things DRY but not explaining why. Does Terraform tend to cause water leaks?
- SgtBastard 3y agoDRY = Don’t Repeat Yourself.
- raffraffraff 3y agoTerraform is supposed to let you write modular, reusable code. But because it's a limited DSL that lacks many "proper language" features (and occasionally breaks the rule of least-suprise). There are several major impediments to fully data-driven terraform. These ultimately result in copy/paste code, or tools like terragrunt which essentially wrap terraform and perform the copy/pasta behind your back by generating that code for you. Some minor examples: - calling a module multiple times using `for_each` to iterate over data works, except if the module contains a "provider" block - if you are deploying two sets of resources by iterating over data, terraform can detect dependency cycles where there are not any
- thunfisch 3y agoWe're using Terragrunt with hundreds of AWS accounts and thousands of Terraform deployments/states. I'll never want to do this without Terragrunt again. The suggested method of referencing remote states, and writing out the backends will fall apart instantly at that scale. It's just way too brittle and unwieldy. Terragrunt with some good defaults that will be included, and separated states for modules (which makes partial applies a breeze) as well as autogenerated backend configs (let Terragrunt inject it for you, with templated values) is the way to go.
- DelightOne 3y agoDo you need to chain multiple Terragrunt executions to first bring the Kubernetes cluster up and then the containers, or does Terragrunt fix that?
- miduil 3y agoYes, with terragrunt you can do a `terragrunt run-all apply` and based on `output` to `variable` in each module data can be passed from one state/module to the next one, terragrunt knows how to run them in the right order so you can bootstrap your EKS cluster by having one module which bootstraps the account, then another one which bootstraps EKS, then one that configures the cluster, installs your "base pods" and then later everything else.
- ckdarby 3y agoHave you spent any time with Pulumi? I've kind of found terraform is dying and encourages a lot of bad practices but everyone agrees with them because HCL and it is transferable as most companies are just using TF.
- RulerOf 3y ago> I've kind of found terraform is dying I don't think it's dying. The hype has worn off. Everybody uses it. It's very mature. There's a module for everything. It's just not new and sexy anymore IMO.
- spicyusername 3y agoEverybody in here is recommending tarragrunt, but I'm not sure what value it provides over regular terraform. After using it for a few months all of the features found in tarragrunt are in terraform.
- jbjohns 3y agoThis is my impression as well. As far as I've understood, terragrunt was made back when terraform was missing a lot of key features (I think it maybe didn't even have modules yet) but when I was asked to evaluate it recently for a client I couldn't find a single reason to justify adding another tool.
- linuxdude314 3y agoThe primary thing terragrunt was designed to do was let you dynamically render providers. Terraform still does not let you this. It becomes very problematic when using providers that are region specific, amongst other scenarios. That being said I don’t like the extra complexity terragrunt adds and instead choose to adopt a hierarchical structure that solves most of the problems being able to dynamically render providers would solve. Each module is stored in its own git repo. Top layer or root module contains one tf file that is ONLY imports with no parameters. The modules being imported are called “tenant modules”. A tenant module contains instantiations of providers and modules with parameters. The nodules imported by the tenant modules at the ones that actually stand up the infrastructure. Variables are used, but no external parameters files are used at any level (except for testing). All of the modules are versioned with git tagged releases so the correct version can easily be imported. Couple this with a single remote state provider in the root module and throw it in a CI/CD pipeline and you have a gitops driven infrastructure as code pipeline.
- OJFord 3y agoWhat do you mean by 'dynamically render providers'? I assume you're aware you can instantiate multiple versions (different params) of a provider and pass them to child modules, e.g. you can instantiate a module once for each of a several AWS regions/accounts? Do you mean that something like the region/account param would be set on the basis of a computed value from some other resource (because we created the account, say, or listed all regions satisfying some filter with a data source)?
- waffletower 3y agoWhile combining the word "best with "Terraform" in a sentence is more than likely to result in an oxymoron, it is counter-productive not to attempt to organize and utilize terraform as elegantly and DRY as possible. We interact with stacks (which we call projects typically) via Terragrunt and have a very large surface of modules as we do have a fair amount of infrastructure pieces. But we also try to expose Terraform infrastructure changes by use of Atlantis; though bulky, github does provide a reasonable means to dialogue and manage changes made by multiple teams. The use of modules also helps us encapsulate infrastructure, and state problems are rare with these approaches, but the data sprawl inherent to Terraform is very unwieldy regardless of so called "best" practices. The language features are weak, awkward and directly encourage repetition and specification bloat. We have had some success via Data Sources to export logic outside of Terraform and provide much needed sanity when interacting with very verbose infrastructure such as Lake Formation.
- ary 3y ago> At scale, many Terraform state files are better than one. But how do you draw the boundaries and decide which resources belong in which state files? What are the best practices for organizing Terraform state files to maximize reliability, minimize the blast-radius of changes, and align with the design of cloud providers? 1000% agree. I put together my version of standing up remote state in AWS into a Github repo. https://github.com/aryounce/terraform-aws-bootstrap https://github.com/aryounce/terraform-aws-bootstrap Our use of Terraform splits state exactly as described primarily to keep the state refresh times reasonable.
- bickfordb 3y agoAside from reducing the blast radius of any Terraform state (split by envs, then by teams as you grow), I highly recommend using cdktf with Python for Terraform projects. Huge timesaver
- 333throwaway342 3y agoI don't know. I kind of think using a language with native JSON support and structural type system would be best. HCL also just works.
- rcarr 3y agoGenuine question for DevOps people: Other than the fact it seems to be an industry standard so it's good for your job prospects, what are the benefits to Terraform over CloudFormation/CDK or whatever the equivalent is for your particular cloud provider? Most companies/people pick a provider and then stick with it and it doesn't seem like there's much portability between configurations if you do decide to switch providers later down the line so I'm not sure what the benefits are. I haven't delved into Terraform yet but I tried doing a project in Pulumi once and felt by the end of it that I might as well have just wrote it in AWS CDK directly.
- jahsome 3y agoThird-party integrations and the universality/reusability across multiple products and familiarity of HCL are big for me.
- x3n0ph3n3 3y agoI have used both terraform and cloudformation substantially and they each have pros and cons. One thing terraform has over cloudformation is its rapid support for new services and features. AWS has done an awful job ensuring that cloudformation support is part of each team's definition of "done" for each release. It just doesn't get the support it really needs from AWS.
- Centigonal 3y agoI worked closely with the folks that wrote our platform's IaC, first in CDK, then in Terraform. I wrote a bit of CDK and zero TF myself, but here are some of the reasons we switched: A big plus is that Terraform works outside of AWS land. CDK is a nightmare to work with. You're writing with programming-language syntax, which tempts you to write dynamic stuff - but everything still compiles down to declarative CFN, which just makes the ergonomics feel limited. The L2 and L3 constructs have a lot of implicit defaults that came back to bite us later. With CDK you get synth and deploy, which felt like a black box. Minor changes would do the same 8 minute long deploy process as large infrastructure refactors. Switching to TF significantly sped up our builds for minor commits. There might be a better way to do this with CDK (maybe deploying separate apps for each part of our infrastructure) and we may have just missed it.
- paulddraper 3y agoWhat does the phrase "stamp out" mean in this context?
- jbjohns 3y agoRapidly create exact duplicates I think.
- aloknnikhil 3y agoIt seems to me, this is trying really hard to shoehorn Terraform into managing at scale. For multi-account, multi-org, multi-region, multi-cloud deployments is Terraform really supposed to be the state of the art? How do you even get visibility into the various deployment workflows?
- Too 3y agoWhat’s the alternative?
- ochoseis 3y agoWe’ve had a pretty good experience with Terraspace at work, which is an opinionated framework/layout for Terraform. It supports hooks and splitting state between regions and accounts.
- RulerOf 3y ago> Using the -target option to terraform plan is discouraged (the Terraform documentation says, “this is for exceptional use only”). Anyway, it’s likely to lead to confusing infrastructure states if changes are applied incrementally with ad hoc, situational boundaries. We've been using -target for years, and while I understand very well the reasons it's discouraged, it is pretty much the only way you can have your cake and eat it too with respect to having "one large terraform project" and not running into terraform refreshes that make your eyes bleed. You end up having to really understand your module structure to use it, but it let us develop some very elegant workflows around tasks like patching. We developed a ruby code base that utilizes rake and the hcl2json tool to automate terraform-based infrastructure workflows, using various libraries to handle and validate that applications are happy with what terraform is doing to their servers while it works. This gives us flexibility to run automations safely against a terraform code base that has been evolving since version 0.3 or so, before most of the mistakes were made often enough to come up with the best practices we have today.
- abledon 3y agoWhy doesn't Hashicorp provide official best practices like this?
- 333throwaway342 3y agoThis isn't a rated E for everyone practice. Hashicorp is focused on CI/CD/cloud/workspace driven workflows over monorepo `chdir` driven.
- datadeft 3y agoI settled on 1 subfolder, 1 “stack” (stage/app, for example dev/login/frontend). This gives us fast deploy time and easy and painless way to destroy/re-create. Databases could go a separate folder if state if we had any. The point of Terraform is to have configuration in version control not to have a giant unmanageable state file.
- smetj 3y agoThe problem with TF is that it lures people into trying to be smart and try to be beat the system after which things often become a personal challenge instead of a business requirement. A true nightmare for the next person in line. Also ... every declaritive language dreams of becoming a programming language ...
- 36chamber 3y agoI'm surprised the blog advocates embedded modules instead of remote ones stored in seperate git repos. This allows you to tag and version them, and therefore progressively update modules.
- de_keyboard 3y agoIf many different "states" are used for one big architecture, how are the boundaries between them managed? Under one state, Terraform will spot issues here during `plan`, but with many states issues will only appear after `apply`.