Skip to content
David Galiata
Go back

The patterns behind infrastructure as code, and how Terraform implements them

The patterns behind infrastructure as code, and how Terraform implements them

Every infrastructure as code tool has to answer the same handful of questions. How do you describe what you want? How does the tool remember what it already built? How does it work out what to change? What happens when somebody edits things by hand in the console? Pulumi, CloudFormation, and Terraform answer those differently, but the questions are fixed, and a lot of what gets called Terraform expertise is really familiarity with the answers.

You describe the end state and Terraform works out the steps

You write resource blocks that say ≈what should exist. Terraform compares that against what it believes exists, computes a diff, and executes whatever operations close the gap.

resource "aws_s3_bucket" "logs" {
bucket = "acme-app-logs"
}
resource "aws_s3_bucket_versioning" "logs" {
bucket = aws_s3_bucket.logs.id
versioning_configuration {
status = "Enabled"
}
}

Nothing in there says create the bucket first and then enable versioning. Terraform derives that because the second resource references the first.

The practical effect is that your config is always the description of the destination. When you delete a block, you’re saying that thing should no longer exist, and Terraform plans a destroy. People coming from scripts trip on this: a deleted line in a shell script does nothing, a deleted block in Terraform is an instruction.

The state file is where most of the trouble lives

Terraform keeps a JSON record mapping each address in your config to the real resource it created. Without it, aws_s3_bucket.logs is just a name, and Terraform has no way to tell a bucket it made from one that was already there.

Three consequences fall out of that, and they cause most of the operational pain:

State has to be shared. Two engineers with two local state files will each think they own the infrastructure. A remote backend fixes it, and S3 is the common choice on AWS.

State has to be locked. Concurrent applies against the same state can corrupt it. The S3 backend used a DynamoDB table for this for years. Native S3 locking went generally available in Terraform 1.11 via the use_lockfile argument, and it’s one less piece of infrastructure to run.

State holds attribute values in plaintext. Every password, key, and generated secret your resources produce is sitting in that file. Encrypt the bucket, lock down who can read it, and treat state with the same care as the credentials inside it.

Plan is a dry run you can put in a pull request

terraform plan refreshes its view of reality, diffs it against your config, and prints what it intends to do. You can save that to a file and apply exactly that plan later, which matters in CI where the config might move between the two steps.

Terminal window
terraform plan -out=tfplan
terraform apply tfplan

The pattern most teams land on is plan on the pull request, apply on merge. The diff becomes the review artifact. A reviewer who can read a plan output catches the accidental destroy line before it runs, which is worth more than any policy engine you bolt on later.

The dependency graph comes from your references

Terraform builds a directed graph from the references between resources and walks it in order, running independent branches in parallel. Most of the time you never think about this because the references are implicit, like aws_s3_bucket.logs.id above.

When the dependency is real but invisible to Terraform, usually an IAM policy that has to exist before something can assume a role, depends_on makes it explicit. Reach for it sparingly. Every depends_on you add is a bit of ordering Terraform can no longer work out for itself, and over-specified graphs get slow and brittle.

Running apply twice should be a no-op

Idempotency is the property that makes the whole model work. Apply a config, apply it again, and the second run should report no changes. When that stops being true you have what people call perpetual diff, and it usually means one of three things: a provider bug, an attribute the API normalizes differently than you wrote it, or something outside Terraform changing the resource.

lifecycle.ignore_changes is the escape hatch, and it’s a trade. You’re telling Terraform to stop managing an attribute, so from then on that attribute drifts silently.

resource "aws_ecs_service" "api" {
# ...
lifecycle {
ignore_changes = [desired_count]
}
}

That one is common when an autoscaler owns the replica count. Write a comment next to every ignore_changes saying who owns the attribute instead, because the next person will not guess.

Modules earn their place when something repeats

A module is a directory of .tf files you call with inputs. The useful ones wrap a pattern you deploy more than once, like a VPC with a standard subnet layout, or an ECS service with the logging and alarms your team always wants.

Pin versions on anything from a registry:

module "vpc" {
source = "terraform-aws-modules/vpc/aws"
version = "~> 5.0"
# ...
}

The most common mistake I see discussed is modularizing too early, wrapping a single resource in a module that takes fifteen variables and passes them straight through. You get indirection without reuse, and debugging means opening two files instead of one. Write it flat first and extract when the second copy shows up.

Workspaces are a poor fit for separating prod from dev

Terraform workspaces give you multiple state files against one configuration and one backend. They work well for short-lived parallel copies of the same thing, like a per-branch test environment.

For real environment separation they fall down, because prod and dev usually differ in more than variable values: different accounts, different backends, different provider credentials, sometimes different resources entirely. HashiCorp’s own documentation says as much. Separate directories with separate backend configs is the boring answer, and it survives the day someone needs prod to have something dev doesn’t.

Replacement is usually safer than mutation

A lot of infrastructure changes can’t happen in place. Changing an EC2 instance’s AMI, or an RDS instance’s engine version in some cases, means the provider destroys and recreates.

create_before_destroy flips the order so the replacement exists before the old one goes away, which is what you want behind a load balancer:

lifecycle {
create_before_destroy = true
}

It isn’t free. The new resource has to be able to coexist with the old one, so anything with a unique name will collide unless you let Terraform generate the name with name_prefix.

Beyond the basics

The features below came in over the last few years, and several of them solve problems that used to mean editing state by hand.

moved, import, and removed blocks keep refactors from destroying things

Renaming a resource in your config used to mean Terraform planned a destroy and a create, because the address changed. The moved block, added in 1.1, tells Terraform the two addresses are the same resource:

moved {
from = aws_s3_bucket.logs
to = aws_s3_bucket.application_logs
}

import blocks arrived in 1.5 and replaced the old terraform import command for most uses. You declare the intent in configuration, so it goes through plan and review like everything else:

import {
to = aws_s3_bucket.legacy
id = "acme-legacy-bucket"
}

Terraform 1.5 also added configuration generation, so terraform plan -generate-config-out=generated.tf writes you a starting resource block for anything you’re importing. The output needs cleanup, but it beats hand-transcribing forty attributes out of the console.

removed blocks came in 1.7 and handle the opposite case: dropping a resource from your configuration without destroying the real thing.

removed {
from = aws_s3_bucket.legacy
lifecycle {
destroy = false
}
}

Between those three, most of the state surgery people used to do with terraform state mv and terraform state rm now happens in reviewable configuration. That’s a real improvement, because state commands leave no trace in git.

You can assert things about infrastructure in the configuration

Preconditions and postconditions landed in 1.2. They halt the run when an assumption breaks, which makes them good for catching bad inputs at the module boundary:

resource "aws_instance" "app" {
# ...
lifecycle {
precondition {
condition = data.aws_ami.selected.architecture == "x86_64"
error_message = "The selected AMI must be x86_64."
}
}
}

check blocks, added in 1.5, work differently. They run after apply and report problems as warnings without failing the run, and they can load their own scoped data source. That makes them suited to things you want to know about but don’t want blocking a deploy, like an endpoint that should be returning 200 after the change.

terraform test runs real plans against your modules

Terraform 1.6 added a native test framework. You write .tftest.hcl files with run blocks, each of which executes a plan or an apply and asserts on the result.

run "bucket_name_is_prefixed" {
command = plan
assert {
condition = aws_s3_bucket.logs.bucket == "acme-app-logs"
error_message = "Bucket name did not match the expected prefix."
}
}

By default a run block with command = apply creates real infrastructure and destroys it afterward, so these are slower and cost money. 1.7 added mock providers, which lets you assert on plan output without touching a cloud API. For module authors publishing to a registry, this is the difference between tested and hoped-for.

Ephemeral values keep secrets out of state

This is the one I’d point security people at. Terraform 1.10 introduced ephemeral resources and ephemeral values, which are read fresh during each phase and never written to the plan or the state file. 1.11 added write-only attributes, where a provider accepts a value on the way in and never returns it for storage.

Before this, pulling a database password from Vault or Secrets Manager into a resource meant that password landed in state in plaintext, and your options were to encrypt the whole state file and restrict access to it. The ephemeral path means the value is used during apply and then gone.

Provider support varies, so check whether the specific resource you care about exposes a write-only variant before you plan around it.

count and for_each diverge the moment a list changes

count addresses instances by numeric index. Remove the second item from a five-item list and everything after it shifts down by one, so Terraform plans to modify or replace four resources when you meant to delete one.

for_each keys instances by a string instead, so removing an item affects only that instance:

resource "aws_iam_user" "team" {
for_each = toset(["alice", "bob", "carol"])
name = each.key
}

Use count for a thing that exists zero or one times, usually behind a feature flag. Use for_each for collections. I’d call this the single most expensive mistake available in Terraform, because the plan output looks reasonable right up until you read it carefully.

Provider aliases handle multi-account and multi-region

One configuration can hold several copies of a provider, which is how you write something that spans regions or accounts:

provider "aws" {
alias = "us_east_1"
region = "us-east-1"
}
resource "aws_acm_certificate" "cdn" {
provider = aws.us_east_1
# ...
}

CloudFront certificates have to live in us-east-1 regardless of where everything else runs, so most people meet aliases through that particular annoyance.

Drift detection belongs on a schedule, not in your memory

Someone will change something in the console. A scheduled plan catches it:

Terminal window
terraform plan -detailed-exitcode

That exits 0 for no changes, 2 for changes pending, and 1 for an error, which is enough for a nightly CI job to open a ticket when reality and configuration separate. Running it on a schedule turns drift into something you find on Tuesday morning rather than during an incident.

Most of these features exist because people were doing the same dangerous thing by hand often enough that HashiCorp built a safer path. If you’re still reaching for terraform state rm, there’s probably a block for what you’re doing now.


Share this post:

Previous Post
GitHub Actions vs GitLab CI vs CircleCI: Jobs, Steps, and Stages Explained