Clicking through the AWS console is fine for learning and terrible for anything you need to repeat. Terraform lets you describe infrastructure in files, review changes like code, and rebuild an environment on demand. This is how I like to structure a Terraform setup for AWS so it stays understandable as it grows.
Why infrastructure as code
Three things make it worth the effort:
- Repeatability. Staging and production come from the same definitions, so they stay alike.
- Review. A change to a security group is a pull request someone can read, not a click nobody noticed.
- Recovery. If an environment is lost or corrupted, you can recreate it instead of remembering how it was built.
A layout that stays readable
Separate reusable building blocks from the environments that use them:
infra/
modules/
network/ # VPC, subnets, routing
app-service/ # compute, load balancer, security groups
database/ # RDS, subnet group, parameter group
environments/
staging/
main.tf
variables.tf
backend.tf
production/
main.tf
variables.tf
backend.tf
Each environment has its own state and its own small main.tf that wires modules together. A mistake in staging can never touch production state.
Remote state from day one
Local state files are a single point of failure and make teamwork impossible. Store state in S3, with locking so two people (or two pipelines) cannot apply at once:
terraform {
required_version = ">= 1.6"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
backend "s3" {
bucket = "example-terraform-state"
key = "staging/terraform.tfstate"
region = "eu-west-1"
dynamodb_table = "terraform-locks"
encrypt = true
}
}
Enable versioning on the bucket so you can recover an earlier state. Newer Terraform releases can also lock directly in S3 with use_lockfile, which removes the need for the DynamoDB table, but the table approach still works everywhere.
Write small, focused modules
A module should do one job and expose a small interface. Here is the shape of a simple one:
# modules/app-service/variables.tf
variable "name" {
type = string
description = "Service name, used in resource names and tags"
}
variable "instance_type" {
type = string
default = "t3.small"
}
variable "subnet_ids" {
type = list(string)
}
# modules/app-service/outputs.tf
output "security_group_id" {
value = aws_security_group.app.id
}
And an environment consumes it:
module "app" {
source = "../../modules/app-service"
name = "orders-api"
instance_type = "t3.medium"
subnet_ids = module.network.private_subnet_ids
}
Resist the urge to wrap every single resource in a module. Premature abstraction produces modules with thirty variables that nobody understands. Extract a module when you have used the same pattern twice.
Plan on every pull request, apply on merge
The safest workflow puts a human in front of every change. Run terraform plan automatically on pull requests so reviewers see exactly what will change, and apply only after merge:
name: terraform
on:
pull_request:
paths: ["infra/**"]
push:
branches: [main]
paths: ["infra/**"]
permissions:
id-token: write # for OIDC
contents: read
pull-requests: write
jobs:
plan:
runs-on: ubuntu-latest
defaults:
run:
working-directory: infra/environments/staging
steps:
- uses: actions/checkout@v4
- uses: hashicorp/setup-terraform@v3
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
aws-region: eu-west-1
- run: terraform init -input=false
- run: terraform fmt -check
- run: terraform validate
- run: terraform plan -input=false -out=tfplan
- if: github.ref == 'refs/heads/main'
run: terraform apply -input=false tfplan
The important detail is the role-to-assume line. GitHub Actions can authenticate to AWS through OpenID Connect, so there are no long-lived access keys stored as secrets. The role is scoped to what the pipeline needs and nothing more.
Habits that prevent pain
- Pin versions. Pin Terraform and provider versions so an upgrade is a deliberate change.
- Tag everything. A consistent
Environment,OwnerandServicetag set makes cost reports and cleanup possible. The provider'sdefault_tagsblock applies tags to every resource at once. - Never edit state by hand unless you understand exactly why. Prefer
terraform importandterraform state mvfor adopting or renaming resources. - Keep secrets out of code. Reference AWS Secrets Manager or SSM Parameter Store, and remember that state can contain sensitive values, so restrict who can read the bucket.
- Watch for drift. If someone changes a resource in the console, the next plan will show it. Treat unexpected diffs as a conversation to have, not an annoyance to apply away.
- Read the plan, every time. Pay special attention to anything marked for destroy or replace.
Starting from existing infrastructure
You do not have to rebuild what exists. Write the resource block, then bring the real resource under management:
terraform import aws_s3_bucket.assets my-existing-bucket-name
terraform plan # should show no changes once your code matches reality
Do this one resource group at a time, running plan until the diff is empty.
The short version
Remote state with locking, one state per environment, small modules, a reviewed plan on every change, and OIDC instead of stored keys. That combination is enough to make infrastructure changes boring, which is exactly what you want from infrastructure.