🏗Terraform
Define your entire cloud estate as code. Learn the plan/apply lifecycle, state management, module composition, remote backends, and how to run production-grade multi-account environments.
Architecture & How It Works 05.1
Terraform builds a dependency graph (DAG) from your configuration. It compares desired state (config) to actual state (state file), calculates the minimal set of changes, then executes them in order.
Providers are Plugins
Providers are downloaded during init. They translate HCL into API calls. Always pin provider versions.
State is Sacred
State is the only record of what Terraform manages. Never edit it manually. Use remote backends with locking.
Core Components & Concepts 05.2
| Element | Purpose | Common Pattern |
|---|---|---|
| resource / data | Manage or read infrastructure | aws_vpc, data "aws_ami" |
| module | Reusable, versioned packages | terraform-aws-modules/vpc/aws |
| variable / output / local | Inputs, outputs, computed values | var.region, output.vpc_id |
| workspace | Isolate state per environment | dev / staging / prod |
| backend "s3" | Remote state + locking | Bucket + DynamoDB table |
| count / for_each / dynamic | Meta-arguments for repetition | Creating N subnets |
WSL Hands-On Lab 05.3
Complete VPC + Subnets + EKS-ready + Workspaces + Import
$ wget -O- https://apt.releases.hashicorp.com/gpg | gpg --dearmor | sudo tee /usr/share/keyrings/hashicorp-archive-keyring.gpg > /dev/null $ echo "deb [signed-by=/usr/share/keyrings/hashicorp-archive-keyring.gpg] https://apt.releases.hashicorp.com $(lsb_release -cs) main" | sudo tee /etc/apt/sources.list.d/hashicorp.list $ sudo apt update && sudo apt install terraform
terraform {
required_version = ">= 1.6.0"
required_providers {
aws = { source = "hashicorp/aws", version = "~> 5.0" }
}
backend "s3" {
bucket = "tf-state-mycompany-98765"
key = "env/${terraform.workspace}/vpc.tfstate"
region = "us-east-1"
dynamodb_table = "tf-locks"
encrypt = true
}
}
provider "aws" { region = var.region }
module "vpc" {
source = "terraform-aws-modules/vpc/aws"
version = "5.5.0"
name = "devops-${terraform.workspace}"
cidr = "10.0.0.0/16"
azs = ["us-east-1a","us-east-1b"]
public_subnets = ["10.0.1.0/24","10.0.2.0/24"]
private_subnets = ["10.0.101.0/24","10.0.102.0/24"]
enable_nat_gateway = true
}$ terraform init $ terraform workspace new dev $ terraform plan -var="region=us-east-1" $ terraform apply -auto-approve $ terraform workspace new prod $ terraform workspace select dev $ # Import an existing SG $ terraform import aws_security_group.existing sg-0123456789abcdef0
- Install Terraform and configure AWS credentials
- Create VPC + public/private subnets using official module
- Add security group + outputs; run plan & apply
- Create and switch between dev and prod workspaces
- Import an existing AWS resource and reconcile
- Use LocalStack or Instruqt for free practice
Real-World Project 05.4
Single EC2 + RDS (free tier eligible)
Minimal VPC, security groups, one t2.micro + Postgres db.t3.micro. All defined in root module.
Production VPC + EKS cluster + RDS + ALB via modules
Separate environments via workspaces + tfvars. Reusable networking module. EKS managed node groups.
Multi-account landing zone with AWS Organizations, Control Tower baseline, SCPs, centralized logging, Transit Gateway
Troubleshooting 05.5
terraform force-unlock <LOCK_ID>Only run when you are sure no other apply is running.
Run terraform refresh then plan. Or use terraform apply -refresh-only.
Pin exact provider versions. Break cycles by splitting modules or using -target for emergency fixes.
30-Day Roadmap 05.6
- Resources, data sources, variables
- plan / apply / destroy cycle
- State inspection commands
- Writing and publishing modules
- Registry modules
- Version constraints
- Remote S3 + DynamoDB
- Workspaces deep dive
- Import, taint, moved
- EKS + RDS modules
- Multi-env tfvars files
- CI drift detection
Deep Dive: Advanced State & Module Patterns 05.7
State Migration Strategies
Use terraform state mv, terraform state rm, and moved blocks in 1.1+. Always backup state before migrations. For large refactors split state files using terraform workspace or separate root modules.
Module Best Practices
Keep modules small and focused. Expose only necessary inputs and outputs. Version modules strictly. Use semantic versioning. Test modules with terraform test (1.6+).
Real Incident: terraform destroy on Production RDS
What happened: A DevOps engineer ran terraform destroy in the wrong terminal window — they were targeting a dev environment but had AWS_PROFILE=production in their shell. The production RDS PostgreSQL (16TB, payments data) began deleting.
# The fateful command prod-term $ terraform destroy -target=aws_db_instance.payments_rds -auto-approve aws_db_instance.payments_rds: Destroying... [id=payments-prod-postgres-15] aws_db_instance.payments_rds: Still destroying... [30s elapsed] ← 11 minutes before DBA noticed via CloudWatch alarm # Recovery: restore from automated snapshot (RDS 7-day retention saved us) $ aws rds restore-db-instance-from-db-snapshot \ --db-instance-identifier payments-prod-restored \ --db-snapshot-identifier rds:payments-prod-postgres-15-2024-03-15-02-00
Prevention pattern:
# ALWAYS set deletion_protection on production databases
resource "aws_db_instance" "payments_rds" {
# ...
deletion_protection = true # terraform destroy will FAIL with this set
final_snapshot_identifier = "payments-rds-final-${formatdate("YYYY-MM-DD", timestamp())}"
backup_retention_period = 30 # 30-day backup retention
skip_final_snapshot = false # Always take final snapshot
}
# Also: use lifecycle prevent_destroy for critical resources
resource "aws_db_instance" "payments_rds" {
lifecycle {
prevent_destroy = true # Terraform plan will error before reaching API
}
}
# Enforce via OPA policy in CI:
# deny { input.resource_changes[_].type == "aws_db_instance"
# input.resource_changes[_].change.actions[_] == "delete" }Terragrunt: DRY Multi-Environment IaC 05.8
Terragrunt wraps Terraform to eliminate code duplication across environments. Instead of copying the same backend.tf and provider.tf into every environment directory, define them once in a root root.hcl and inherit everywhere.
infrastructure/
├── root.hcl # Shared backend + provider config
├── modules/
│ ├── eks/ # Reusable EKS module
│ ├── rds/ # Reusable RDS module
│ └── vpc/ # Reusable VPC module
└── environments/
├── dev/
│ ├── vpc/terragrunt.hcl
│ ├── eks/terragrunt.hcl
│ └── rds/terragrunt.hcl
├── staging/
│ └── ... (same structure)
└── prod/
├── vpc/terragrunt.hcl
├── eks/terragrunt.hcl
└── rds/terragrunt.hcllocals {
account_vars = read_terragrunt_config(find_in_parent_folders("account.hcl"))
env_vars = read_terragrunt_config(find_in_parent_folders("env.hcl"))
account_id = local.account_vars.locals.account_id
environment = local.env_vars.locals.environment
aws_region = local.env_vars.locals.aws_region
}
generate "provider" {
path = "provider.tf"
if_exists = "overwrite_terragrunt"
contents = <<EOF
provider "aws" {
region = "${local.aws_region}"
assume_role {
role_arn = "arn:aws:iam::${local.account_id}:role/TerraformRole"
}
default_tags {
tags = {
Environment = "${local.environment}"
ManagedBy = "terraform"
Repository = "infrastructure"
}
}
}
EOF
}
remote_state {
backend = "s3"
config = {
bucket = "tf-state-${local.account_id}-${local.aws_region}"
key = "${path_relative_to_include()}/terraform.tfstate"
region = local.aws_region
encrypt = true
dynamodb_table = "terraform-state-locks"
}
generate = {
path = "backend.tf"
if_exists = "overwrite_terragrunt"
}
}# Each environment module is just a thin config
include "root" {
path = find_in_parent_folders()
}
terraform {
source = "../../../modules//eks" # double-slash = module root
}
inputs = {
cluster_name = "prod-eks-us-east-1"
cluster_version = "1.29"
vpc_id = dependency.vpc.outputs.vpc_id
subnet_ids = dependency.vpc.outputs.private_subnet_ids
node_groups = {
general = {
instance_types = ["m5.xlarge"]
min_size = 3
max_size = 20
desired_size = 5
}
}
}
dependency "vpc" {
config_path = "../vpc"
mock_outputs = { # Allows plan without vpc deployed yet
vpc_id = "vpc-fake123"
private_subnet_ids = ["subnet-fake1", "subnet-fake2"]
}
}Atlantis + CI Drift Detection 05.9
Atlantis: Pull Request Automation
Atlantis runs terraform plan on every PR and comments the diff. Only applies after PR approval. No human needs terminal access to production infra.
# In your PR, comment one of: atlantis plan ← runs terraform plan, posts diff atlantis apply ← applies after approval atlantis unlock ← releases state lock
Drift Detection in CI
name: Terraform Drift Detection
on:
schedule:
- cron: '0 6 * * *' # Daily at 6am
jobs:
drift:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: hashicorp/setup-terraform@v3
- name: Configure AWS
uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::123456:role/TerraformReadOnly
aws-region: us-east-1
- name: Detect drift
run: |
terraform init
if ! terraform plan -detailed-exitcode -out=drift.tfplan 2>&1; then
echo "DRIFT DETECTED" | \
curl -s -X POST $SLACK_WEBHOOK \
-d '{"text":"🚨 Terraform drift detected in prod!"}'
fideletion_protection = true and prevent_destroy lifecycle to a production RDS resource