Skip to content
DevOps Architect

    Syllabus / Infrastructure / 05

    STARTUP TO ENTERPRISE
    05 / 14 • INFRASTRUCTURE AS CODE

    🏗Terraform

    Define your entire cloud estate as code. Learn the plan/apply lifecycle, state management, module composition, remote backends, and how to run production-grade multi-account environments.

    Free Tier EC2+RDS EKS + Modules Landing Zone

    Architecture & How It Works 05.1

    Terraform builds a dependency graph (DAG) from your configuration. It compares desired state (config) to actual state (state file), calculates the minimal set of changes, then executes them in order.

    graph TD Init[terraform init] --> Plan[terraform plan] Plan --> Refresh[Refresh state from real world] Refresh --> Graph[Build DAG] Graph --> Apply[terraform apply] Apply --> Create[Create/Update/Delete] Create --> Persist[Write state to remote backend] Persist --> Lock[DynamoDB state lock]

    Providers are Plugins

    Providers are downloaded during init. They translate HCL into API calls. Always pin provider versions.

    State is Sacred

    State is the only record of what Terraform manages. Never edit it manually. Use remote backends with locking.

    Core Components & Concepts 05.2

    ElementPurposeCommon Pattern
    resource / dataManage or read infrastructureaws_vpc, data "aws_ami"
    moduleReusable, versioned packagesterraform-aws-modules/vpc/aws
    variable / output / localInputs, outputs, computed valuesvar.region, output.vpc_id
    workspaceIsolate state per environmentdev / staging / prod
    backend "s3"Remote state + lockingBucket + DynamoDB table
    count / for_each / dynamicMeta-arguments for repetitionCreating N subnets

    WSL Hands-On Lab 05.3

    Complete VPC + Subnets + EKS-ready + Workspaces + Import

    install terraform
    $ wget -O- https://apt.releases.hashicorp.com/gpg | gpg --dearmor | sudo tee /usr/share/keyrings/hashicorp-archive-keyring.gpg > /dev/null
    $ echo "deb [signed-by=/usr/share/keyrings/hashicorp-archive-keyring.gpg] https://apt.releases.hashicorp.com $(lsb_release -cs) main" | sudo tee /etc/apt/sources.list.d/hashicorp.list
    $ sudo apt update && sudo apt install terraform
    main.tf (VPC + subnets)
    terraform {
      required_version = ">= 1.6.0"
      required_providers {
        aws = { source = "hashicorp/aws", version = "~> 5.0" }
      }
      backend "s3" {
        bucket         = "tf-state-mycompany-98765"
        key            = "env/${terraform.workspace}/vpc.tfstate"
        region         = "us-east-1"
        dynamodb_table = "tf-locks"
        encrypt        = true
      }
    }
    provider "aws" { region = var.region }
    
    module "vpc" {
      source  = "terraform-aws-modules/vpc/aws"
      version = "5.5.0"
      name = "devops-${terraform.workspace}"
      cidr = "10.0.0.0/16"
      azs  = ["us-east-1a","us-east-1b"]
      public_subnets  = ["10.0.1.0/24","10.0.2.0/24"]
      private_subnets = ["10.0.101.0/24","10.0.102.0/24"]
      enable_nat_gateway = true
    }
    full workflow
    $ terraform init
    $ terraform workspace new dev
    $ terraform plan -var="region=us-east-1"
    $ terraform apply -auto-approve
    $ terraform workspace new prod
    $ terraform workspace select dev
    $ # Import an existing SG
    $ terraform import aws_security_group.existing sg-0123456789abcdef0
    • Install Terraform and configure AWS credentials
    • Create VPC + public/private subnets using official module
    • Add security group + outputs; run plan & apply
    • Create and switch between dev and prod workspaces
    • Import an existing AWS resource and reconcile
    • Use LocalStack or Instruqt for free practice

    Real-World Project 05.4

    Single EC2 + RDS (free tier eligible)

    Minimal VPC, security groups, one t2.micro + Postgres db.t3.micro. All defined in root module.

    Production VPC + EKS cluster + RDS + ALB via modules

    Separate environments via workspaces + tfvars. Reusable networking module. EKS managed node groups.

    Multi-account landing zone with AWS Organizations, Control Tower baseline, SCPs, centralized logging, Transit Gateway

    Troubleshooting 05.5

    terraform force-unlock <LOCK_ID>

    Only run when you are sure no other apply is running.

    Run terraform refresh then plan. Or use terraform apply -refresh-only.

    Pin exact provider versions. Break cycles by splitting modules or using -target for emergency fixes.

    30-Day Roadmap 05.6

    WEEK 1
    Core HCL + Workflow
    • Resources, data sources, variables
    • plan / apply / destroy cycle
    • State inspection commands
    WEEK 2
    Modules & Composition
    • Writing and publishing modules
    • Registry modules
    • Version constraints
    WEEK 3
    State & Backends
    • Remote S3 + DynamoDB
    • Workspaces deep dive
    • Import, taint, moved
    WEEK 4
    AWS Production Patterns
    • EKS + RDS modules
    • Multi-env tfvars files
    • CI drift detection

    Deep Dive: Advanced State & Module Patterns 05.7

    State Migration Strategies

    Use terraform state mv, terraform state rm, and moved blocks in 1.1+. Always backup state before migrations. For large refactors split state files using terraform workspace or separate root modules.

    Module Best Practices

    Keep modules small and focused. Expose only necessary inputs and outputs. Version modules strictly. Use semantic versioning. Test modules with terraform test (1.6+).

     Real Incident: terraform destroy on Production RDS

    What happened: A DevOps engineer ran terraform destroy in the wrong terminal window — they were targeting a dev environment but had AWS_PROFILE=production in their shell. The production RDS PostgreSQL (16TB, payments data) began deleting.

    # The fateful command
    prod-term $ terraform destroy -target=aws_db_instance.payments_rds -auto-approve
    aws_db_instance.payments_rds: Destroying... [id=payments-prod-postgres-15]
    aws_db_instance.payments_rds: Still destroying... [30s elapsed]
    ← 11 minutes before DBA noticed via CloudWatch alarm
    
    # Recovery: restore from automated snapshot (RDS 7-day retention saved us)
    $ aws rds restore-db-instance-from-db-snapshot \
      --db-instance-identifier payments-prod-restored \
      --db-snapshot-identifier rds:payments-prod-postgres-15-2024-03-15-02-00

    Prevention pattern:

    main.tf — deletion_protection
    # ALWAYS set deletion_protection on production databases
    resource "aws_db_instance" "payments_rds" {
      # ...
      deletion_protection = true    # terraform destroy will FAIL with this set
      final_snapshot_identifier = "payments-rds-final-${formatdate("YYYY-MM-DD", timestamp())}"
      backup_retention_period = 30  # 30-day backup retention
      skip_final_snapshot = false   # Always take final snapshot
    }
    
    # Also: use lifecycle prevent_destroy for critical resources
    resource "aws_db_instance" "payments_rds" {
      lifecycle {
        prevent_destroy = true  # Terraform plan will error before reaching API
      }
    }
    
    # Enforce via OPA policy in CI:
    # deny { input.resource_changes[_].type == "aws_db_instance"
    #        input.resource_changes[_].change.actions[_] == "delete" }

    Terragrunt: DRY Multi-Environment IaC 05.8

    Terragrunt wraps Terraform to eliminate code duplication across environments. Instead of copying the same backend.tf and provider.tf into every environment directory, define them once in a root root.hcl and inherit everywhere.

    Repository layout
    infrastructure/
    ├── root.hcl                    # Shared backend + provider config
    ├── modules/
    │   ├── eks/                    # Reusable EKS module
    │   ├── rds/                    # Reusable RDS module
    │   └── vpc/                    # Reusable VPC module
    └── environments/
        ├── dev/
        │   ├── vpc/terragrunt.hcl
        │   ├── eks/terragrunt.hcl
        │   └── rds/terragrunt.hcl
        ├── staging/
        │   └── ... (same structure)
        └── prod/
            ├── vpc/terragrunt.hcl
            ├── eks/terragrunt.hcl
            └── rds/terragrunt.hcl
    root.hcl (defined once, inherited everywhere)
    locals {
      account_vars = read_terragrunt_config(find_in_parent_folders("account.hcl"))
      env_vars     = read_terragrunt_config(find_in_parent_folders("env.hcl"))
      account_id   = local.account_vars.locals.account_id
      environment  = local.env_vars.locals.environment
      aws_region   = local.env_vars.locals.aws_region
    }
    
    generate "provider" {
      path      = "provider.tf"
      if_exists = "overwrite_terragrunt"
      contents  = <<EOF
    provider "aws" {
      region = "${local.aws_region}"
      assume_role {
        role_arn = "arn:aws:iam::${local.account_id}:role/TerraformRole"
      }
      default_tags {
        tags = {
          Environment = "${local.environment}"
          ManagedBy   = "terraform"
          Repository  = "infrastructure"
        }
      }
    }
    EOF
    }
    
    remote_state {
      backend = "s3"
      config = {
        bucket         = "tf-state-${local.account_id}-${local.aws_region}"
        key            = "${path_relative_to_include()}/terraform.tfstate"
        region         = local.aws_region
        encrypt        = true
        dynamodb_table = "terraform-state-locks"
      }
      generate = {
        path      = "backend.tf"
        if_exists = "overwrite_terragrunt"
      }
    }
    environments/prod/eks/terragrunt.hcl
    # Each environment module is just a thin config
    include "root" {
      path = find_in_parent_folders()
    }
    
    terraform {
      source = "../../../modules//eks"  # double-slash = module root
    }
    
    inputs = {
      cluster_name    = "prod-eks-us-east-1"
      cluster_version = "1.29"
      vpc_id          = dependency.vpc.outputs.vpc_id
      subnet_ids      = dependency.vpc.outputs.private_subnet_ids
      node_groups = {
        general = {
          instance_types = ["m5.xlarge"]
          min_size       = 3
          max_size       = 20
          desired_size   = 5
        }
      }
    }
    
    dependency "vpc" {
      config_path = "../vpc"
      mock_outputs = {   # Allows plan without vpc deployed yet
        vpc_id            = "vpc-fake123"
        private_subnet_ids = ["subnet-fake1", "subnet-fake2"]
      }
    }

    Atlantis + CI Drift Detection 05.9

    Atlantis: Pull Request Automation

    Atlantis runs terraform plan on every PR and comments the diff. Only applies after PR approval. No human needs terminal access to production infra.

    # In your PR, comment one of:
    atlantis plan   ← runs terraform plan, posts diff
    atlantis apply  ← applies after approval
    atlantis unlock ← releases state lock

    Drift Detection in CI

    .github/workflows/drift-detect.yml
    name: Terraform Drift Detection
    on:
      schedule:
        - cron: '0 6 * * *'  # Daily at 6am
    jobs:
      drift:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - uses: hashicorp/setup-terraform@v3
          - name: Configure AWS
            uses: aws-actions/configure-aws-credentials@v4
            with:
              role-to-assume: arn:aws:iam::123456:role/TerraformReadOnly
              aws-region: us-east-1
          - name: Detect drift
            run: |
              terraform init
              if ! terraform plan -detailed-exitcode -out=drift.tfplan 2>&1; then
                echo "DRIFT DETECTED" | \
                curl -s -X POST $SLACK_WEBHOOK \
                  -d '{"text":"🚨 Terraform drift detected in prod!"}'
              fi
  1. Set up Terragrunt with root.hcl, create dev and prod environment configs sharing same VPC module
  2. Add deletion_protection = true and prevent_destroy lifecycle to a production RDS resource
  3. Write OPA/Conftest policy that blocks any Terraform plan that would delete a database resource
  4. Configure drift detection workflow that runs daily and sends Slack alert on changes
  5. Integrate Infracost into CI pipeline — block PRs that increase monthly cost by >$200
  6. Run: wget -O- https://apt.releases.hashicorp.com/gpg | gpg --dearmor | sudo tee /usr/share/keyrings/hashicorp-archive-keyring.gpg > /dev/null

    Extra commands from this lesson (13) are kept out of this page. Quizzes were not in the source HTML.