Skip to content
DevOps Architect

    Syllabus / Advanced / 19

    SRE RELIABILITY ENGINEERING
    19 / 19 • SRE PLAYBOOKS & TOIL ELIMINATION

    📋SRE Playbooks & Toil Automation

    The best incident is the one that never pages you. Master LitmusChaos experiments to find failure modes before production does, automated OOMKill remediation runbooks, error budget burn-rate alerting, PagerDuty escalation policies, chaos game days, and blameless postmortem culture that actually improves reliability over time.

    SLO + Alertmanager LitmusChaos + Auto-Remediation Chaos GameDay + Error Budget Policy

     Real Postmortem: 47-Minute Payments Outage

    Date: 2024-03-15 14:22 UTC  |  Impact: 100% payment failures, ~$280K revenue loss  |  Duration: 47 minutes

    Time (UTC)Event
    14:22Deploy payments-api v2.3.1 — Redis connection pool size reduced from 50→5 (misconfiguration)
    14:24Redis connection exhaustion under normal load. Requests queue. P99 latency climbs to 15s.
    14:26SLO error budget burn alert fires (1h window). PagerDuty pages on-call SRE.
    14:31On-call SRE joins. Reviews Grafana. Sees Redis connection errors in logs. Suspects Redis cluster issue.
    14:38Redis cluster scaled up (wrong fix). No improvement.
    14:44Second SRE joins. Reviews deployment diff. Finds pool size change.
    14:48Hotfix deployed. Connection pool restored to 50. Recovery begins.
    15:09Full recovery confirmed. P99 back to 120ms.

    Root Cause: Redis connection pool config was in a separate values file from the main chart. PR review missed it. No automated smoke test for connection pool behavior.

    Action Items: (1) Add integration test verifying Redis connectivity under load. (2) Add config diff check in CI. (3) Implement canary deployment for payments-api (10% traffic first). (4) Build automated rollback on SLO breach.

    SLO, SLI & Error Budget Enforcement 19.1

    slo-payments-api.yaml (OpenSLO format)
    apiVersion: openslo/v1
    kind: SLO
    metadata:
      name: payments-api-availability
      namespace: production
    spec:
      service: payments-api
      description: "Payment API must be available 99.9% of the time"
      timeWindow:
        - duration: 30d
          isRolling: true
      indicator:
        metadata:
          name: payment-success-rate
        spec:
          ratioMetric:
            good:
              metricSource:
                type: Prometheus
                spec:
                  query: |
                    sum(rate(http_request_total{
                      service="payments-api",
                      status_code!~"5.."
                    }[5m]))
            total:
              metricSource:
                type: Prometheus
                spec:
                  query: |
                    sum(rate(http_request_total{
                      service="payments-api"
                    }[5m]))
      objectives:
        - target: 0.999         # 99.9% = 43.2 min/month error budget
          timeWindow: 30d
    
    ---
    # Prometheus recording rules for SLO tracking
    # In prometheus-rules.yaml:
    # - record: slo:payments_availability:ratio_rate5m
    #   expr: |
    #     sum(rate(http_requests_total{job="payments-api",status!~"5.."}[5m]))
    #     /
    #     sum(rate(http_requests_total{job="payments-api"}[5m]))
    burn-rate-alerting-rules.yaml
    # Multi-window burn-rate alerting (Google SRE Workbook pattern)
    # Catches fast burns early AND slow burns over time
    groups:
      - name: payments-api-slo-alerts
        rules:
          # 1h window: burning 14x faster than budget allows
          # → pages immediately (fast burn = big incident)
          - alert: PaymentsSLOBurnRateFastHigh
            expr: |
              (
                1 - slo:payments_availability:ratio_rate1h
              ) / (1 - 0.999) > 14
            for: 2m
            labels:
              severity: critical
              team: payments
            annotations:
              summary: "Payments SLO burning 14x fast (1h window)"
              description: >
                Error budget consuming at 14x normal rate over 1h.
                At this rate, 30-day budget exhausted in {{ printf "%.1f" (30 / 14) }} days.
                Current error rate: {{ $value | humanizePercentage }}
    
          # 6h window: burning 6x — page but less urgently
          - alert: PaymentsSLOBurnRateMedium
            expr: |
              (
                1 - slo:payments_availability:ratio_rate6h
              ) / (1 - 0.999) > 6
            for: 15m
            labels:
              severity: warning
              team: payments
    
          # Both windows must fire to reduce false positives
          - alert: PaymentsSLOBurnRateFastConfirmed
            expr: |
              (
                (1 - slo:payments_availability:ratio_rate1h) / (1-0.999) > 14
                and
                (1 - slo:payments_availability:ratio_rate5m) / (1-0.999) > 14
              )
            for: 0m
            labels:
              severity: critical
              pagerduty: true
    error-budget-policy.md
    # Error Budget Policy — Payments Team
    # Adopted: 2024-01-01 | Owner: Platform SRE Team
    
    ## Error Budget Thresholds and Actions
    
    ### Green (>50% budget remaining)
    - Normal deployment velocity permitted
    - Feature releases: up to 3 per week
    - Risky changes: allowed with testing
    
    ### Yellow (25-50% remaining)
    - Review all upcoming risky changes with SRE
    - No infrastructure maintenance windows
    - Increase release gate testing coverage
    
    ### Red (10-25% remaining)
    - Feature deployments FROZEN
    - Only bugfixes and reliability improvements
    - Mandatory SRE review of every deployment
    - Weekly error budget review meeting
    
    ### Critical (<10% remaining)
    - All deployments BLOCKED (requires VP approval)
    - War room activated to identify root cause
    - Incident retrospective within 48h
    - Error budget thaw only after reliability improvements shipped
    
    ## Automated Enforcement
    The following CI gate checks error budget via API:
      curl https://sloth.monitoring.acme.com/api/v1/slo/payments-api \
        | jq '.errorBudgetPercent' | awk '{if($1<10) exit 1}'

    Auto-Remediation Runbooks 19.2

    oomkill-remediation-runbook.sh
    #!/bin/bash
    # Auto-remediation: OOMKill detected → increase memory limit by 25%
    # Triggered by Alertmanager webhook → Lambda → this script runs in K8s Job
    set -euo pipefail
    
    NAMESPACE="${1:-production}"
    DEPLOYMENT="${2:-}"
    ALERT_THRESHOLD=3  # Only auto-remediate if OOMKilled > 3 times in 1h
    
    if [[ -z "$DEPLOYMENT" ]]; then
      # Find deployments with recent OOMKills
      DEPLOYMENT=$(kubectl get pods -n "$NAMESPACE" \
        --field-selector=status.phase=Running \
        -o json | jq -r '
          .items[] |
          select(.status.containerStatuses[]?.lastState.terminated.reason == "OOMKilled") |
          .metadata.labels."app"
        ' | sort -u | head -1)
    fi
    
    echo "=== OOMKill Auto-Remediation: $DEPLOYMENT in $NAMESPACE ==="
    
    # Count OOMKills in last hour
    OOMKILL_COUNT=$(kubectl get events -n "$NAMESPACE" \
      --field-selector=reason=OOMKilling \
      --sort-by='.lastTimestamp' \
      -o json | \
      jq --arg app "$DEPLOYMENT" '
        .items[] |
        select(.involvedObject.name | test($app)) |
        select(.lastTimestamp > (now - 3600 | todate))
      ' | jq -s 'length')
    
    echo "OOMKill count in last 1h: $OOMKILL_COUNT"
    
    if [[ $OOMKILL_COUNT -lt $ALERT_THRESHOLD ]]; then
      echo "Below threshold ($ALERT_THRESHOLD). No auto-remediation."
      exit 0
    fi
    
    # Get current memory limit
    CURRENT_LIMIT=$(kubectl get deployment "$DEPLOYMENT" -n "$NAMESPACE" \
      -o jsonpath='{.spec.template.spec.containers[0].resources.limits.memory}')
    
    echo "Current memory limit: $CURRENT_LIMIT"
    
    # Parse and increase by 25%
    CURRENT_MI=$(echo "$CURRENT_LIMIT" | sed 's/Mi//')
    NEW_MI=$(echo "scale=0; $CURRENT_MI * 1.25 / 1" | bc)
    NEW_LIMIT="${NEW_MI}Mi"
    
    echo "Increasing to: $NEW_LIMIT"
    
    # Apply patch
    kubectl patch deployment "$DEPLOYMENT" -n "$NAMESPACE" \
      --type='json' \
      -p="[{
        \"op\": \"replace\",
        \"path\": \"/spec/template/spec/containers/0/resources/limits/memory\",
        \"value\": \"${NEW_LIMIT}\"
      }]"
    
    # Post remediation event to Slack
    curl -s -X POST "$SLACK_WEBHOOK_URL" -H 'Content-type: application/json' \
      -d "{
        \"text\": \"🤖 *Auto-Remediation* | $DEPLOYMENT in $NAMESPACE\",
        \"attachments\": [{
          \"color\": \"#f59e0b\",
          \"fields\": [
            {\"title\": \"Action\", \"value\": \"Memory limit increased: $CURRENT_LIMIT → $NEW_LIMIT\"},
            {\"title\": \"OOMKill Count\", \"value\": \"$OOMKILL_COUNT in last 1h\"},
            {\"title\": \"Next Step\", \"value\": \"Review with VPA recommendation and open ticket to right-size properly\"}
          ]
        }]
      }"
    
    echo "Auto-remediation complete. Deployment rolling update started."
    disk-full-remediation.sh
    #!/bin/bash
    # Auto-remediation: node disk pressure → clean up unused images + logs
    # Triggered when node hits 85% disk usage
    
    NODE="${1:-}"
    THRESHOLD=85
    
    echo "=== Disk Pressure Auto-Remediation: $NODE ==="
    
    # Check which node is at risk
    if [[ -z "$NODE" ]]; then
      NODE=$(kubectl get nodes -o json | jq -r '
        .items[] |
        select(.status.conditions[] |
          select(.type == "DiskPressure" and .status == "True")
        ) | .metadata.name' | head -1)
    fi
    
    if [[ -z "$NODE" ]]; then
      echo "No node in DiskPressure state. Checking by usage..."
      NODE=$(kubectl top nodes --no-headers | \
        awk '{gsub(/%/,"",$5); if($5>'$THRESHOLD') print $1}' | head -1)
    fi
    
    echo "Target node: $NODE"
    
    # Run cleanup job on that specific node
    cat <100MB)
                  find /var/log -name "*.log" -size +100M -exec gzip {} \;
                  # Remove old journal logs older than 3 days
                  journalctl --vacuum-time=3d
                  echo "Cleanup complete"
              securityContext:
                privileged: true
              volumeMounts:
                - name: docker-sock
                  mountPath: /var/run/docker.sock
                - name: var-log
                  mountPath: /var/log
          volumes:
            - name: docker-sock
              hostPath:
                path: /var/run/docker.sock
            - name: var-log
              hostPath:
                path: /var/log
          restartPolicy: Never
    EOF
    
    echo "Cleanup job dispatched to node: $NODE"
    alertmanager-webhook-receiver.yaml
    # Alertmanager routes alerts to webhook receivers
    # which trigger automated runbooks
    route:
      group_by: ['alertname', 'deployment', 'namespace']
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 4h
      receiver: default-slack
    
      routes:
        # OOMKill alerts → auto-remediation first, then Slack
        - matchers:
            - alertname =~ "OOMKill.*"
          receiver: oomkill-remediation
          continue: true   # Also send to default-slack after
    
        # Disk pressure → auto-cleanup
        - matchers:
            - alertname = "NodeDiskPressure"
          receiver: disk-remediation
    
        # SLO burn → PagerDuty
        - matchers:
            - pagerduty = "true"
          receiver: pagerduty
    
    receivers:
      - name: oomkill-remediation
        webhook_configs:
          - url: 'https://remediation-api.acme.internal/oomkill'
            send_resolved: false
            http_config:
              bearer_token_file: /var/run/secrets/remediation-token
    
      - name: disk-remediation
        webhook_configs:
          - url: 'https://remediation-api.acme.internal/disk-pressure'
    
      - name: pagerduty
        pagerduty_configs:
          - routing_key: '{{ secrets.PAGERDUTY_KEY }}'
            severity: '{{ .GroupLabels.severity }}'
            description: '{{ .CommonAnnotations.summary }}'
    
      - name: default-slack
        slack_configs:
          - api_url: 'https://hooks.slack.com/YOUR_WEBHOOK'
            channel: '#sre-alerts'
            title: '{{ .Status | toUpper }} {{ .GroupLabels.alertname }}'
            text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'

    LitmusChaos: Production Readiness Testing 19.3

    The only way to know your system handles failures is to deliberately inject them before production does. LitmusChaos is a CNCF project that runs chaos experiments as Kubernetes CRDs — pod kills, network delays, CPU hogs, and disk pressure.

    chaos-pod-delete.yaml
    apiVersion: litmuschaos.io/v1alpha1
    kind: ChaosEngine
    metadata:
      name: payments-pod-delete
      namespace: production
    spec:
      appinfo:
        appns: production
        applabel: "app=payments-api"
        appkind: deployment
      # Chaos will stop if SLO is violated
      jobCleanUpPolicy: delete
      monitoring: true
      annotationCheck: 'false'
      engineState: 'active'
      experiments:
        - name: pod-delete
          spec:
            components:
              env:
                # Kill 1 pod every 30s for 5 minutes
                - name: TOTAL_CHAOS_DURATION
                  value: "300"
                - name: CHAOS_INTERVAL
                  value: "30"
                # Kill 30% of pods (not all at once)
                - name: FORCE
                  value: "false"
                - name: PODS_AFFECTED_PERC
                  value: "30"
            probe:
              # Abort chaos if SLO drops below 99%
              - name: payments-availability-probe
                type: promProbe
                promProbe/inputs:
                  endpoint: http://prometheus.monitoring.svc.cluster.local:9090
                  query: |
                    sum(rate(http_request_total{service="payments-api",status!~"5.."}[1m]))
                    /
                    sum(rate(http_request_total{service="payments-api"}[1m]))
                mode: Continuous
                runProperties:
                  probeTimeout: 5
                  interval: 5
                  retry: 1
                comparator:
                  type: float
                  criteria: ">="
                  value: "0.99"   # Abort if availability drops below 99%
    chaos-network-latency.yaml
    apiVersion: litmuschaos.io/v1alpha1
    kind: ChaosEngine
    metadata:
      name: inventory-network-latency
      namespace: production
    spec:
      appinfo:
        appns: production
        applabel: "app=inventory-service"
        appkind: deployment
      experiments:
        - name: pod-network-latency
          spec:
            components:
              env:
                # Add 500ms latency to simulate degraded upstream
                - name: NETWORK_LATENCY
                  value: "500"    # milliseconds
                - name: JITTER
                  value: "50"     # ±50ms jitter
                # Only affect traffic to specific destination
                - name: DESTINATION_IPS
                  value: "10.0.0.0/8"
                - name: TOTAL_CHAOS_DURATION
                  value: "180"
                - name: PODS_AFFECTED_PERC
                  value: "50"
            probe:
              # Expect circuit breaker to open → verify fallback works
              - name: fallback-working
                type: httpProbe
                httpProbe/inputs:
                  url: https://checkout.acme.com/health
                  insecureSkipVerify: false
                  responseTimeout: 2000
                  method:
                    get:
                      criteria: ==
                      responseCode: "200"
                mode: Continuous
                runProperties:
                  probeTimeout: 2
                  interval: 5
                  retry: 1
    chaos-cpu-stress.yaml
    apiVersion: litmuschaos.io/v1alpha1
    kind: ChaosEngine
    metadata:
      name: payments-cpu-stress
      namespace: production
    spec:
      appinfo:
        appns: production
        applabel: "app=payments-api"
        appkind: deployment
      experiments:
        - name: pod-cpu-hog
          spec:
            components:
              env:
                # Consume 80% of CPU limit for 2 minutes
                - name: CPU_CORES
                  value: "0"        # 0 = use % of request
                - name: CPU_LOAD
                  value: "80"       # 80% CPU load
                - name: TOTAL_CHAOS_DURATION
                  value: "120"
                - name: PODS_AFFECTED_PERC
                  value: "50"
            probe:
              # HPA should scale out — verify it does
              - name: hpa-scale-check
                type: k8sProbe
                k8sProbe/inputs:
                  group: autoscaling
                  version: v2
                  resource: horizontalpodautoscalers
                  namespace: production
                  resourceNames: "payments-api"
                  operation: present
                mode: OnChaos
                runProperties:
                  probeTimeout: 5
                  interval: 10
                  retry: 3
    chaos-gameday-runbook.md
    # Chaos Game Day Runbook — Payments Platform
    # Frequency: Monthly | Duration: 3h | Owner: Platform SRE
    
    ## Pre-Game Day Checklist
    - [ ] Notify all teams in #engineering Slack channel (3 days before)
    - [ ] Confirm on-call rotations are staffed
    - [ ] Verify Grafana dashboards are accessible
    - [ ] Document current SLO baseline and error budget
    - [ ] Confirm rollback procedures are documented
    - [ ] Get product owner approval for traffic reduction
    
    ## Game Day Agenda (All Times UTC)
    
    ### 14:00 — Kick-off
    - Confirm participant attendance (Zoom call)
    - Review today's scenarios and abort conditions
    - Verify monitoring is working (send test alert)
    
    ### 14:15 — Scenario 1: Pod Delete (Low Risk)
    - Run: pod-delete chaos on payments-api (30% pods, 5 min)
    - Observe: HPA response time, alert firing, PDB protection
    - Expected: no SLO impact, HPA scales within 2 min
    - Abort if: error rate > 1% for > 30 seconds
    
    ### 14:45 — Scenario 2: Network Latency (Medium Risk)
    - Run: 500ms latency on inventory-service calls
    - Observe: circuit breaker trip, fallback behavior
    - Expected: checkout continues with degraded inventory check
    - Abort if: checkout error rate > 5%
    
    ### 15:15 — Scenario 3: AZ Failure Simulation (High Risk)
    - Cordon all nodes in us-east-1a
    - Drain workloads to remaining AZs
    - Observe: anti-affinity rules, cross-AZ traffic
    - Abort if: total service downtime > 2 minutes
    
    ### 15:45 — Review & Retrospective
    - Document findings per scenario
    - Identify gaps in alerting
    - Create reliability improvement tickets
    - Update runbooks with learnings
    
    ## Abort Criteria (Any of these = stop immediately)
    - SLO error budget consumes > 20% of monthly budget
    - Customer-visible error rate > 2% for > 60 seconds
    - On-call team loses control of incident

    Measuring & Reducing Toil 19.4

    What Counts as Toil

    • Manual, repetitive operational work
    • Running the same script every deploy
    • Manually scaling resources reactively
    • Copy-pasting secrets between systems
    • Answering the same ticket type >2x/week
    • Manually rotating SSL certificates
    • Clicking through consoles to check status

    Toil Elimination Targets

    • Google SRE target: <50% time on toil
    • Track toil weekly in team retrospectives
    • Automate any task done >2x identical
    • Self-service > ticket for all dev requests
    • Auto-cert renewal via cert-manager
    • Auto-scaling replacing manual scaling
    • Runbook automation via SSM/Lambda
    # Track toil via PagerDuty API — how many alerts are "noisy" vs actionable
    $ curl -H "Authorization: Token token=$PD_TOKEN" \
      "https://api.pagerduty.com/incidents?time_zone=UTC&since=2024-01-01&until=2024-02-01&limit=100" \
      | jq '.incidents[] | {
          service: .service.summary,
          resolved_by: .resolved_by_user.summary,
          duration_min: ((.last_status_change_at | fromdateiso8601) -
                         (.created_at | fromdateiso8601)) / 60 | round,
          acknowledged: (.acknowledgements | length)
        }' | jq -s 'group_by(.service) | .[] | {
          service: .[0].service,
          total_pages: length,
          avg_duration_min: ([.[].duration_min] | add / length | round)
        }'
    {
      "service": "payments-api",
      "total_pages": 47,
      "avg_duration_min": 12
    }
    ← 47 pages in 1 month = investigate noisy alert or recurring bug

    Hands-On Labs 19.5

    SRE Reliability Labs

    • Define SLO for a sample service using OpenSLO spec, implement recording rules in Prometheus
    • Configure multi-window burn-rate alerting (1h 14x + 6h 6x) in Alertmanager
    • Install LitmusChaos on kind, run pod-delete experiment on a sample deployment
    • Write OOMKill auto-remediation script and test it by manually creating OOM conditions
    • Configure Alertmanager webhook route that calls a remediation endpoint for OOMKill alerts
    • Write a complete blameless postmortem for the Redis connection pool incident (use template)
    • Plan and document a Chaos Game Day agenda for a 3-service system with clear abort criteria
    • Track your own toil for 1 week — log every manual task and identify top 3 automation opportunities

    Troubleshooting 19.6

    Use the two-window approach: both the fast (1h) and slow (6h) windows must fire simultaneously. Add a for: 2m to avoid transient spikes. Review your SLI query — if it includes health checks or internal traffic, exclude them. Use job_name != "blackbox-health" filters.

    Check: (1) The target pod has the correct label matching applabel in ChaosEngine. (2) ChaosServiceAccount has RBAC permissions to delete pods. (3) annotationCheck: 'false' — if true, pods need litmuschaos.io/chaos: "true" annotation. Check ChaosEngine status: kubectl describe chaosengine <name> -n <ns>.

    Add a cooldown period — track last remediation timestamp in a ConfigMap or DynamoDB table. Only trigger remediation if last run was >30 minutes ago. Cap auto-escalation: allow max 2 automatic increases, then create a Jira ticket for human review. Always Slack-notify with a clear "MANUAL ACTION REQUIRED IF PERSISTS" message.

    Run: curl -H "Authorization: Token token=$PD_TOKEN" \

    Extra commands from this lesson (12) are kept out of this page. Quizzes were not in the source HTML.