📋SRE Playbooks & Toil Automation
The best incident is the one that never pages you. Master LitmusChaos experiments to find failure modes before production does, automated OOMKill remediation runbooks, error budget burn-rate alerting, PagerDuty escalation policies, chaos game days, and blameless postmortem culture that actually improves reliability over time.
Real Postmortem: 47-Minute Payments Outage
Date: 2024-03-15 14:22 UTC | Impact: 100% payment failures, ~$280K revenue loss | Duration: 47 minutes
| Time (UTC) | Event |
|---|---|
| 14:22 | Deploy payments-api v2.3.1 — Redis connection pool size reduced from 50→5 (misconfiguration) |
| 14:24 | Redis connection exhaustion under normal load. Requests queue. P99 latency climbs to 15s. |
| 14:26 | SLO error budget burn alert fires (1h window). PagerDuty pages on-call SRE. |
| 14:31 | On-call SRE joins. Reviews Grafana. Sees Redis connection errors in logs. Suspects Redis cluster issue. |
| 14:38 | Redis cluster scaled up (wrong fix). No improvement. |
| 14:44 | Second SRE joins. Reviews deployment diff. Finds pool size change. |
| 14:48 | Hotfix deployed. Connection pool restored to 50. Recovery begins. |
| 15:09 | Full recovery confirmed. P99 back to 120ms. |
Root Cause: Redis connection pool config was in a separate values file from the main chart. PR review missed it. No automated smoke test for connection pool behavior.
Action Items: (1) Add integration test verifying Redis connectivity under load. (2) Add config diff check in CI. (3) Implement canary deployment for payments-api (10% traffic first). (4) Build automated rollback on SLO breach.
SLO, SLI & Error Budget Enforcement 19.1
apiVersion: openslo/v1
kind: SLO
metadata:
name: payments-api-availability
namespace: production
spec:
service: payments-api
description: "Payment API must be available 99.9% of the time"
timeWindow:
- duration: 30d
isRolling: true
indicator:
metadata:
name: payment-success-rate
spec:
ratioMetric:
good:
metricSource:
type: Prometheus
spec:
query: |
sum(rate(http_request_total{
service="payments-api",
status_code!~"5.."
}[5m]))
total:
metricSource:
type: Prometheus
spec:
query: |
sum(rate(http_request_total{
service="payments-api"
}[5m]))
objectives:
- target: 0.999 # 99.9% = 43.2 min/month error budget
timeWindow: 30d
---
# Prometheus recording rules for SLO tracking
# In prometheus-rules.yaml:
# - record: slo:payments_availability:ratio_rate5m
# expr: |
# sum(rate(http_requests_total{job="payments-api",status!~"5.."}[5m]))
# /
# sum(rate(http_requests_total{job="payments-api"}[5m]))# Multi-window burn-rate alerting (Google SRE Workbook pattern)
# Catches fast burns early AND slow burns over time
groups:
- name: payments-api-slo-alerts
rules:
# 1h window: burning 14x faster than budget allows
# → pages immediately (fast burn = big incident)
- alert: PaymentsSLOBurnRateFastHigh
expr: |
(
1 - slo:payments_availability:ratio_rate1h
) / (1 - 0.999) > 14
for: 2m
labels:
severity: critical
team: payments
annotations:
summary: "Payments SLO burning 14x fast (1h window)"
description: >
Error budget consuming at 14x normal rate over 1h.
At this rate, 30-day budget exhausted in {{ printf "%.1f" (30 / 14) }} days.
Current error rate: {{ $value | humanizePercentage }}
# 6h window: burning 6x — page but less urgently
- alert: PaymentsSLOBurnRateMedium
expr: |
(
1 - slo:payments_availability:ratio_rate6h
) / (1 - 0.999) > 6
for: 15m
labels:
severity: warning
team: payments
# Both windows must fire to reduce false positives
- alert: PaymentsSLOBurnRateFastConfirmed
expr: |
(
(1 - slo:payments_availability:ratio_rate1h) / (1-0.999) > 14
and
(1 - slo:payments_availability:ratio_rate5m) / (1-0.999) > 14
)
for: 0m
labels:
severity: critical
pagerduty: true# Error Budget Policy — Payments Team
# Adopted: 2024-01-01 | Owner: Platform SRE Team
## Error Budget Thresholds and Actions
### Green (>50% budget remaining)
- Normal deployment velocity permitted
- Feature releases: up to 3 per week
- Risky changes: allowed with testing
### Yellow (25-50% remaining)
- Review all upcoming risky changes with SRE
- No infrastructure maintenance windows
- Increase release gate testing coverage
### Red (10-25% remaining)
- Feature deployments FROZEN
- Only bugfixes and reliability improvements
- Mandatory SRE review of every deployment
- Weekly error budget review meeting
### Critical (<10% remaining)
- All deployments BLOCKED (requires VP approval)
- War room activated to identify root cause
- Incident retrospective within 48h
- Error budget thaw only after reliability improvements shipped
## Automated Enforcement
The following CI gate checks error budget via API:
curl https://sloth.monitoring.acme.com/api/v1/slo/payments-api \
| jq '.errorBudgetPercent' | awk '{if($1<10) exit 1}'Auto-Remediation Runbooks 19.2
#!/bin/bash
# Auto-remediation: OOMKill detected → increase memory limit by 25%
# Triggered by Alertmanager webhook → Lambda → this script runs in K8s Job
set -euo pipefail
NAMESPACE="${1:-production}"
DEPLOYMENT="${2:-}"
ALERT_THRESHOLD=3 # Only auto-remediate if OOMKilled > 3 times in 1h
if [[ -z "$DEPLOYMENT" ]]; then
# Find deployments with recent OOMKills
DEPLOYMENT=$(kubectl get pods -n "$NAMESPACE" \
--field-selector=status.phase=Running \
-o json | jq -r '
.items[] |
select(.status.containerStatuses[]?.lastState.terminated.reason == "OOMKilled") |
.metadata.labels."app"
' | sort -u | head -1)
fi
echo "=== OOMKill Auto-Remediation: $DEPLOYMENT in $NAMESPACE ==="
# Count OOMKills in last hour
OOMKILL_COUNT=$(kubectl get events -n "$NAMESPACE" \
--field-selector=reason=OOMKilling \
--sort-by='.lastTimestamp' \
-o json | \
jq --arg app "$DEPLOYMENT" '
.items[] |
select(.involvedObject.name | test($app)) |
select(.lastTimestamp > (now - 3600 | todate))
' | jq -s 'length')
echo "OOMKill count in last 1h: $OOMKILL_COUNT"
if [[ $OOMKILL_COUNT -lt $ALERT_THRESHOLD ]]; then
echo "Below threshold ($ALERT_THRESHOLD). No auto-remediation."
exit 0
fi
# Get current memory limit
CURRENT_LIMIT=$(kubectl get deployment "$DEPLOYMENT" -n "$NAMESPACE" \
-o jsonpath='{.spec.template.spec.containers[0].resources.limits.memory}')
echo "Current memory limit: $CURRENT_LIMIT"
# Parse and increase by 25%
CURRENT_MI=$(echo "$CURRENT_LIMIT" | sed 's/Mi//')
NEW_MI=$(echo "scale=0; $CURRENT_MI * 1.25 / 1" | bc)
NEW_LIMIT="${NEW_MI}Mi"
echo "Increasing to: $NEW_LIMIT"
# Apply patch
kubectl patch deployment "$DEPLOYMENT" -n "$NAMESPACE" \
--type='json' \
-p="[{
\"op\": \"replace\",
\"path\": \"/spec/template/spec/containers/0/resources/limits/memory\",
\"value\": \"${NEW_LIMIT}\"
}]"
# Post remediation event to Slack
curl -s -X POST "$SLACK_WEBHOOK_URL" -H 'Content-type: application/json' \
-d "{
\"text\": \"🤖 *Auto-Remediation* | $DEPLOYMENT in $NAMESPACE\",
\"attachments\": [{
\"color\": \"#f59e0b\",
\"fields\": [
{\"title\": \"Action\", \"value\": \"Memory limit increased: $CURRENT_LIMIT → $NEW_LIMIT\"},
{\"title\": \"OOMKill Count\", \"value\": \"$OOMKILL_COUNT in last 1h\"},
{\"title\": \"Next Step\", \"value\": \"Review with VPA recommendation and open ticket to right-size properly\"}
]
}]
}"
echo "Auto-remediation complete. Deployment rolling update started."#!/bin/bash
# Auto-remediation: node disk pressure → clean up unused images + logs
# Triggered when node hits 85% disk usage
NODE="${1:-}"
THRESHOLD=85
echo "=== Disk Pressure Auto-Remediation: $NODE ==="
# Check which node is at risk
if [[ -z "$NODE" ]]; then
NODE=$(kubectl get nodes -o json | jq -r '
.items[] |
select(.status.conditions[] |
select(.type == "DiskPressure" and .status == "True")
) | .metadata.name' | head -1)
fi
if [[ -z "$NODE" ]]; then
echo "No node in DiskPressure state. Checking by usage..."
NODE=$(kubectl top nodes --no-headers | \
awk '{gsub(/%/,"",$5); if($5>'$THRESHOLD') print $1}' | head -1)
fi
echo "Target node: $NODE"
# Run cleanup job on that specific node
cat <100MB)
find /var/log -name "*.log" -size +100M -exec gzip {} \;
# Remove old journal logs older than 3 days
journalctl --vacuum-time=3d
echo "Cleanup complete"
securityContext:
privileged: true
volumeMounts:
- name: docker-sock
mountPath: /var/run/docker.sock
- name: var-log
mountPath: /var/log
volumes:
- name: docker-sock
hostPath:
path: /var/run/docker.sock
- name: var-log
hostPath:
path: /var/log
restartPolicy: Never
EOF
echo "Cleanup job dispatched to node: $NODE" # Alertmanager routes alerts to webhook receivers
# which trigger automated runbooks
route:
group_by: ['alertname', 'deployment', 'namespace']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receiver: default-slack
routes:
# OOMKill alerts → auto-remediation first, then Slack
- matchers:
- alertname =~ "OOMKill.*"
receiver: oomkill-remediation
continue: true # Also send to default-slack after
# Disk pressure → auto-cleanup
- matchers:
- alertname = "NodeDiskPressure"
receiver: disk-remediation
# SLO burn → PagerDuty
- matchers:
- pagerduty = "true"
receiver: pagerduty
receivers:
- name: oomkill-remediation
webhook_configs:
- url: 'https://remediation-api.acme.internal/oomkill'
send_resolved: false
http_config:
bearer_token_file: /var/run/secrets/remediation-token
- name: disk-remediation
webhook_configs:
- url: 'https://remediation-api.acme.internal/disk-pressure'
- name: pagerduty
pagerduty_configs:
- routing_key: '{{ secrets.PAGERDUTY_KEY }}'
severity: '{{ .GroupLabels.severity }}'
description: '{{ .CommonAnnotations.summary }}'
- name: default-slack
slack_configs:
- api_url: 'https://hooks.slack.com/YOUR_WEBHOOK'
channel: '#sre-alerts'
title: '{{ .Status | toUpper }} {{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'LitmusChaos: Production Readiness Testing 19.3
The only way to know your system handles failures is to deliberately inject them before production does. LitmusChaos is a CNCF project that runs chaos experiments as Kubernetes CRDs — pod kills, network delays, CPU hogs, and disk pressure.
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: payments-pod-delete
namespace: production
spec:
appinfo:
appns: production
applabel: "app=payments-api"
appkind: deployment
# Chaos will stop if SLO is violated
jobCleanUpPolicy: delete
monitoring: true
annotationCheck: 'false'
engineState: 'active'
experiments:
- name: pod-delete
spec:
components:
env:
# Kill 1 pod every 30s for 5 minutes
- name: TOTAL_CHAOS_DURATION
value: "300"
- name: CHAOS_INTERVAL
value: "30"
# Kill 30% of pods (not all at once)
- name: FORCE
value: "false"
- name: PODS_AFFECTED_PERC
value: "30"
probe:
# Abort chaos if SLO drops below 99%
- name: payments-availability-probe
type: promProbe
promProbe/inputs:
endpoint: http://prometheus.monitoring.svc.cluster.local:9090
query: |
sum(rate(http_request_total{service="payments-api",status!~"5.."}[1m]))
/
sum(rate(http_request_total{service="payments-api"}[1m]))
mode: Continuous
runProperties:
probeTimeout: 5
interval: 5
retry: 1
comparator:
type: float
criteria: ">="
value: "0.99" # Abort if availability drops below 99%apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: inventory-network-latency
namespace: production
spec:
appinfo:
appns: production
applabel: "app=inventory-service"
appkind: deployment
experiments:
- name: pod-network-latency
spec:
components:
env:
# Add 500ms latency to simulate degraded upstream
- name: NETWORK_LATENCY
value: "500" # milliseconds
- name: JITTER
value: "50" # ±50ms jitter
# Only affect traffic to specific destination
- name: DESTINATION_IPS
value: "10.0.0.0/8"
- name: TOTAL_CHAOS_DURATION
value: "180"
- name: PODS_AFFECTED_PERC
value: "50"
probe:
# Expect circuit breaker to open → verify fallback works
- name: fallback-working
type: httpProbe
httpProbe/inputs:
url: https://checkout.acme.com/health
insecureSkipVerify: false
responseTimeout: 2000
method:
get:
criteria: ==
responseCode: "200"
mode: Continuous
runProperties:
probeTimeout: 2
interval: 5
retry: 1apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: payments-cpu-stress
namespace: production
spec:
appinfo:
appns: production
applabel: "app=payments-api"
appkind: deployment
experiments:
- name: pod-cpu-hog
spec:
components:
env:
# Consume 80% of CPU limit for 2 minutes
- name: CPU_CORES
value: "0" # 0 = use % of request
- name: CPU_LOAD
value: "80" # 80% CPU load
- name: TOTAL_CHAOS_DURATION
value: "120"
- name: PODS_AFFECTED_PERC
value: "50"
probe:
# HPA should scale out — verify it does
- name: hpa-scale-check
type: k8sProbe
k8sProbe/inputs:
group: autoscaling
version: v2
resource: horizontalpodautoscalers
namespace: production
resourceNames: "payments-api"
operation: present
mode: OnChaos
runProperties:
probeTimeout: 5
interval: 10
retry: 3# Chaos Game Day Runbook — Payments Platform
# Frequency: Monthly | Duration: 3h | Owner: Platform SRE
## Pre-Game Day Checklist
- [ ] Notify all teams in #engineering Slack channel (3 days before)
- [ ] Confirm on-call rotations are staffed
- [ ] Verify Grafana dashboards are accessible
- [ ] Document current SLO baseline and error budget
- [ ] Confirm rollback procedures are documented
- [ ] Get product owner approval for traffic reduction
## Game Day Agenda (All Times UTC)
### 14:00 — Kick-off
- Confirm participant attendance (Zoom call)
- Review today's scenarios and abort conditions
- Verify monitoring is working (send test alert)
### 14:15 — Scenario 1: Pod Delete (Low Risk)
- Run: pod-delete chaos on payments-api (30% pods, 5 min)
- Observe: HPA response time, alert firing, PDB protection
- Expected: no SLO impact, HPA scales within 2 min
- Abort if: error rate > 1% for > 30 seconds
### 14:45 — Scenario 2: Network Latency (Medium Risk)
- Run: 500ms latency on inventory-service calls
- Observe: circuit breaker trip, fallback behavior
- Expected: checkout continues with degraded inventory check
- Abort if: checkout error rate > 5%
### 15:15 — Scenario 3: AZ Failure Simulation (High Risk)
- Cordon all nodes in us-east-1a
- Drain workloads to remaining AZs
- Observe: anti-affinity rules, cross-AZ traffic
- Abort if: total service downtime > 2 minutes
### 15:45 — Review & Retrospective
- Document findings per scenario
- Identify gaps in alerting
- Create reliability improvement tickets
- Update runbooks with learnings
## Abort Criteria (Any of these = stop immediately)
- SLO error budget consumes > 20% of monthly budget
- Customer-visible error rate > 2% for > 60 seconds
- On-call team loses control of incidentMeasuring & Reducing Toil 19.4
What Counts as Toil
- Manual, repetitive operational work
- Running the same script every deploy
- Manually scaling resources reactively
- Copy-pasting secrets between systems
- Answering the same ticket type >2x/week
- Manually rotating SSL certificates
- Clicking through consoles to check status
Toil Elimination Targets
- Google SRE target: <50% time on toil
- Track toil weekly in team retrospectives
- Automate any task done >2x identical
- Self-service > ticket for all dev requests
- Auto-cert renewal via cert-manager
- Auto-scaling replacing manual scaling
- Runbook automation via SSM/Lambda
# Track toil via PagerDuty API — how many alerts are "noisy" vs actionable $ curl -H "Authorization: Token token=$PD_TOKEN" \ "https://api.pagerduty.com/incidents?time_zone=UTC&since=2024-01-01&until=2024-02-01&limit=100" \ | jq '.incidents[] | { service: .service.summary, resolved_by: .resolved_by_user.summary, duration_min: ((.last_status_change_at | fromdateiso8601) - (.created_at | fromdateiso8601)) / 60 | round, acknowledged: (.acknowledgements | length) }' | jq -s 'group_by(.service) | .[] | { service: .[0].service, total_pages: length, avg_duration_min: ([.[].duration_min] | add / length | round) }' { "service": "payments-api", "total_pages": 47, "avg_duration_min": 12 } ← 47 pages in 1 month = investigate noisy alert or recurring bug
Hands-On Labs 19.5
SRE Reliability Labs
- Define SLO for a sample service using OpenSLO spec, implement recording rules in Prometheus
- Configure multi-window burn-rate alerting (1h 14x + 6h 6x) in Alertmanager
- Install LitmusChaos on kind, run pod-delete experiment on a sample deployment
- Write OOMKill auto-remediation script and test it by manually creating OOM conditions
- Configure Alertmanager webhook route that calls a remediation endpoint for OOMKill alerts
- Write a complete blameless postmortem for the Redis connection pool incident (use template)
- Plan and document a Chaos Game Day agenda for a 3-service system with clear abort criteria
- Track your own toil for 1 week — log every manual task and identify top 3 automation opportunities
Troubleshooting 19.6
Use the two-window approach: both the fast (1h) and slow (6h) windows must fire simultaneously. Add a for: 2m to avoid transient spikes. Review your SLI query — if it includes health checks or internal traffic, exclude them. Use job_name != "blackbox-health" filters.
Check: (1) The target pod has the correct label matching applabel in ChaosEngine. (2) ChaosServiceAccount has RBAC permissions to delete pods. (3) annotationCheck: 'false' — if true, pods need litmuschaos.io/chaos: "true" annotation. Check ChaosEngine status: kubectl describe chaosengine <name> -n <ns>.
Add a cooldown period — track last remediation timestamp in a ConfigMap or DynamoDB table. Only trigger remediation if last run was >30 minutes ago. Cap auto-escalation: allow max 2 automatic increases, then create a Jira ticket for human review. Always Slack-notify with a clear "MANUAL ACTION REQUIRED IF PERSISTS" message.