Skip to content
DevOps Architect

    Syllabus / Operations / 08

    STARTUP TO ENTERPRISE
    08 / 14 • METRICS & ALERTING

    🔥Prometheus & Alerting

    Pull-based time series database. Master PromQL, recording rules, alerting rules, AlertManager, exporters, and production SLO-based monitoring on Kubernetes.

    node-exporter kube-prom-stack Federation + SLOs

    Architecture & How It Works 08.1

    Prometheus scrapes /metrics endpoints on a fixed interval. Data is stored locally in TSDB. AlertManager receives alerts and handles routing, grouping, inhibition and silencing.

    graph TD P[Prometheus] -- scrape --> Node[node-exporter] P -- scrape --> KSM[kube-state-metrics] P -- scrape --> App[Application metrics] P -- alerts --> AM[AlertManager] AM --> Slack AM --> PagerDuty

    Core Components 08.2

    TypeBehaviorPromQL Example
    CounterMonotonically increasingrate(http_requests_total[5m])
    GaugeCan go up or downnode_memory_MemAvailable_bytes
    HistogramBucketed observationshistogram_quantile(0.95, rate(...))
    SummaryPrecomputed quantilesquantile(0.9, ...)

    WSL Hands-On Lab 08.3

    Full Observability Stack: Prometheus + AlertManager + Grafana + Custom Metrics

    stack
    $ docker compose up -d
    $ curl http://localhost:9090/api/v1/query?query=up
    prometheus.yml
    global:
      scrape_interval: 15s
    scrape_configs:
      - job_name: 'node'
        static_configs:
          - targets: ['node-exporter:9100']
      - job_name: 'app'
        static_configs:
          - targets: ['host.docker.internal:9464']
    • Deploy full Prometheus + Grafana + AlertManager stack via Compose
    • Instrument a simple Express app with prom-client and expose /metrics
    • Write recording rule and alerting rule for high CPU usage
    • Configure AlertManager to route to Slack or email
    • Complete Killercoda Prometheus and Grafana Play labs

    Real-World Project 08.4

    node-exporter + basic application metrics dashboard

    kube-prometheus-stack on EKS + custom ServiceMonitors + PagerDuty

    Multi-cluster federation + SLO recording rules + HA AlertManager

    Troubleshooting 08.5

    Verify scrape config, network, labels, and relabeling rules.

    Use query inspector in Prometheus UI. Check for clause and vector matching.

    Drop unwanted labels early with relabel_configs or metric_relabel_configs.

    30-Day Roadmap 08.6

    WEEK 1
    Metrics & PromQL
    • Four metric types
    • rate/irate/increase
    • by/without/group_left
    WEEK 2
    Alerting Rules
    • for: clause
    • labels + annotations
    • AlertManager config
    WEEK 3
    Kubernetes Monitoring
    • ServiceMonitor/PodMonitor
    • kube-state-metrics
    • Recording rules
    WEEK 4
    SLOs & Advanced
    • Error budget alerts
    • Federation
    • HA AlertManager

    Deep Dive: SLOs and Advanced Alerting 08.7

    Defining SLOs

    Availability = good events / total events. Error budget = 1 - SLO target. Alert when burn rate is too high for the window. Use multi-window, multi-burn-rate alerts.

    Recording Rules for Performance

    Pre-compute expensive queries (rate over long windows, complex joins). Store results as new time series that are fast to query from Grafana.

    Run: docker compose up -d

    Extra commands from this lesson (1) are kept out of this page. Quizzes were not in the source HTML.