🔥Prometheus & Alerting
Pull-based time series database. Master PromQL, recording rules, alerting rules, AlertManager, exporters, and production SLO-based monitoring on Kubernetes.
Architecture & How It Works 08.1
Prometheus scrapes /metrics endpoints on a fixed interval. Data is stored locally in TSDB. AlertManager receives alerts and handles routing, grouping, inhibition and silencing.
Core Components 08.2
| Type | Behavior | PromQL Example |
|---|---|---|
| Counter | Monotonically increasing | rate(http_requests_total[5m]) |
| Gauge | Can go up or down | node_memory_MemAvailable_bytes |
| Histogram | Bucketed observations | histogram_quantile(0.95, rate(...)) |
| Summary | Precomputed quantiles | quantile(0.9, ...) |
WSL Hands-On Lab 08.3
Full Observability Stack: Prometheus + AlertManager + Grafana + Custom Metrics
$ docker compose up -d $ curl http://localhost:9090/api/v1/query?query=up
global:
scrape_interval: 15s
scrape_configs:
- job_name: 'node'
static_configs:
- targets: ['node-exporter:9100']
- job_name: 'app'
static_configs:
- targets: ['host.docker.internal:9464']- Deploy full Prometheus + Grafana + AlertManager stack via Compose
- Instrument a simple Express app with prom-client and expose /metrics
- Write recording rule and alerting rule for high CPU usage
- Configure AlertManager to route to Slack or email
- Complete Killercoda Prometheus and Grafana Play labs
Real-World Project 08.4
node-exporter + basic application metrics dashboard
kube-prometheus-stack on EKS + custom ServiceMonitors + PagerDuty
Multi-cluster federation + SLO recording rules + HA AlertManager
Troubleshooting 08.5
Verify scrape config, network, labels, and relabeling rules.
Use query inspector in Prometheus UI. Check for clause and vector matching.
Drop unwanted labels early with relabel_configs or metric_relabel_configs.
30-Day Roadmap 08.6
- Four metric types
- rate/irate/increase
- by/without/group_left
- for: clause
- labels + annotations
- AlertManager config
- ServiceMonitor/PodMonitor
- kube-state-metrics
- Recording rules
- Error budget alerts
- Federation
- HA AlertManager
Deep Dive: SLOs and Advanced Alerting 08.7
Defining SLOs
Availability = good events / total events. Error budget = 1 - SLO target. Alert when burn rate is too high for the window. Use multi-window, multi-burn-rate alerts.
Recording Rules for Performance
Pre-compute expensive queries (rate over long windows, complex joins). Store results as new time series that are fast to query from Grafana.