🔭OpenTelemetry & Observability
Stop flying blind in microservices. Master OpenTelemetry auto-instrumentation, Collector pipeline design, Grafana Tempo distributed tracing, Loki log correlation, eBPF zero-code observability with Beyla, and span-level SLOs. From "we have logs" to "we have answers in 30 seconds" for any production incident.
Real Incident: "Checkout is slow for 8% of users"
Symptom: P99 checkout latency >8s (SLO breach). Prometheus shows average latency fine. No errors. 3 teams pointing fingers.
Without tracing: 4 engineers spending 3 hours looking at logs, metrics, guessing. Eventually blame shifts to "must be the database."
With OpenTelemetry + Tempo:
# Query Tempo for slow traces (>3s) in checkout service # TraceQL — Grafana Tempo's query language { .service.name = "checkout-api" && duration > 3s } # Result: 8% of traces show a specific span slow: # checkout-api → inventory-service → [SLOW: 6.2s avg] # └─ span: inventory.check_availability # └─ db.query: SELECT * FROM inventory WHERE sku_id IN (...) # └─ db.rows_returned: 84,000 rows ← no index on sku_id! # Cross-correlate with Loki logs using trace_id {app="inventory-service"} |= "trace_id=abc123def456" 2024-01-15T14:22:33Z WARN slow query detected duration=6.2s query="SELECT..." rows=84000 2024-01-15T14:22:33Z INFO trace_id=abc123def456 span_id=xyz789
Resolution: 22 minutes to identify root cause. Added index on sku_id. P99 dropped from 8s to 180ms. No more guessing, no more blaming teams.
OTel Architecture: The Three Pillars 17.1
OTel SDK Auto-Instrument] App2[Node.js Service
OTel SDK Auto-Instrument] App3[Java Service
OTel Java Agent] Beyla[eBPF Beyla
Zero-code sidecar] end subgraph "OTel Collector (DaemonSet)" Recv[OTLP Receiver
:4317 gRPC / :4318 HTTP] Proc[Processors
Batch, Filter, Tail-Sampling] Export[Exporters] end subgraph "Backends" Tempo[Grafana Tempo
Traces] Loki[Grafana Loki
Logs] Prom[Prometheus
Metrics] Grafana[Grafana
Unified Dashboard + Alerts] end App1 -->|OTLP| Recv App2 -->|OTLP| Recv App3 -->|OTLP| Recv Beyla -->|OTLP| Recv Recv --> Proc Proc -->|traces| Tempo Proc -->|logs| Loki Proc -->|metrics| Prom Tempo --> Grafana Loki --> Grafana Prom --> Grafana
OTel Collector: Enterprise Pipeline 17.2
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
cors:
allowed_origins: ["https://*.acme.com"]
prometheus:
config:
scrape_configs:
- job_name: 'k8s-pods'
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
processors:
batch:
timeout: 5s
send_batch_size: 1024
memory_limiter:
check_interval: 1s
limit_mib: 512
spike_limit_mib: 128
resource:
attributes:
- action: upsert
key: cluster.name
value: prod-eks-us-east-1
- action: upsert
key: deployment.environment
from_attribute: k8s.namespace.name
# Add trace IDs to logs for correlation
transform/logs:
log_statements:
- context: log
statements:
- set(attributes["trace_id"], trace_id.string) where trace_id != SpanID(0x0000000000000000)
exporters:
otlp/tempo:
endpoint: tempo.monitoring.svc.cluster.local:4317
tls:
insecure: true
loki:
endpoint: http://loki.monitoring.svc.cluster.local:3100/loki/api/v1/push
labels:
resource:
service.name: "service_name"
k8s.namespace.name: "namespace"
prometheusremotewrite:
endpoint: http://prometheus.monitoring.svc.cluster.local:9090/api/v1/write
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch, resource]
exporters: [otlp/tempo]
logs:
receivers: [otlp]
processors: [memory_limiter, batch, transform/logs]
exporters: [loki]
metrics:
receivers: [otlp, prometheus]
processors: [memory_limiter, batch]
exporters: [prometheusremotewrite]# Tail-based sampling — make decisions AFTER seeing full trace
# Critical for high-volume services (>10k req/s)
# Head sampling (random %) loses important error traces
processors:
tail_sampling:
decision_wait: 10s
num_traces: 100000
expected_new_traces_per_sec: 10000
policies:
# Always keep error traces
- name: error-policy
type: status_code
status_code: {status_codes: [ERROR]}
# Always keep slow traces (>500ms)
- name: latency-policy
type: latency
latency: {threshold_ms: 500}
# Keep 1% of healthy fast traces (for baseline)
- name: probabilistic-policy
type: probabilistic
probabilistic: {sampling_percentage: 1}
# Always keep traces with specific attributes
# (e.g., payment transactions, new user signup)
- name: always-sample-payments
type: string_attribute
string_attribute:
key: transaction.type
values: [payment, refund, chargeback]
# Composite: slow OR error in checkout service
- name: checkout-composite
type: composite
composite:
max_total_spans_per_second: 1000
policy_order: [error-policy, latency-policy]
composite_sub_policy:
- name: checkout-service-filter
type: string_attribute
string_attribute:
key: service.name
values: [checkout-api]apiVersion: apps/v1
kind: DaemonSet
metadata:
name: otel-collector
namespace: monitoring
spec:
selector:
matchLabels:
app: otel-collector
template:
metadata:
labels:
app: otel-collector
spec:
serviceAccountName: otel-collector
containers:
- name: otel-collector
image: otel/opentelemetry-collector-contrib:0.95.0
args: ["--config=/conf/otel-collector-config.yaml"]
resources:
requests:
cpu: 200m
memory: 400Mi
limits:
cpu: 1000m
memory: 1Gi
ports:
- containerPort: 4317 # OTLP gRPC
- containerPort: 4318 # OTLP HTTP
- containerPort: 8888 # Prometheus metrics (self-monitoring)
volumeMounts:
- name: config
mountPath: /conf
env:
- name: KUBE_NODE_NAME
valueFrom:
fieldRef:
fieldPath: spec.nodeName
volumes:
- name: config
configMap:
name: otel-collector-config# OpenTelemetry Operator — inject instrumentation automatically
# No code changes needed. Just add an annotation to the Deployment.
apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
name: acme-instrumentation
namespace: production
spec:
exporter:
endpoint: http://otel-collector.monitoring.svc.cluster.local:4317
propagators:
- tracecontext
- baggage
- b3
sampler:
type: parentbased_traceidratio
argument: "1.0" # 100% sampling (use tail sampling in collector)
python:
image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-python:0.43b0
java:
image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-java:1.32.0
nodejs:
image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-nodejs:0.41.0
go:
image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-go:0.8.0
---
# Just add this annotation to any Deployment — zero code changes!
apiVersion: apps/v1
kind: Deployment
metadata:
name: checkout-api
namespace: production
spec:
template:
metadata:
annotations:
instrumentation.opentelemetry.io/inject-python: "acme-instrumentation"
# For Java: instrumentation.opentelemetry.io/inject-java: "acme-instrumentation"Beyla: eBPF Zero-Code Instrumentation 17.3
Grafana Beyla uses eBPF to capture HTTP/gRPC traces, RED metrics (Rate, Errors, Duration), and TCP connections from any application without code changes, SDKs, or language-specific agents. Ideal for legacy apps, compiled binaries, or polyglot environments.
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: beyla
namespace: monitoring
spec:
template:
spec:
hostPID: true # Required for eBPF process inspection
containers:
- name: beyla
image: grafana/beyla:1.4.0
securityContext:
privileged: true # Required for eBPF
runAsUser: 0
env:
- name: BEYLA_KUBE_METADATA_ENABLE
value: "true"
- name: BEYLA_OTEL_EXPORTER_ENDPOINT
value: "http://otel-collector.monitoring.svc.cluster.local:4317"
- name: BEYLA_SERVICE_NAMESPACE
value: "production"
volumeMounts:
- name: beyla-config
mountPath: /config
args: ["--config=/config/beyla-config.yaml"]
volumes:
- name: beyla-config
configMap:
name: beyla-configdiscovery:
services:
# Instrument ALL HTTP services on standard ports
- k8s_namespace: production
k8s_deployment_name: .* # regex: all deployments
open_ports: 8080,8443,3000,9090
trace:
print_traces: false
otel_traces_export:
endpoint: http://otel-collector.monitoring.svc.cluster.local:4317
otel_metrics_export:
endpoint: http://otel-collector.monitoring.svc.cluster.local:4317
features:
# Export RED metrics: http_server_request_duration_seconds histogram
- otel-metrics-net
- application
# Filter out health check noise
filters:
application:
- metadata.name: health.*
action: excludeGrafana Tempo + TraceQL 17.4
# Find all slow traces in checkout service
{ .service.name = "checkout-api" && duration > 2s }
# Find traces with errors in any service
{ status = error }
# Find traces involving both payment-api and fraud-service
{ .service.name = "payment-api" } >> { .service.name = "fraud-service" }
# Count slow DB queries per service (span-level analysis)
{ .db.system = "postgresql" && duration > 500ms }
| by(.service.name)
| count() > 10
# P99 latency per service — use in alerting
{ .service.name =~ ".*-api" }
| by(.service.name)
| p99(duration)
# Find traces where inventory span is slow but parent is healthy
{ .service.name = "inventory-service" && duration > 1s && parent.duration < 200ms }# Step 1: Find slow trace IDs in Tempo (TraceQL)
{ .service.name = "checkout-api" && duration > 3s }
# → Returns trace IDs like: abc123def456, xyz789abc012
# Step 2: Jump to Loki logs for that trace (Grafana does this automatically)
# In Loki data source, Derived Fields config:
# Name: TraceID
# Regex: "trace_id=(\w+)"
# Internal Link: Tempo data source
# Query: ${__value.raw}
# Step 3: LogQL to find all logs in that trace
{namespace="production"} |= "trace_id=abc123def456"
# Step 4: LogQL with JSON parsing to find the actual error
{app="checkout-api"} | json | level="error" | duration_ms > 3000
# Step 5: Histogram of log volume correlated with error rate
sum by (app) (
rate({namespace="production", level="error"}[5m])
)# Python FastAPI — Prometheus exemplars linking metrics to traces
from prometheus_client import Histogram
from opentelemetry import trace
REQUEST_DURATION = Histogram(
'http_request_duration_seconds',
'HTTP request duration',
['method', 'endpoint', 'status']
)
@app.middleware("http")
async def metrics_middleware(request: Request, call_next):
start = time.time()
response = await call_next(request)
duration = time.time() - start
# Get current trace context for exemplar
span = trace.get_current_span()
trace_id = format(span.get_span_context().trace_id, '032x')
# Record histogram with exemplar (trace_id attached to data point)
REQUEST_DURATION.labels(
method=request.method,
endpoint=request.url.path,
status=response.status_code
).observe(
duration,
exemplar={'TraceID': trace_id} # Links metric spike to trace
)
return response
# In Grafana: enable exemplars on Prometheus data source
# When you see a latency spike, click the diamond = jump to traceHands-On Labs 17.5
Observability Labs — Grafana Stack on kind
- Deploy Grafana stack (Tempo, Loki, Prometheus, Grafana) via Helm on kind cluster
- Deploy OTel Collector as DaemonSet with the enterprise config — verify OTLP receiver working
- Install OTel Operator, create Instrumentation CR, annotate a Python app — verify traces in Tempo
- Write TraceQL query to find all error traces with duration >1s in last 1 hour
- Configure Loki Derived Fields to create trace ID links — jump from log line to trace
- Build a tail-sampling policy that keeps 100% of errors + slow spans, 1% of healthy spans
- Deploy Grafana Beyla as DaemonSet — verify RED metrics appear without code changes
Troubleshooting 17.6
Check OTel Collector logs: kubectl logs -n monitoring ds/otel-collector. Common issues: application not exporting to correct endpoint (check OTEL_EXPORTER_OTLP_ENDPOINT env var), network policy blocking port 4317, Tempo not receiving via its OTLP ingestion endpoint. Use otelcol --feature-gates=telemetry.useOtelForInternalMetrics to debug collector internals.
Check: (1) Instrumentation CR namespace matches Deployment namespace. (2) The correct annotation key used (e.g., inject-python vs inject-java). (3) OTel Operator webhook is running: kubectl get pods -n opentelemetry-operator-system. (4) Pod must be restarted after annotation added. Check operator logs for admission webhook errors.
Add memory_limiter processor (first in pipeline). Increase DaemonSet memory limits. Consider scaling to a Gateway Collector deployment for high-volume scenarios. Enable tail sampling to reduce volume before exporting. Monitor collector's own metrics via :8888/metrics.