Skip to content
DevOps Architect

    Syllabus / Operations / 17

    OTEL + GRAFANA STACK
    17 / 19 • DISTRIBUTED TRACING & FULL-STACK OBSERVABILITY

    🔭OpenTelemetry & Observability

    Stop flying blind in microservices. Master OpenTelemetry auto-instrumentation, Collector pipeline design, Grafana Tempo distributed tracing, Loki log correlation, eBPF zero-code observability with Beyla, and span-level SLOs. From "we have logs" to "we have answers in 30 seconds" for any production incident.

    OTel SDK + Jaeger Tempo + Loki + Exemplars eBPF Beyla + Tail Sampling

     Real Incident: "Checkout is slow for 8% of users"

    Symptom: P99 checkout latency >8s (SLO breach). Prometheus shows average latency fine. No errors. 3 teams pointing fingers.

    Without tracing: 4 engineers spending 3 hours looking at logs, metrics, guessing. Eventually blame shifts to "must be the database."

    With OpenTelemetry + Tempo:

    # Query Tempo for slow traces (>3s) in checkout service
    # TraceQL — Grafana Tempo's query language
    { .service.name = "checkout-api" && duration > 3s }
    
    # Result: 8% of traces show a specific span slow:
    # checkout-api → inventory-service → [SLOW: 6.2s avg]
    #   └─ span: inventory.check_availability
    #      └─ db.query: SELECT * FROM inventory WHERE sku_id IN (...)
    #         └─ db.rows_returned: 84,000 rows ← no index on sku_id!
    
    # Cross-correlate with Loki logs using trace_id
    {app="inventory-service"} |= "trace_id=abc123def456"
    
    2024-01-15T14:22:33Z WARN slow query detected duration=6.2s query="SELECT..." rows=84000
    2024-01-15T14:22:33Z INFO trace_id=abc123def456 span_id=xyz789

    Resolution: 22 minutes to identify root cause. Added index on sku_id. P99 dropped from 8s to 180ms. No more guessing, no more blaming teams.

    OTel Architecture: The Three Pillars 17.1

    graph TD subgraph "Application Layer" App1[FastAPI Service
    OTel SDK Auto-Instrument] App2[Node.js Service
    OTel SDK Auto-Instrument] App3[Java Service
    OTel Java Agent] Beyla[eBPF Beyla
    Zero-code sidecar] end subgraph "OTel Collector (DaemonSet)" Recv[OTLP Receiver
    :4317 gRPC / :4318 HTTP] Proc[Processors
    Batch, Filter, Tail-Sampling] Export[Exporters] end subgraph "Backends" Tempo[Grafana Tempo
    Traces] Loki[Grafana Loki
    Logs] Prom[Prometheus
    Metrics] Grafana[Grafana
    Unified Dashboard + Alerts] end App1 -->|OTLP| Recv App2 -->|OTLP| Recv App3 -->|OTLP| Recv Beyla -->|OTLP| Recv Recv --> Proc Proc -->|traces| Tempo Proc -->|logs| Loki Proc -->|metrics| Prom Tempo --> Grafana Loki --> Grafana Prom --> Grafana

    OTel Collector: Enterprise Pipeline 17.2

    otel-collector-config.yaml
    receivers:
      otlp:
        protocols:
          grpc:
            endpoint: 0.0.0.0:4317
          http:
            endpoint: 0.0.0.0:4318
            cors:
              allowed_origins: ["https://*.acme.com"]
      prometheus:
        config:
          scrape_configs:
            - job_name: 'k8s-pods'
              kubernetes_sd_configs:
                - role: pod
              relabel_configs:
                - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
                  action: keep
                  regex: true
    
    processors:
      batch:
        timeout: 5s
        send_batch_size: 1024
      memory_limiter:
        check_interval: 1s
        limit_mib: 512
        spike_limit_mib: 128
      resource:
        attributes:
          - action: upsert
            key: cluster.name
            value: prod-eks-us-east-1
          - action: upsert
            key: deployment.environment
            from_attribute: k8s.namespace.name
      # Add trace IDs to logs for correlation
      transform/logs:
        log_statements:
          - context: log
            statements:
              - set(attributes["trace_id"], trace_id.string) where trace_id != SpanID(0x0000000000000000)
    
    exporters:
      otlp/tempo:
        endpoint: tempo.monitoring.svc.cluster.local:4317
        tls:
          insecure: true
      loki:
        endpoint: http://loki.monitoring.svc.cluster.local:3100/loki/api/v1/push
        labels:
          resource:
            service.name: "service_name"
            k8s.namespace.name: "namespace"
      prometheusremotewrite:
        endpoint: http://prometheus.monitoring.svc.cluster.local:9090/api/v1/write
    
    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [memory_limiter, batch, resource]
          exporters: [otlp/tempo]
        logs:
          receivers: [otlp]
          processors: [memory_limiter, batch, transform/logs]
          exporters: [loki]
        metrics:
          receivers: [otlp, prometheus]
          processors: [memory_limiter, batch]
          exporters: [prometheusremotewrite]
    tail-sampling-policy.yaml
    # Tail-based sampling — make decisions AFTER seeing full trace
    # Critical for high-volume services (>10k req/s)
    # Head sampling (random %) loses important error traces
    processors:
      tail_sampling:
        decision_wait: 10s
        num_traces: 100000
        expected_new_traces_per_sec: 10000
        policies:
          # Always keep error traces
          - name: error-policy
            type: status_code
            status_code: {status_codes: [ERROR]}
    
          # Always keep slow traces (>500ms)
          - name: latency-policy
            type: latency
            latency: {threshold_ms: 500}
    
          # Keep 1% of healthy fast traces (for baseline)
          - name: probabilistic-policy
            type: probabilistic
            probabilistic: {sampling_percentage: 1}
    
          # Always keep traces with specific attributes
          # (e.g., payment transactions, new user signup)
          - name: always-sample-payments
            type: string_attribute
            string_attribute:
              key: transaction.type
              values: [payment, refund, chargeback]
    
          # Composite: slow OR error in checkout service
          - name: checkout-composite
            type: composite
            composite:
              max_total_spans_per_second: 1000
              policy_order: [error-policy, latency-policy]
              composite_sub_policy:
                - name: checkout-service-filter
                  type: string_attribute
                  string_attribute:
                    key: service.name
                    values: [checkout-api]
    otel-collector-daemonset.yaml
    apiVersion: apps/v1
    kind: DaemonSet
    metadata:
      name: otel-collector
      namespace: monitoring
    spec:
      selector:
        matchLabels:
          app: otel-collector
      template:
        metadata:
          labels:
            app: otel-collector
        spec:
          serviceAccountName: otel-collector
          containers:
            - name: otel-collector
              image: otel/opentelemetry-collector-contrib:0.95.0
              args: ["--config=/conf/otel-collector-config.yaml"]
              resources:
                requests:
                  cpu: 200m
                  memory: 400Mi
                limits:
                  cpu: 1000m
                  memory: 1Gi
              ports:
                - containerPort: 4317  # OTLP gRPC
                - containerPort: 4318  # OTLP HTTP
                - containerPort: 8888  # Prometheus metrics (self-monitoring)
              volumeMounts:
                - name: config
                  mountPath: /conf
              env:
                - name: KUBE_NODE_NAME
                  valueFrom:
                    fieldRef:
                      fieldPath: spec.nodeName
          volumes:
            - name: config
              configMap:
                name: otel-collector-config
    otel-auto-instrumentation.yaml (zero-code!)
    # OpenTelemetry Operator — inject instrumentation automatically
    # No code changes needed. Just add an annotation to the Deployment.
    apiVersion: opentelemetry.io/v1alpha1
    kind: Instrumentation
    metadata:
      name: acme-instrumentation
      namespace: production
    spec:
      exporter:
        endpoint: http://otel-collector.monitoring.svc.cluster.local:4317
      propagators:
        - tracecontext
        - baggage
        - b3
      sampler:
        type: parentbased_traceidratio
        argument: "1.0"   # 100% sampling (use tail sampling in collector)
      python:
        image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-python:0.43b0
      java:
        image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-java:1.32.0
      nodejs:
        image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-nodejs:0.41.0
      go:
        image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-go:0.8.0
    ---
    # Just add this annotation to any Deployment — zero code changes!
    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: checkout-api
      namespace: production
    spec:
      template:
        metadata:
          annotations:
            instrumentation.opentelemetry.io/inject-python: "acme-instrumentation"
            # For Java: instrumentation.opentelemetry.io/inject-java: "acme-instrumentation"

    Beyla: eBPF Zero-Code Instrumentation 17.3

    Grafana Beyla uses eBPF to capture HTTP/gRPC traces, RED metrics (Rate, Errors, Duration), and TCP connections from any application without code changes, SDKs, or language-specific agents. Ideal for legacy apps, compiled binaries, or polyglot environments.

    beyla-daemonset.yaml
    apiVersion: apps/v1
    kind: DaemonSet
    metadata:
      name: beyla
      namespace: monitoring
    spec:
      template:
        spec:
          hostPID: true   # Required for eBPF process inspection
          containers:
            - name: beyla
              image: grafana/beyla:1.4.0
              securityContext:
                privileged: true   # Required for eBPF
                runAsUser: 0
              env:
                - name: BEYLA_KUBE_METADATA_ENABLE
                  value: "true"
                - name: BEYLA_OTEL_EXPORTER_ENDPOINT
                  value: "http://otel-collector.monitoring.svc.cluster.local:4317"
                - name: BEYLA_SERVICE_NAMESPACE
                  value: "production"
              volumeMounts:
                - name: beyla-config
                  mountPath: /config
              args: ["--config=/config/beyla-config.yaml"]
          volumes:
            - name: beyla-config
              configMap:
                name: beyla-config
    beyla-config.yaml
    discovery:
      services:
        # Instrument ALL HTTP services on standard ports
        - k8s_namespace: production
          k8s_deployment_name: .*   # regex: all deployments
          open_ports: 8080,8443,3000,9090
    
    trace:
      print_traces: false
    
    otel_traces_export:
      endpoint: http://otel-collector.monitoring.svc.cluster.local:4317
    
    otel_metrics_export:
      endpoint: http://otel-collector.monitoring.svc.cluster.local:4317
      features:
        # Export RED metrics: http_server_request_duration_seconds histogram
        - otel-metrics-net
        - application
    
    # Filter out health check noise
    filters:
      application:
        - metadata.name: health.*
          action: exclude

    Grafana Tempo + TraceQL 17.4

    traceql-examples.txt
    # Find all slow traces in checkout service
    { .service.name = "checkout-api" && duration > 2s }
    
    # Find traces with errors in any service
    { status = error }
    
    # Find traces involving both payment-api and fraud-service
    { .service.name = "payment-api" } >> { .service.name = "fraud-service" }
    
    # Count slow DB queries per service (span-level analysis)
    { .db.system = "postgresql" && duration > 500ms }
    | by(.service.name)
    | count() > 10
    
    # P99 latency per service — use in alerting
    { .service.name =~ ".*-api" }
    | by(.service.name)
    | p99(duration)
    
    # Find traces where inventory span is slow but parent is healthy
    { .service.name = "inventory-service" && duration > 1s && parent.duration < 200ms }
    loki-trace-correlation.logql
    # Step 1: Find slow trace IDs in Tempo (TraceQL)
    { .service.name = "checkout-api" && duration > 3s }
    # → Returns trace IDs like: abc123def456, xyz789abc012
    
    # Step 2: Jump to Loki logs for that trace (Grafana does this automatically)
    # In Loki data source, Derived Fields config:
    #   Name: TraceID
    #   Regex: "trace_id=(\w+)"
    #   Internal Link: Tempo data source
    #   Query: ${__value.raw}
    
    # Step 3: LogQL to find all logs in that trace
    {namespace="production"} |= "trace_id=abc123def456"
    
    # Step 4: LogQL with JSON parsing to find the actual error
    {app="checkout-api"} | json | level="error" | duration_ms > 3000
    
    # Step 5: Histogram of log volume correlated with error rate
    sum by (app) (
      rate({namespace="production", level="error"}[5m])
    )
    exemplars-config.py
    # Python FastAPI — Prometheus exemplars linking metrics to traces
    from prometheus_client import Histogram
    from opentelemetry import trace
    
    REQUEST_DURATION = Histogram(
        'http_request_duration_seconds',
        'HTTP request duration',
        ['method', 'endpoint', 'status']
    )
    
    @app.middleware("http")
    async def metrics_middleware(request: Request, call_next):
        start = time.time()
        response = await call_next(request)
        duration = time.time() - start
    
        # Get current trace context for exemplar
        span = trace.get_current_span()
        trace_id = format(span.get_span_context().trace_id, '032x')
    
        # Record histogram with exemplar (trace_id attached to data point)
        REQUEST_DURATION.labels(
            method=request.method,
            endpoint=request.url.path,
            status=response.status_code
        ).observe(
            duration,
            exemplar={'TraceID': trace_id}   # Links metric spike to trace
        )
    
        return response
    
    # In Grafana: enable exemplars on Prometheus data source
    # When you see a latency spike, click the diamond = jump to trace

    Hands-On Labs 17.5

    Observability Labs — Grafana Stack on kind

    • Deploy Grafana stack (Tempo, Loki, Prometheus, Grafana) via Helm on kind cluster
    • Deploy OTel Collector as DaemonSet with the enterprise config — verify OTLP receiver working
    • Install OTel Operator, create Instrumentation CR, annotate a Python app — verify traces in Tempo
    • Write TraceQL query to find all error traces with duration >1s in last 1 hour
    • Configure Loki Derived Fields to create trace ID links — jump from log line to trace
    • Build a tail-sampling policy that keeps 100% of errors + slow spans, 1% of healthy spans
    • Deploy Grafana Beyla as DaemonSet — verify RED metrics appear without code changes

    Troubleshooting 17.6

    Check OTel Collector logs: kubectl logs -n monitoring ds/otel-collector. Common issues: application not exporting to correct endpoint (check OTEL_EXPORTER_OTLP_ENDPOINT env var), network policy blocking port 4317, Tempo not receiving via its OTLP ingestion endpoint. Use otelcol --feature-gates=telemetry.useOtelForInternalMetrics to debug collector internals.

    Check: (1) Instrumentation CR namespace matches Deployment namespace. (2) The correct annotation key used (e.g., inject-python vs inject-java). (3) OTel Operator webhook is running: kubectl get pods -n opentelemetry-operator-system. (4) Pod must be restarted after annotation added. Check operator logs for admission webhook errors.

    Add memory_limiter processor (first in pipeline). Increase DaemonSet memory limits. Consider scaling to a Gateway Collector deployment for high-volume scenarios. Enable tail sampling to reduce volume before exporting. Monitor collector's own metrics via :8888/metrics.

    Run: { .service.name = "checkout-api" && duration > 3s }

    Extra commands from this lesson (1) are kept out of this page. Quizzes were not in the source HTML.