Skip to content
DevOps Architect

    Syllabus / Advanced / 14

    2026 HOT MARKET SKILL
    14 / 14 • CLOUD NATIVE AI INFRASTRUCTURE

    🤖AI & LLMOps Infrastructure

    Scale generative AI models from prototype to enterprise production. Master Kubernetes GPU scheduling, vLLM continuous batching, PagedAttention, Karpenter Spot GPU pooling, KubeRay, and OpenTelemetry LLM tracing.

    Ollama + WSL vLLM + Karpenter Spot KubeRay + MIG + Triton Killercoda

    Architecture & How It Works 14.1

    Running LLMs in production differs fundamentally from traditional microservices. LLMs require memory-bound computing where VRAM bandwidth is the primary bottleneck. Kubernetes orchestrates physical GPUs via the NVIDIA Container Toolkit and NVIDIA GPU Operator, while serving engines (vLLM, Triton) implement PagedAttention to eliminate memory fragmentation.

    graph TD subgraph "Client Tier" App[Agentic Apps / Users] Gateway[Envoy AI Gateway • Rate Limiting & Prompt Caching] end subgraph "Kubernetes AI Cluster (EKS / Bare Metal)" subgraph "Serving Tier" vLLM1[vLLM Pod 1 • Llama-3-70B • PagedAttention] vLLM2[vLLM Pod 2 • Mistral-Large • vGPU MIG 3g.40gb] end subgraph "Orchestration & Autoscaling" KEDA[KEDA Prometheus Scaler • Queue Length Trigger] Karpenter[Karpenter GPU NodePool • Multi-AZ Spot Pooling] end subgraph "Hardware & Kernel Layer" NV[NVIDIA GPU Operator • driver/container-toolkit] GPU[(Physical GPUs: NVIDIA H100 / A100 / L40S)] end subgraph "Observability" OTel[OpenTelemetry Collector • GenAI Semantic Conventions] Grafana[Grafana Dashboard: TTFT, TPS, KV Cache Util] end end App --> Gateway Gateway --> vLLM1 Gateway --> vLLM2 vLLM1 --> NV vLLM2 --> NV NV --> GPU vLLM1 -.->|Metrics| KEDA KEDA --> Karpenter vLLM1 -.->|Traces: Time-to-First-Token| OTel OTel --> Grafana

    PagedAttention (vLLM)

    Traditional serving wastes up to 70% of GPU memory due to internal/external fragmentation and reserved space for future tokens. PagedAttention allocates KV-cache memory in non-contiguous pages like virtual memory in modern OS kernels.

    MIG vs Time-Slicing

    Multi-Instance GPU (MIG) creates hardware-isolated GPU instances (dedicated compute, memory, and memory bandwidth) on NVIDIA A100/H100. Time-slicing shares compute without memory isolation (suitable only for dev environments).

    Core Components & 2026 Tools 14.2

    ComponentTechnology2026 Production RoleStaff+ Trade-off
    High-Throughput Serving vLLM / SGLang Continuous batching, PagedAttention, Tensor Parallelism across multi-GPU Higher throughput than HuggingFace TGI; requires tuning max model length to prevent OOM.
    Enterprise Multi-Model Triton Inference Server Concurrent model execution (PyTorch, ONNX, TensorRT-LLM, vLLM backend) Extremely versatile for mixed vision/LLM pipelines; steep C++ configuration learning curve.
    Distributed Training KubeRay (Ray on K8s) Multi-node distributed fine-tuning (LoRA / QLoRA / Full fine-tune) and RLHF Automates Ray cluster lifecycle; requires NCCL network tuning over AWS EFA / InfiniBand.
    GPU Elasticity Karpenter GPU NodePool Sub-minute provisioning of Spot GPU nodes (`g5.12xlarge`, `p4de.24xlarge`) 60-70% cost savings on AWS; requires interruption draining with graceful token completion.
    AI Observability OpenTelemetry LLM Traces Tracks TTFT (Time-to-First-Token), TPOT (Time-per-Output-Token), prompt tokens, cost Standardized OpenInference schema; avoids proprietary vendor lock-in.

    WSL2 Hands-On Lab: Deploy Local Ollama & vLLM with Prometheus 14.3

    Run Local LLM Inference API + Prometheus Metrics Scraping in WSL2 (100% Free)

    wsl-llmops-lab.sh
    $ # 1. Install Ollama local runtime on WSL2
    $ curl -fsSL https://ollama.com/install.sh | sh
    
    $ # 2. Pull lean lightweight Llama-3.2-1B model
    $ ollama pull llama3.2:1b
    
    $ # 3. Spin up Ollama in background & expose OpenAI-compatible REST API on port 11434
    $ OLLAMA_HOST=0.0.0.0:11434 ollama serve &
    
    $ # 4. Test inference request via curl
    $ curl -s http://localhost:11434/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "llama3.2:1b",
        "messages": [{"role": "user", "content": "Explain Kubernetes eBPF in one sentence."}],
        "temperature": 0.2
      }' | jq .choices[0].message.content
    vllm-deployment.yaml
    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: vllm-inference
      namespace: ai-serving
    spec:
      replicas: 2
      selector:
        matchLabels:
          app: vllm-llama3
      template:
        metadata:
          labels:
            app: vllm-llama3
          annotations:
            prometheus.io/scrape: "true"
            prometheus.io/port: "8000"
            prometheus.io/path: "/metrics"
        spec:
          containers:
          - name: vllm
            image: vllm/vllm-openai:v0.6.3
            args:
            - "--model=meta-llama/Llama-3.2-1B-Instruct"
            - "--gpu-memory-utilization=0.90"
            - "--max-model-len=4096"
            - "--port=8000"
            resources:
              limits:
                nvidia.com/gpu: "1"
                memory: 16Gi
              requests:
                nvidia.com/gpu: "1"
                memory: 8Gi
            ports:
            - containerPort: 8000
            readinessProbe:
              httpGet:
                path: /health
                port: 8000
              initialDelaySeconds: 30
              periodSeconds: 10
    • Install Ollama and run a test completion on WSL2
    • Send an OpenAI-compatible REST API request via curl to port 11434
    • Review the production vLLM PagedAttention Kubernetes deployment manifest
    • Verify Prometheus metrics endpoints for TTFT (Time-to-First-Token)
    • Configure Karpenter GPU NodePool for spot auto-recovery
    • Complete Killercoda AI Infrastructure lab

    Real-World Scale Progression 14.4

    Startup Strategy: Local WSL2 + Serverless API Fallback

    Run Ollama and quantized small models (Llama 3.2 1B/3B, Mistral 7B) locally on WSL2 for development and CI testing. In production, route to serverless OpenAI/Anthropic/DeepSeek endpoints with prompt caching to avoid upfront GPU reservation costs.

    SME Strategy: EKS + Karpenter Spot GPUs + vLLM

    Deploy self-hosted vLLM on AWS EKS using Karpenter Spot instances (`g5.4xlarge`, `g5.12xlarge`). Implement KEDA scaling based on average pending prompt queue depth. Saves up to 70% compared to on-demand instances.

    Enterprise Strategy: Multi-Node KubeRay + Triton + H100 Fleet + OTel

    Multi-cluster Kubernetes with NVIDIA GPU Operator, hardware MIG slicing, KubeRay distributed fine-tuning, Triton Inference Server with TensorRT-LLM, and end-to-end OpenTelemetry LLM tracing for token cost attribution across business units.

    Troubleshooting & Production Postmortems 14.5

    Cause: Concurrent requests with long prompts exceed the dynamic KV-cache allocation pool.

    Fix: Set --gpu-memory-utilization 0.90 in vLLM to leave a 10% safety cushion. Enable prefix caching (--enable-prefix-caching) so identical system prompts share KV cache blocks.

    Cause: AWS VPC network MTU mismatch or missing EFA (Elastic Fabric Adapter) kernel interfaces between worker pods.

    Fix: Verify jumbo frames (MTU 9001) across VPC subnets and set NCCL_DEBUG=INFO NCCL_SOCKET_IFNAME=eth0 in container environment variables.

    2026 LLMOps Career Roadmap & Interview Tips 14.6

    Key 2026 Interview Soundbites

    • Time-to-First-Token (TTFT): Governed by Prefill phase compute throughput.
    • Time-per-Output-Token (TPOT): Governed by Decode phase memory bandwidth.
    • PagedAttention: Virtual memory paging for GPU KV-cache blocks.

    Next Step: The War Room

    Ready to test your live incident triage skills? Head into the Staff/Principal War Room Simulator.

    Enter War Room Simulator

    Run: curl -fsSL https://ollama.com/install.sh | sh

    Extra commands from this lesson (9) are kept out of this page. Quizzes were not in the source HTML.