🤖AI & LLMOps Infrastructure
Scale generative AI models from prototype to enterprise production. Master Kubernetes GPU scheduling, vLLM continuous batching, PagedAttention, Karpenter Spot GPU pooling, KubeRay, and OpenTelemetry LLM tracing.
Architecture & How It Works 14.1
Running LLMs in production differs fundamentally from traditional microservices. LLMs require memory-bound computing where VRAM bandwidth is the primary bottleneck. Kubernetes orchestrates physical GPUs via the NVIDIA Container Toolkit and NVIDIA GPU Operator, while serving engines (vLLM, Triton) implement PagedAttention to eliminate memory fragmentation.
PagedAttention (vLLM)
Traditional serving wastes up to 70% of GPU memory due to internal/external fragmentation and reserved space for future tokens. PagedAttention allocates KV-cache memory in non-contiguous pages like virtual memory in modern OS kernels.
MIG vs Time-Slicing
Multi-Instance GPU (MIG) creates hardware-isolated GPU instances (dedicated compute, memory, and memory bandwidth) on NVIDIA A100/H100. Time-slicing shares compute without memory isolation (suitable only for dev environments).
Core Components & 2026 Tools 14.2
| Component | Technology | 2026 Production Role | Staff+ Trade-off |
|---|---|---|---|
| High-Throughput Serving | vLLM / SGLang | Continuous batching, PagedAttention, Tensor Parallelism across multi-GPU | Higher throughput than HuggingFace TGI; requires tuning max model length to prevent OOM. |
| Enterprise Multi-Model | Triton Inference Server | Concurrent model execution (PyTorch, ONNX, TensorRT-LLM, vLLM backend) | Extremely versatile for mixed vision/LLM pipelines; steep C++ configuration learning curve. |
| Distributed Training | KubeRay (Ray on K8s) | Multi-node distributed fine-tuning (LoRA / QLoRA / Full fine-tune) and RLHF | Automates Ray cluster lifecycle; requires NCCL network tuning over AWS EFA / InfiniBand. |
| GPU Elasticity | Karpenter GPU NodePool | Sub-minute provisioning of Spot GPU nodes (`g5.12xlarge`, `p4de.24xlarge`) | 60-70% cost savings on AWS; requires interruption draining with graceful token completion. |
| AI Observability | OpenTelemetry LLM Traces | Tracks TTFT (Time-to-First-Token), TPOT (Time-per-Output-Token), prompt tokens, cost | Standardized OpenInference schema; avoids proprietary vendor lock-in. |
WSL2 Hands-On Lab: Deploy Local Ollama & vLLM with Prometheus 14.3
Run Local LLM Inference API + Prometheus Metrics Scraping in WSL2 (100% Free)
$ # 1. Install Ollama local runtime on WSL2 $ curl -fsSL https://ollama.com/install.sh | sh $ # 2. Pull lean lightweight Llama-3.2-1B model $ ollama pull llama3.2:1b $ # 3. Spin up Ollama in background & expose OpenAI-compatible REST API on port 11434 $ OLLAMA_HOST=0.0.0.0:11434 ollama serve & $ # 4. Test inference request via curl $ curl -s http://localhost:11434/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "llama3.2:1b", "messages": [{"role": "user", "content": "Explain Kubernetes eBPF in one sentence."}], "temperature": 0.2 }' | jq .choices[0].message.content
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-inference
namespace: ai-serving
spec:
replicas: 2
selector:
matchLabels:
app: vllm-llama3
template:
metadata:
labels:
app: vllm-llama3
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8000"
prometheus.io/path: "/metrics"
spec:
containers:
- name: vllm
image: vllm/vllm-openai:v0.6.3
args:
- "--model=meta-llama/Llama-3.2-1B-Instruct"
- "--gpu-memory-utilization=0.90"
- "--max-model-len=4096"
- "--port=8000"
resources:
limits:
nvidia.com/gpu: "1"
memory: 16Gi
requests:
nvidia.com/gpu: "1"
memory: 8Gi
ports:
- containerPort: 8000
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 30
periodSeconds: 10- Install Ollama and run a test completion on WSL2
- Send an OpenAI-compatible REST API request via curl to port 11434
- Review the production vLLM PagedAttention Kubernetes deployment manifest
- Verify Prometheus metrics endpoints for TTFT (Time-to-First-Token)
- Configure Karpenter GPU NodePool for spot auto-recovery
- Complete Killercoda AI Infrastructure lab
Real-World Scale Progression 14.4
Startup Strategy: Local WSL2 + Serverless API Fallback
Run Ollama and quantized small models (Llama 3.2 1B/3B, Mistral 7B) locally on WSL2 for development and CI testing. In production, route to serverless OpenAI/Anthropic/DeepSeek endpoints with prompt caching to avoid upfront GPU reservation costs.
SME Strategy: EKS + Karpenter Spot GPUs + vLLM
Deploy self-hosted vLLM on AWS EKS using Karpenter Spot instances (`g5.4xlarge`, `g5.12xlarge`). Implement KEDA scaling based on average pending prompt queue depth. Saves up to 70% compared to on-demand instances.
Enterprise Strategy: Multi-Node KubeRay + Triton + H100 Fleet + OTel
Multi-cluster Kubernetes with NVIDIA GPU Operator, hardware MIG slicing, KubeRay distributed fine-tuning, Triton Inference Server with TensorRT-LLM, and end-to-end OpenTelemetry LLM tracing for token cost attribution across business units.
Troubleshooting & Production Postmortems 14.5
Cause: Concurrent requests with long prompts exceed the dynamic KV-cache allocation pool.
Fix: Set --gpu-memory-utilization 0.90 in vLLM to leave a 10% safety cushion. Enable prefix caching (--enable-prefix-caching) so identical system prompts share KV cache blocks.
Cause: AWS VPC network MTU mismatch or missing EFA (Elastic Fabric Adapter) kernel interfaces between worker pods.
Fix: Verify jumbo frames (MTU 9001) across VPC subnets and set NCCL_DEBUG=INFO NCCL_SOCKET_IFNAME=eth0 in container environment variables.
2026 LLMOps Career Roadmap & Interview Tips 14.6
Key 2026 Interview Soundbites
- Time-to-First-Token (TTFT): Governed by Prefill phase compute throughput.
- Time-per-Output-Token (TPOT): Governed by Decode phase memory bandwidth.
- PagedAttention: Virtual memory paging for GPU KV-cache blocks.
Next Step: The War Room
Ready to test your live incident triage skills? Head into the Staff/Principal War Room Simulator.
Enter War Room Simulator