Kubernetes — Container Orchestration at Scale
From Pod scheduling internals and eBPF networking to RBAC hardening, HPA, and production multi-node clusters. Learn on Kind locally, graduate to EKS on AWS.
Under-The-Hood Architecture & Networking 01
Control Plane → Data Plane Flow
When you run kubectl apply -f deployment.yaml, here's the exact journey your manifest takes through the Kubernetes control plane before a Pod lands on a Node.
- API Server validation: kubectl sends a POST to
/apis/apps/v1/namespaces/default/deployments. API server authenticates (x509/OIDC), runs admission webhooks (mutating → validating), validates the schema. - etcd persistence: API server writes the Deployment object to etcd. etcd uses Raft consensus — the write is committed when a quorum of etcd members acknowledges it.
- Deployment Controller: watches for new Deployment objects via API server's watch stream, creates ReplicaSet objects to match
spec.replicas. - ReplicaSet Controller: creates Pod objects in
Pendingstate. Pods have nonodeNameassigned yet. - Scheduler: watches for Pending Pods. Runs Filtering (which nodes can run this pod?) and Scoring (which node is best?) phases. Writes the chosen
nodeNameto the Pod object via a binding. - kubelet on target node: watches for Pods assigned to its node. Calls the container runtime (containerd → runc) to pull the image and start containers. Updates Pod status back to API server.
Felix agent programs IP routing tables on the host. In VXLAN mode, packets are encapsulated in UDP (port 4789) adding ~50 bytes of header, reducing effective MTU to ~1450. Increases CPU due to kernel-space encap/decap.| Mode | Lookup Complexity | Max Services | Load Balancing Algorithms |
|---|---|---|---|
| IPTables | O(N) — sequential rules | ~5,000 before degradation | Random, Round-Robin only |
| IPVS | O(1) — hash table lookups | 100,000+ | LeastConn, SourceHash, DNAT, WRR |
CoreDNS uses a pipeline: health → errors → cache → kubernetes → forward.
- Standard Services: Returns a single A record → ClusterIP → iptables/IPVS routes to a pod.
- Headless Services (
clusterIP: None): Returns A records directly mapping to all backing Pod IPs. Used for StatefulSets (Cassandra, Kafka, Elasticsearch). - SRV records: Headless services also emit SRV records for named ports — essential for service discovery protocols like gRPC.
Control Plane Architecture Diagram
Core Components & Objects 02
| Object | Purpose | Key Fields | Best Practice |
|---|---|---|---|
| Pod | Smallest deployable unit. Shares network namespace and volumes. | containers, volumes, nodeSelector | Never run Pods directly — use Deployments. |
| Deployment | Declarative Pod management with rolling updates and rollbacks. | replicas, strategy, selector | Always set maxSurge and maxUnavailable for zero-downtime. |
| StatefulSet | Ordered Pod deployment with stable network identity and persistent storage. | serviceName, volumeClaimTemplates | Use for databases, Kafka, Elasticsearch. |
| Service | Stable virtual IP (ClusterIP) for Pod discovery and load balancing. | type, selector, ports | Use LoadBalancer only when cloud provider provides one; prefer Ingress. |
| HPA | Horizontal Pod Autoscaler — scales replicas based on CPU/memory/custom metrics. | scaleTargetRef, minReplicas, maxReplicas, metrics | Combine with KEDA for event-driven scaling (queue depth, etc.). |
| PVC/PV | Persistent Volume Claim/Volume for durable storage beyond Pod lifecycle. | storageClassName, accessModes, capacity | Use dynamic provisioning (StorageClass) — avoid static PVs in production. |
| NetworkPolicy | L3/L4 firewall rules for Pod-to-Pod communication (requires CNI support). | podSelector, policyTypes, ingress, egress | Default-deny all, then allow only required paths explicitly. |
| RBAC | Role-Based Access Control for API resource permissions. | Role, ClusterRole, RoleBinding, ClusterRoleBinding | Principle of least privilege. Never use verbs: ["*"] in production. |
WSL Hands-On Lab 03
wsl-setup.sh first. This lab requires: kubectl, kind, helm, k9s. Estimated time: 60–90 minutes.Step 1 — Create a Multi-Node Kind Cluster
apiVersion: kind.x-k8s.io/v1alpha4 kind: Cluster nodes: - role: control-plane extraPortMappings: - containerPort: 30080 hostPort: 8080 protocol: TCP - role: worker - role: worker
# Create the cluster (takes ~2 min) $ kind create cluster --config kind-cluster.yaml --name devops-lab Creating cluster "devops-lab" ... ✓ Ensuring node image (kindest/node:v1.29.2) ✓ Preparing nodes 📦 📦 📦 ✓ Writing configuration 📜 ✓ Starting control-plane 🕹️ ✓ Installing CNI 🔌 ✓ Joining worker nodes 🤝 Set kubectl context to "kind-devops-lab" # Verify the cluster $ kubectl get nodes NAME STATUS ROLES AGE VERSION devops-lab-control-plane Ready control-plane 2m v1.29.2 devops-lab-worker Ready <none> 90s v1.29.2 devops-lab-worker2 Ready <none> 90s v1.29.2 # Launch k9s TUI for visual cluster management $ k9s
Step 2 — Deploy a 3-Tier Application
apiVersion: v1 kind: Namespace metadata: name: three-tier-app labels: pod-security.kubernetes.io/enforce: restricted --- apiVersion: v1 kind: ServiceAccount metadata: name: backend-sa namespace: three-tier-app automountServiceAccountToken: false --- apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: backend-reader namespace: three-tier-app rules: - apiGroups: [""] resources: ["configmaps"] verbs: ["get", "list"]
apiVersion: apps/v1 kind: Deployment metadata: name: backend namespace: three-tier-app spec: replicas: 2 strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 maxUnavailable: 0 # Zero-downtime rolling update selector: matchLabels: app: backend template: metadata: labels: app: backend spec: serviceAccountName: backend-sa securityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 2000 containers: - name: backend image: nginx:alpine ports: - containerPort: 8080 securityContext: allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: drop: ["ALL"] resources: requests: { cpu: "100m", memory: "128Mi" } limits: { cpu: "500m", memory: "256Mi" } readinessProbe: httpGet: { path: /healthz, port: 8080 } initialDelaySeconds: 5 livenessProbe: httpGet: { path: /healthz, port: 8080 } initialDelaySeconds: 15
apiVersion: v1 kind: Service metadata: name: backend-svc namespace: three-tier-app spec: selector: app: backend ports: - port: 80 targetPort: 8080 --- apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: backend-ingress namespace: three-tier-app annotations: nginx.ingress.kubernetes.io/rewrite-target: / spec: ingressClassName: nginx rules: - host: app.local http: paths: - path: / pathType: Prefix backend: service: name: backend-svc port: { number: 80 }
apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: backend-hpa namespace: three-tier-app spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: backend minReplicas: 2 maxReplicas: 10 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70 - type: Resource resource: name: memory target: type: AverageValue averageValue: 200Mi
# Apply all manifests $ kubectl apply -f namespace-rbac.yaml backend-deployment.yaml service-ingress.yaml hpa.yaml # Watch pods come up $ kubectl get pods -n three-tier-app -w # Check HPA status $ kubectl get hpa -n three-tier-app # Inspect a pod's security context $ kubectl get pod backend-xxxx -n three-tier-app -o jsonpath='{.spec.securityContext}' # Test rollout: update image and watch rolling update $ kubectl set image deployment/backend backend=nginx:latest -n three-tier-app $ kubectl rollout status deployment/backend -n three-tier-app # Rollback if needed $ kubectl rollout undo deployment/backend -n three-tier-app
Real-World Project: 3-Tier App 04
🌱 Startup: Local Kind Cluster
You're a solo developer deploying a React frontend + Node.js API + PostgreSQL database. Total infrastructure cost: $0 (Kind on WSL2).
# Deploy PostgreSQL with a simple StatefulSet $ kubectl create secret generic pg-secret \ --from-literal=POSTGRES_PASSWORD=mypassword \ -n three-tier-app # Deploy using helm (bitnami postgresql chart) $ helm repo add bitnami https://charts.bitnami.com/bitnami $ helm install postgres bitnami/postgresql \ --set auth.existingSecret=pg-secret \ -n three-tier-app # Port-forward to test locally $ kubectl port-forward svc/backend-svc 8080:80 -n three-tier-app
🏢 SME: EKS on AWS with Terraform
You have 20 engineers. Infrastructure needs to be automated, reliable, and multi-AZ. Uses eksctl for cluster creation and Helm for application deployment.
# Create EKS cluster (takes ~15 min) $ eksctl create cluster \ --name sme-production \ --region us-east-1 \ --nodegroup-name workers \ --node-type t3.medium \ --nodes-min 2 \ --nodes-max 10 \ --managed # Install AWS Load Balancer Controller $ helm install aws-load-balancer-controller \ eks/aws-load-balancer-controller \ --set clusterName=sme-production \ -n kube-system # Enable IRSA for pod-level AWS permissions $ eksctl create iamserviceaccount \ --cluster sme-production \ --namespace three-tier-app \ --name backend-sa \ --attach-policy-arn arn:aws:iam::aws:policy/AmazonS3ReadOnlyAccess \ --approve
🏛 Enterprise: Multi-Cluster + GitOps
500+ engineers across 3 regions. Each team owns a namespace. Infrastructure as Code (Terraform), GitOps (ArgoCD), policy enforcement (OPA Gatekeeper), and multi-cluster federation.
Troubleshooting & Disaster Recovery 05
# Step 1: Describe the pod — look for Events section $ kubectl describe pod <pod-name> -n <namespace> # Common causes from Events: # "0/3 nodes are available: 3 Insufficient cpu" → resource requests too high # "no nodes matched node affinity" → nodeSelector mismatch # "did not match pod anti-affinity rules" → topology constraint violation # "PVC not bound" → storage class issue # Step 2: Check node capacity $ kubectl describe nodes | grep -A 5 "Allocated resources" # Step 3: Check pending PVCs $ kubectl get pvc -n <namespace>
# Check rollout status $ kubectl rollout status deployment/<name> -n <ns> --timeout=60s # View ReplicaSets to see old vs new $ kubectl get rs -n <ns> # Check pod logs of new pods (likely failing readinessProbe) $ kubectl logs <new-pod-name> --previous # Emergency rollback $ kubectl rollout undo deployment/<name> -n <ns>
# Check if Service has endpoints (if NONE, label selector is wrong) $ kubectl get endpoints <service-name> -n <ns> # Run a debug pod to test connectivity from inside cluster $ kubectl run debug --image=nicolaka/netshoot --rm -it -- bash debug# curl http://<service-name>.<namespace>.svc.cluster.local debug# nslookup <service-name>.<namespace> # Check if NetworkPolicy is blocking traffic $ kubectl get networkpolicies -n <ns>
Interactive Lab Checklist 06
- Create a 3-node Kind cluster and verify all nodes are ReadyLab ↗
- Deploy a Deployment with 3 replicas and verify rolling update worksLab ↗
- Create a Role allowing only pod listing in a specific namespace, bind to a ServiceAccountLab ↗
- Apply a NetworkPolicy that isolates a namespace (deny all ingress/egress by default)Lab ↗
- Configure HPA on a deployment and generate load to trigger auto-scalingLab ↗
- Debug a broken service: fix label selector mismatch causing empty EndpointsSadServers ↗
30-Day Learning Roadmap 07
- Pod, ReplicaSet, Deployment lifecycle
- Services: ClusterIP, NodePort, LoadBalancer
- ConfigMaps & Secrets
- kubectl cheat sheet mastery
- PV/PVC/StorageClass dynamic provisioning
- StatefulSets for databases
- Node affinity & taints/tolerations
- Topology Spread Constraints
- RBAC: Roles, ClusterRoles, Bindings
- NetworkPolicies (Calico/Cilium)
- Pod Security Standards (restricted)
- Ingress + TLS cert-manager
- HPA + KEDA event-driven scaling
- PodDisruptionBudgets
- Multi-cluster management
- EKS on AWS with eksctl