💰FinOps & Cloud Cost Engineering
Stop guessing and start governing. Master Kubecost per-namespace chargeback, Infracost PR cost gates, Karpenter Spot automation, right-sizing with VPA, Reserved Instance strategy, and real enterprise techniques that save $100K+/month. Every command, every query, every decision a senior engineer makes.
Real Incident: The $54,000 Monthly Surprise
Context: E-commerce platform, 200-node EKS cluster. AWS bill arrives: $54K over budget. CFO calls VP Engineering at 9pm.
Root Cause (discovered in 3 hours): A data engineering team had launched 40× r5.4xlarge On-Demand nodes for a Spark job — and forgotten to terminate them 6 weeks ago. No tagging, no namespace quotas, no budget alerts.
Blast Radius: $54,200 wasted. 6-week billing cycle. Team only noticed when CFO escalated.
# How we found it in 10 minutes with Kubecost CLI $ kubectl cost namespace --window 30d --show-cpu --show-memory --show-efficiency +-------------------+----------+-------+--------+------------+ | Namespace | CPU Cost | Mem | Total | Efficiency | +-------------------+----------+-------+--------+------------+ | spark-jobs | $41,200 | $9,800| $51,000| 3.2% | ← HERE | production | $12,400 | $3,100| $15,500| 67.4% | | staging | $1,200 | $400| $1,600| 54.1% | +-------------------+----------+-------+--------+------------+ $ kubectl get nodes -l spark=true --no-headers | wc -l 40 $ kubectl get nodes -l spark=true -o jsonpath='{.items[0].metadata.creationTimestamp}' 2024-12-03T14:22:00Z ← 6 weeks ago
Prevention implemented: ResourceQuota per namespace, Kubecost Budget alerts at 80% threshold, Karpenter TTL on Spark NodePool, mandatory cost-center tags enforced via OPA Gatekeeper.
FinOps Architecture & Toolchain 15.1
Enterprise FinOps operates across three layers: Visibility (what are we spending?), Optimization (how do we reduce waste?), and Governance (how do we prevent overspend?). The toolchain maps directly to these layers.
Daily billing breakdown] Kubecost[Kubecost
Per namespace/pod cost] Cloudwatch[CloudWatch
Resource utilization] end subgraph "OPTIMIZATION" Karpenter[Karpenter
Spot + consolidation] VPA[VPA
Right-size containers] Infracost[Infracost
Pre-deploy cost estimate] RightSizing[AWS Compute
Optimizer + Graviton] end subgraph "GOVERNANCE" Budgets[AWS Budgets
Alert + SNS → Slack] OPA[OPA Gatekeeper
Enforce cost labels] ResourceQuota[K8s ResourceQuota
Namespace spend caps] RISP[RI/SP Automation
Commitment purchasing] end CostExplorer --> Kubecost Kubecost --> Karpenter Kubecost --> Budgets Infracost --> ResourceQuota
Visibility First
You cannot optimize what you cannot see. Kubecost gives you per-Pod, per-Deployment, per-Namespace cost breakdown with CPU/Memory/Network/Storage split. Enable cost allocation tags in AWS before anything else.
Governance = Prevention
OPA policies that block untagged resources, K8s ResourceQuotas per team namespace, and Infracost PR gates are the only way to prevent the next $54K surprise. Reactive alerting alone is not enough.
Kubecost: Per-Namespace Chargeback 15.2
Kubecost is the de-facto standard for Kubernetes cost allocation. It runs as a Deployment in your cluster, scrapes Prometheus metrics, and maps CPU/memory requests to node pricing from cloud APIs.
# Install Kubecost on EKS with Spot node support $ helm repo add kubecost https://kubecost.github.io/cost-analyzer/ $ helm repo update $ helm install kubecost kubecost/cost-analyzer \ --namespace kubecost --create-namespace \ --set kubecostToken="YOUR_TOKEN" \ --set global.prometheus.enabled=true \ --set global.grafana.enabled=true \ --set clusterName="prod-eks-us-east-1" \ --set serviceAccount.annotations."eks\.amazonaws\.com/role-arn"=arn:aws:iam::123456789:role/kubecost-sa NAME: kubecost STATUS: deployed NOTES: Visit http://localhost:9090 after port-forwarding # Port-forward to access UI $ kubectl port-forward svc/kubecost-cost-analyzer 9090:9090 -n kubecost # Install kubectl-cost plugin $ kubectl krew install cost
# Query via kubectl-cost (CLI) — daily cost per namespace
kubectl cost namespace \
--window 7d \
--show-cpu --show-memory --show-network --show-efficiency \
--show-pv
# Query via Kubecost HTTP API — last 30 days, by deployment
curl "http://localhost:9090/model/allocation?window=30d&aggregate=deployment&accumulate=false" \
| jq '.data[0] | to_entries | sort_by(-.value.totalCost) | .[:10]'
# Get cost per label (e.g., team=payments)
kubectl cost label --label team=payments --window 30d
# Show idle (wasted) resources per namespace
kubectl cost namespace --show-efficiency --window 30d \
| awk 'NR>2 && $NF < 0.30 {print "LOW EFFICIENCY: "$0}'apiVersion: v1
kind: ConfigMap
metadata:
name: kubecost-alerts
namespace: kubecost
data:
alerts.json: |
{
"alerts": [
{
"type": "budget",
"threshold": 5000,
"window": "30d",
"aggregation": "namespace",
"filter": "namespace=payments",
"slackWebhookUrl": "https://hooks.slack.com/YOUR_WEBHOOK",
"alertNames": ["payments-namespace-budget"]
},
{
"type": "efficiency",
"efficiencyThreshold": 0.25,
"window": "7d",
"aggregation": "deployment",
"slackWebhookUrl": "https://hooks.slack.com/YOUR_WEBHOOK"
}
]
}#!/bin/bash
# Monthly chargeback report — send to finance team
KUBECOST_URL="http://kubecost-cost-analyzer.kubecost:9090"
MONTH=$(date -d "last month" +%Y-%m)
START="${MONTH}-01T00:00:00Z"
END=$(date -d "${MONTH}-01 +1 month -1 second" +%Y-%m-%dT%H:%M:%SZ)
curl -s "${KUBECOST_URL}/model/allocation?window=${START},${END}&aggregate=namespace&accumulate=true" \
| jq -r '
.data[0] | to_entries[] |
{
namespace: .key,
team: (.value.properties.labels.team // "untagged"),
cost_center: (.value.properties.labels."cost-center" // "UNKNOWN"),
total_usd: (.value.totalCost | . * 100 | round / 100),
cpu_usd: (.value.cpuCost | . * 100 | round / 100),
memory_usd: (.value.ramCost | . * 100 | round / 100),
efficiency_pct: ((.value.totalEfficiency // 0) * 100 | round)
}' | \
jq -r '[.namespace, .team, .cost_center, .total_usd, .efficiency_pct] | @csv' \
> /tmp/chargeback-${MONTH}.csv
echo "Chargeback report: /tmp/chargeback-${MONTH}.csv"
# Upload to S3 for Finance
aws s3 cp /tmp/chargeback-${MONTH}.csv s3://finops-reports/chargeback/- Install Kubecost on kind/EKS and access the UI dashboard
- Run
kubectl cost namespace --window 7dand identify the highest-cost namespace - Configure a Kubecost budget alert to fire when namespace cost exceeds $1000/month
- Generate and export a monthly chargeback CSV report
Infracost: Pre-Deploy Cost Gates in CI 15.3
Infracost parses Terraform plans and calculates the monthly cost delta before any infrastructure is deployed. Integrate it as a PR gate — block merges that increase cost by more than a threshold.
name: Infracost Cost Check
on:
pull_request:
paths:
- '**.tf'
- '**.tfvars'
jobs:
infracost:
name: Estimate infrastructure cost
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
# Generate plan for baseline (base branch)
- name: Setup Infracost
uses: infracost/actions/setup@v3
with:
api-key: ${{ secrets.INFRACOST_API_KEY }}
- name: Checkout base branch
uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.base.ref }}
path: base
- name: Generate Infracost baseline
run: |
infracost breakdown --path=base/terraform \
--format=json \
--out-file=/tmp/infracost-base.json
- name: Generate Infracost for PR
run: |
infracost breakdown --path=terraform \
--format=json \
--out-file=/tmp/infracost-pr.json
- name: Generate diff and post PR comment
run: |
infracost diff \
--path=/tmp/infracost-pr.json \
--compare-to=/tmp/infracost-base.json \
--format=json \
--out-file=/tmp/infracost-diff.json
infracost comment github \
--path=/tmp/infracost-diff.json \
--repo=$GITHUB_REPOSITORY \
--pull-request=${{ github.event.pull_request.number }} \
--github-token=${{ github.token }} \
--behavior=update
- name: Enforce cost policy
run: |
infracost comment github \
--path=/tmp/infracost-diff.json \
--policy-path=infracost-policy.rego \
--github-token=${{ github.token }} \
--repo=$GITHUB_REPOSITORY \
--pull-request=${{ github.event.pull_request.number }}version: 0.1
# Custom price overrides for reserved instances
# or private pricing agreements
projects:
- path: terraform/
name: prod-eks-cluster
terraform_vars:
environment: prod
region: us-east-1
# Exclude resources from cost check (e.g., tagging-only changes)
exclude_resources:
- aws_iam_role
- aws_iam_policy
- aws_s3_bucket_policypackage infracost
# Block PRs that increase monthly cost by more than $500
deny[msg] {
monthlyCostDelta := input.diffTotalMonthlyCost
to_number(monthlyCostDelta) > 500
msg := sprintf(
"PR increases monthly cost by $%v (limit: $500). Requires FinOps team approval.",
[monthlyCostDelta]
)
}
# Warn on any new On-Demand instance larger than r5.2xlarge
warn[msg] {
resource := input.projects[_].breakdownDiff.resources[_]
resource.resourceType == "aws_instance"
contains(resource.name, "r5.4xlarge")
resource.monthlyCost > 0
msg := sprintf(
"New large On-Demand instance %v may be a candidate for Spot or Reserved Instance.",
[resource.name]
)
}Karpenter: Intelligent Spot Automation 15.4
Karpenter replaces the Cluster Autoscaler with a faster, smarter node provisioner. It provisions the exact right instance type for your workload — using Spot first, falling back to On-Demand, and consolidating idle nodes automatically.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general-spot
spec:
template:
metadata:
labels:
billing-team: platform
spec:
nodeClassRef:
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
name: default
requirements:
# Prefer Spot, fall back to On-Demand automatically
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
# Allow multiple instance families for better Spot availability
- key: node.kubernetes.io/instance-type
operator: In
values:
- m5.xlarge
- m5.2xlarge
- m5a.xlarge
- m5a.2xlarge
- m6i.xlarge
- m6i.2xlarge
- m6a.xlarge
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"] # Include Graviton for ~20% savings
- key: topology.kubernetes.io/zone
operator: In
values: ["us-east-1a", "us-east-1b", "us-east-1c"]
limits:
cpu: "1000" # Max 1000 vCPUs in this NodePool
memory: 4000Gi
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 60s
budgets:
- nodes: "10%" # Never disrupt more than 10% of nodes at once
---
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
name: default
spec:
amiFamily: AL2023
role: "KarpenterNodeRole-prod"
subnetSelectorTerms:
- tags:
karpenter.sh/discovery: prod-eks-cluster
securityGroupSelectorTerms:
- tags:
karpenter.sh/discovery: prod-eks-cluster
blockDeviceMappings:
- deviceName: /dev/xvda
ebs:
volumeSize: 100Gi
volumeType: gp3
iops: 3000
throughput: 125
tags:
Environment: prod
ManagedBy: karpenter
CostCenter: platform# NodePool for Spark/batch jobs with auto-TTL
# Nodes self-terminate after job completion
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: spark-batch
spec:
template:
metadata:
labels:
billing-team: data-engineering
workload-type: batch
spec:
nodeClassRef:
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
name: compute-optimized
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot"] # Batch is always Spot
- key: node.kubernetes.io/instance-type
operator: In
values: ["c5.4xlarge", "c5.8xlarge", "c5a.4xlarge", "c6i.4xlarge"]
taints:
- key: workload-type
value: batch
effect: NoSchedule
limits:
cpu: "500"
disruption:
consolidationPolicy: WhenEmpty
consolidateAfter: 30s # Aggressive consolidation for batch nodes
expireAfter: 24h # Hard TTL — nodes die after 24h max# Karpenter v1 Disruption Budget — protect critical workloads
# during Spot consolidation
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general-spot
spec:
disruption:
budgets:
# During business hours (9-5 UTC): very conservative
- schedule: "0 9 * * mon-fri"
duration: 8h
nodes: "5%"
# Off-hours: aggressive consolidation for savings
- schedule: "0 17 * * mon-fri"
duration: 16h
nodes: "20%"
# Weekends: most aggressive
- schedule: "0 0 * * sat"
duration: 48h
nodes: "30%"
# Also add PodDisruptionBudget for critical services
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payments-api-pdb
namespace: production
spec:
minAvailable: "80%"
selector:
matchLabels:
app: payments-apiTypical Spot Savings
- → On-Demand
m5.2xlarge: ~$276/month - → Spot
m5.2xlarge: ~$83/month (70% savings) - → 100-node cluster: $19,300/month saved
- → + Graviton arm64 additional 20%: $23,000/month saved
Spot Interruption Handling
AWS Spot Interruption Handler (helm chart) watches for 2-minute interruption notices and gracefully cordons/drains the node before reclamation. Use karpenter.sh/interruption-queue SQS integration for zero-data-loss workloads.
Right-Sizing: VPA + AWS Compute Optimizer 15.5
# VPA in Recommendation-only mode (safe for production)
# Never auto-apply without testing Off mode first
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: payments-api-vpa
namespace: production
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: payments-api
updatePolicy:
updateMode: "Off" # "Off" = recommend only; "Auto" = apply + evict
resourcePolicy:
containerPolicies:
- containerName: "*"
minAllowed:
cpu: 50m
memory: 64Mi
maxAllowed:
cpu: 4
memory: 8Gi
controlledResources: ["cpu", "memory"]
---
# Query VPA recommendations
# kubectl get vpa -n production -o json | jq '.items[].status.recommendation'# List EC2 right-sizing recommendations $ aws compute-optimizer get-ec2-instance-recommendations \ --region us-east-1 \ --filters name=Finding,values=OVER_PROVISIONED \ --query 'instanceRecommendations[*].{ Instance:instanceArn, Current:currentInstanceType, Recommended:recommendationOptions[0].instanceType, Savings:recommendationOptions[0].estimatedMonthlySavings.value }' --output table +------------------+-------------+--------------+---------+ | Instance | Current | Recommended | Savings | +------------------+-------------+--------------+---------+ | i-0a1b2c3d4e | m5.4xlarge | m5.2xlarge | $142.50 | | i-0f1e2d3c4b | r5.2xlarge | r5.xlarge | $98.00 | +------------------+-------------+--------------+---------+ # Get RDS right-sizing recommendations $ aws compute-optimizer get-rds-database-recommendations \ --region us-east-1 \ --filters name=Finding,values=OVER_PROVISIONED # Export all recommendations to S3 for monthly review $ aws compute-optimizer export-ec2-instance-recommendations \ --s3-destination-config bucket=finops-reports,keyPrefix=optimizer/
package kubernetes.admission
# OPA Gatekeeper ConstraintTemplate — enforce cost labels on all pods
# Required labels: team, cost-center, environment
deny[msg] {
input.request.kind.kind == "Pod"
pod := input.request.object
required_labels := {"team", "cost-center", "environment"}
provided := {label | pod.metadata.labels[label]}
missing := required_labels - provided
count(missing) > 0
msg := sprintf(
"Pod missing required cost labels: %v. All pods must have team, cost-center, environment.",
[missing]
)
}
# Block pods without resource limits (will cause OOM and waste)
deny[msg] {
input.request.kind.kind == "Pod"
container := input.request.object.spec.containers[_]
not container.resources.limits
msg := sprintf(
"Container '%v' missing resource limits. Set cpu and memory limits to prevent cost overruns.",
[container.name]
)
}RI/SP Strategy: Commitment Purchasing 15.6
When to Buy Reserved Instances
- Baseline compute stable for >6 months → 1yr Standard RI (40% savings)
- DB instances (RDS/ElastiCache) → Reserved DB Instance
- Multi-AZ NAT Gateways → Savings Plans cover egress
- Use Convertible RIs when instance family may change
- Never commit 100% — keep 20-30% On-Demand for spikes
Compute Savings Plans (Preferred)
- Applies to EC2 and Fargate and Lambda
- Flexible across region, instance family, OS
- Up to 66% discount vs On-Demand
- 1-year No-Upfront: best for cash-constrained teams
- 3-year All-Upfront: maximum savings if budget allows
# Analyze your Savings Plan coverage and utilization $ aws ce get-savings-plans-coverage \ --time-period Start=2024-01-01,End=2024-02-01 \ --granularity MONTHLY \ --filter '{"Dimensions":{"Key":"SERVICE","Values":["Amazon EC2"]}}' \ --query 'SavingsPlansCoverages[*].Coverage' # Purchase recommendation: how much to commit $ aws ce get-savings-plans-purchase-recommendation \ --savings-plans-type COMPUTE_SP \ --term-in-years ONE_YEAR \ --payment-option NO_UPFRONT \ --lookback-period-in-days SIXTY_DAYS { "RecommendedSpend": "47.23", ← commit $47.23/hr = ~$34K/month "EstimatedSavings": "$12,450", ← monthly savings vs On-Demand "CurrentCoverage": "34.2%" ← only 34% currently covered }
Hands-On Enterprise Labs 15.7
FinOps Mastery Labs
- Deploy Kubecost on kind cluster, generate cost report for 3 sample namespaces
- Write Infracost policy that blocks any PR adding an instance > $300/month
- Configure Karpenter NodePool with Spot + arm64 Graviton support, verify consolidation
- Apply VPA in recommendation mode to a busy deployment, read its suggestions
- Create AWS Budget with SNS → Slack alert for 80% threshold breach
- Write OPA Gatekeeper constraint requiring cost-center label on all Pods
- Run AWS Cost Anomaly Detection, analyze and explain the anomaly report
- Calculate break-even point for 1-year Compute Savings Plan vs On-Demand for a given workload
Troubleshooting & Gotchas 15.8
Kubecost uses list price by default. Enable AWS Cost and Usage Reports (CUR) integration in Kubecost settings to pull actual billing data including discounts, EDP, and reserved instance rates. CUR integration requires an S3 bucket with Athena queries.
Check for daemon sets (can't evict), PodDisruptionBudgets with minAvailable=1 blocking eviction, or karpenter.sh/do-not-disrupt: "true" annotations on pods. Increase consolidateAfter window. Check disruption budget percentages.
Set karpenter.sh/capacity-type: spot only on non-critical workloads. Use node selectors and taints to separate Spot-tolerant workloads. Deploy AWS Node Termination Handler (helm chart). Ensure PodDisruptionBudgets exist for all critical Deployments. Set terminationGracePeriodSeconds: 120 to handle 2-min Spot warning.
The baseline must be generated from the target branch state, not the current PR. Always checkout the base ref first, run terraform init, then generate baseline JSON. Ensure both runs use the same --terraform-var-file flags.