Kubernetes Advanced Topics
Advanced Kubernetes patterns, resources, and architectures for production-grade cluster management.
Kubernetes Advanced Topics
Advanced Kubernetes patterns, resources, and architectures for production-grade cluster management.
Overview
This cheatsheet covers advanced Kubernetes concepts beyond basic pod and deployment management, including custom resources, operators, advanced scheduling, security policies, service meshes, and multi-cluster architectures. These topics are essential for building scalable, secure, and resilient production Kubernetes environments.
graph TB
subgraph "Advanced K8s Ecosystem"
CRD[Custom Resources<br/>& CRDs]
OP[Operators]
ADM[Admission<br/>Controllers]
SCHED[Advanced<br/>Scheduling]
SEC[Security<br/>Policies]
MESH[Service Mesh]
MC[Multi-Cluster]
CRD --> OP
ADM --> SEC
SCHED --> OP
MESH --> MC
end
subgraph "Core K8s"
API[API Server]
end
CRD --> API
ADM --> API
OP --> API
Custom Resource Definitions (CRDs)
CRDs extend the Kubernetes API to create custom resource types, allowing you to define domain-specific objects that behave like native Kubernetes resources.
Key Concepts
- CustomResourceDefinition: Defines the schema and validation for a new resource type
- Custom Resource (CR): Instance of a CRD
- API Group/Version: Organises CRDs into logical groups with versioning
- Structural Schema: OpenAPI v3 schema defining the resource structure
- Subresources: Optional status and scale subresources
flowchart LR
A[Define CRD] --> B[Apply CRD to Cluster]
B --> C[Create CR Instances]
C --> D[Controller Watches CRs]
D --> E[Reconcile Desired State]
E --> D
Creating a CRD
# crontab-crd.yaml
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: crontabs.stable.example.com
spec:
group: stable.example.com
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
cronSpec:
type: string
pattern: '^(\d+|\*)(/\d+)?(\s+(\d+|\*)(/\d+)?){4}$'
image:
type: string
replicas:
type: integer
minimum: 1
maximum: 10
required: ["cronSpec", "image"]
status:
type: object
properties:
lastScheduleTime:
type: string
format: date-time
subresources:
status: {}
scope: Namespaced
names:
plural: crontabs
singular: crontab
kind: CronTab
shortNames:
- ct
# Apply the CRD
kubectl apply -f crontab-crd.yaml
# Verify CRD creation
kubectl get crds
kubectl describe crd crontabs.stable.example.com
# Create a custom resource instance
cat <<EOF | kubectl apply -f -
apiVersion: stable.example.com/v1
kind: CronTab
metadata:
name: my-cron-job
spec:
cronSpec: "*/5 * * * *"
image: my-cron-image:v1
replicas: 3
EOF
# Interact with custom resources like native resources
kubectl get crontabs
kubectl get ct
kubectl describe crontab my-cron-job
kubectl delete crontab my-cron-job
Advanced CRD Features
# CRD with validation, defaults, and printer columns
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: databases.db.example.com
spec:
group: db.example.com
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
engine:
type: string
enum: ["postgresql", "mysql", "mongodb"]
default: "postgresql"
version:
type: string
storageSize:
type: string
pattern: '^\d+Gi$'
backupEnabled:
type: boolean
default: true
additionalPrinterColumns:
- name: Engine
type: string
jsonPath: .spec.engine
- name: Version
type: string
jsonPath: .spec.version
- name: Storage
type: string
jsonPath: .spec.storageSize
- name: Age
type: date
jsonPath: .metadata.creationTimestamp
scope: Namespaced
names:
plural: databases
singular: database
kind: Database
Operators and Operator Lifecycle Manager (OLM)
Operators are Kubernetes extensions that use custom resources to manage applications and their components, encoding operational knowledge in software.
Key Concepts
- Operator Pattern: Controller + CRD that manages complex applications
- Operator SDK: Framework for building operators (Go, Ansible, Helm)
- Operator Lifecycle Manager (OLM): Manages operator installation, updates, and dependencies
- OperatorHub: Repository of certified operators
- Reconciliation Loop: Continuous process ensuring desired state matches actual state
sequenceDiagram
participant User
participant CRD as Custom Resource
participant Operator
participant K8s as Kubernetes API
participant App as Application
User->>CRD: Create/Update CR
Operator->>CRD: Watch for changes
CRD-->>Operator: Event triggered
Operator->>K8s: Reconcile resources
K8s->>App: Create/Update pods, services, etc.
App-->>Operator: Status update
Operator->>CRD: Update CR status
Installing OLM
# Install OLM on cluster
curl -sL https://github.com/operator-framework/operator-lifecycle-manager/releases/download/v0.25.0/install.sh | bash -s v0.25.0
# Verify OLM installation
kubectl get pods -n olm
kubectl get crds | grep operators
# List available operators in OperatorHub
kubectl get packagemanifests -n olm
Installing an Operator via OLM
# subscription.yaml - Install Prometheus Operator
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: prometheus
namespace: operators
spec:
channel: beta
name: prometheus
source: operatorhubio-catalog
sourceNamespace: olm
# Create namespace and operator group
kubectl create ns operators
cat <<EOF | kubectl apply -f -
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
name: operator-group
namespace: operators
spec:
targetNamespaces:
- operators
EOF
# Install operator via subscription
kubectl apply -f subscription.yaml
# Check operator status
kubectl get csv -n operators
kubectl get subscription -n operators
kubectl get installplan -n operators
Simple Operator Example (Conceptual)
// Simplified operator reconciliation logic
func (r *DatabaseReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
// Fetch the Database instance
var database dbv1.Database
if err := r.Get(ctx, req.NamespacedName, &database); err != nil {
return ctrl.Result{}, client.IgnoreNotFound(err)
}
// Define desired state (e.g., StatefulSet for database)
desired := constructStatefulSet(&database)
// Get current state
var current appsv1.StatefulSet
err := r.Get(ctx, types.NamespacedName{Name: desired.Name, Namespace: desired.Namespace}, ¤t)
if err != nil && errors.IsNotFound(err) {
// Create if doesn't exist
return ctrl.Result{}, r.Create(ctx, desired)
} else if err != nil {
return ctrl.Result{}, err
}
// Update if current differs from desired
if !equality.Semantic.DeepEqual(current.Spec, desired.Spec) {
current.Spec = desired.Spec
return ctrl.Result{}, r.Update(ctx, ¤t)
}
// Update status
return ctrl.Result{}, r.updateStatus(ctx, &database)
}
Admission Controllers and Webhooks
Admission controllers intercept API requests before they're persisted to etcd, enabling validation, mutation, and policy enforcement.
Key Concepts
- Validating Webhooks: Reject requests that violate policies
- Mutating Webhooks: Modify requests before persistence (e.g., inject sidecars)
- Admission Controller Chain: Built-in controllers that process requests
- Webhook Configuration: Defines when and how webhooks are invoked
- Failure Policy: Determines behaviour when webhook is unavailable
flowchart LR
A[kubectl apply] --> B[API Server]
B --> C[Authentication]
C --> D[Authorisation]
D --> E[Mutating Admission]
E --> F[Object Schema Validation]
F --> G[Validating Admission]
G --> H{All Checks Pass?}
H -->|Yes| I[Persist to etcd]
H -->|No| J[Reject Request]
Enabling Admission Controllers
# Check enabled admission controllers
kubectl exec -it kube-apiserver-master -n kube-system -- kube-apiserver -h | grep enable-admission-plugins
# Edit kube-apiserver manifest to enable controllers
# /etc/kubernetes/manifests/kube-apiserver.yaml
--enable-admission-plugins=NodeRestriction,MutatingAdmissionWebhook,ValidatingAdmissionWebhook,PodSecurity
Mutating Webhook Example (Sidecar Injection)
# mutating-webhook.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: MutatingWebhookConfiguration
metadata:
name: sidecar-injector
webhooks:
- name: sidecar-injector.example.com
clientConfig:
service:
name: sidecar-injector
namespace: default
path: "/mutate"
caBundle: LS0tLS1CRUdJTi... # Base64 encoded CA cert
rules:
- operations: ["CREATE"]
apiGroups: [""]
apiVersions: ["v1"]
resources: ["pods"]
admissionReviewVersions: ["v1", "v1beta1"]
sideEffects: None
timeoutSeconds: 5
reinvocationPolicy: Never
failurePolicy: Ignore # or Fail
namespaceSelector:
matchLabels:
sidecar-injection: enabled
Validating Webhook Example (Resource Quota)
# validating-webhook.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
name: resource-validator
webhooks:
- name: validate.resources.example.com
clientConfig:
service:
name: resource-validator
namespace: validation
path: "/validate"
caBundle: LS0tLS1CRUdJTi...
rules:
- operations: ["CREATE", "UPDATE"]
apiGroups: ["apps"]
apiVersions: ["v1"]
resources: ["deployments"]
admissionReviewVersions: ["v1"]
sideEffects: None
failurePolicy: Fail
matchPolicy: Equivalent
# Apply webhook configurations
kubectl apply -f mutating-webhook.yaml
kubectl apply -f validating-webhook.yaml
# Test webhook by creating a pod in labeled namespace
kubectl label namespace default sidecar-injection=enabled
kubectl run test --image=nginx
# Check webhook rejections
kubectl describe validatingwebhookconfiguration resource-validator
Advanced Scheduling
Control pod placement using node affinity, pod affinity/anti-affinity, taints, and tolerations.
Key Concepts
- Node Affinity: Schedule pods to specific nodes based on labels
- Pod Affinity: Co-locate pods based on labels
- Pod Anti-Affinity: Spread pods across nodes/zones
- Taints: Mark nodes to repel pods
- Tolerations: Allow pods to schedule on tainted nodes
- Priority Classes: Define pod scheduling priority
graph TB
subgraph "Scheduling Decision"
POD[Pod to Schedule]
POD --> FILT[Filter Nodes]
FILT --> AFF[Apply Affinity Rules]
AFF --> TAINT[Check Taints/Tolerations]
TAINT --> SCORE[Score Nodes]
SCORE --> SELECT[Select Best Node]
end
subgraph "Node Pool"
N1[Node 1<br/>zone=us-east-1a<br/>type=compute]
N2[Node 2<br/>zone=us-east-1b<br/>type=memory]
N3[Node 3<br/>zone=us-east-1a<br/>type=gpu<br/>taint=gpu]
end
SELECT -.-> N1
SELECT -.-> N2
SELECT -.-> N3
Node Affinity
# Node affinity example
apiVersion: v1
kind: Pod
metadata:
name: with-node-affinity
spec:
affinity:
nodeAffinity:
# Hard requirement - must match
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/arch
operator: In
values:
- amd64
- arm64
- key: node-type
operator: NotIn
values:
- spot
# Soft preference - prefers but not required
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 80
preference:
matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values:
- us-east-1a
- weight: 20
preference:
matchExpressions:
- key: instance-type
operator: In
values:
- c5.xlarge
containers:
- name: app
image: nginx
Pod Affinity and Anti-Affinity
# Pod affinity and anti-affinity
apiVersion: apps/v1
kind: Deployment
metadata:
name: web-cache
spec:
replicas: 3
selector:
matchLabels:
app: web-cache
template:
metadata:
labels:
app: web-cache
spec:
affinity:
# Co-locate with web servers
podAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: app
operator: In
values:
- web-server
topologyKey: kubernetes.io/hostname
# Spread across availability zones
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchExpressions:
- key: app
operator: In
values:
- web-cache
topologyKey: topology.kubernetes.io/zone
containers:
- name: redis
image: redis:7
Taints and Tolerations
# Add taint to node
kubectl taint nodes node1 gpu=true:NoSchedule
kubectl taint nodes node2 maintenance=true:NoExecute
kubectl taint nodes node3 environment=production:PreferNoSchedule
# Remove taint
kubectl taint nodes node1 gpu=true:NoSchedule-
# View node taints
kubectl describe node node1 | grep Taints
# Pod with tolerations
apiVersion: v1
kind: Pod
metadata:
name: gpu-workload
spec:
tolerations:
# Exact match toleration
- key: "gpu"
operator: "Equal"
value: "true"
effect: "NoSchedule"
# Tolerate any value for key
- key: "maintenance"
operator: "Exists"
effect: "NoExecute"
tolerationSeconds: 3600 # Evict after 1 hour
# Wildcard toleration (tolerates all taints)
- operator: "Exists"
containers:
- name: tensorflow
image: tensorflow/tensorflow:latest-gpu
Priority Classes
# priority-class.yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: high-priority
value: 1000000
globalDefault: false
description: "High priority for critical services"
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: low-priority
value: 100
globalDefault: false
description: "Low priority for batch jobs"
# Use priority class in pod
apiVersion: v1
kind: Pod
metadata:
name: critical-app
spec:
priorityClassName: high-priority
containers:
- name: app
image: critical-app:v1
# List priority classes
kubectl get priorityclasses
# Preemption in action (high priority pod evicts low priority)
kubectl describe pod critical-app | grep -A 5 Events
StatefulSets and DaemonSets
Specialised workload controllers for stateful applications and node-level services.
StatefulSets
StatefulSets manage stateful applications requiring stable network identities, persistent storage, and ordered deployment/scaling.
graph LR
subgraph StatefulSet
A[Pod: web-0<br/>PVC: data-web-0]
B[Pod: web-1<br/>PVC: data-web-1]
C[Pod: web-2<br/>PVC: data-web-2]
end
subgraph Headless Service
SVC[web.default.svc]
end
A --> SVC
B --> SVC
C --> SVC
A -.DNS.-> DNS1[web-0.web.default.svc]
B -.DNS.-> DNS2[web-1.web.default.svc]
C -.DNS.-> DNS3[web-2.web.default.svc]
# statefulset.yaml - MySQL cluster
apiVersion: v1
kind: Service
metadata:
name: mysql
labels:
app: mysql
spec:
ports:
- port: 3306
name: mysql
clusterIP: None # Headless service
selector:
app: mysql
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: mysql
spec:
serviceName: mysql
replicas: 3
selector:
matchLabels:
app: mysql
template:
metadata:
labels:
app: mysql
spec:
containers:
- name: mysql
image: mysql:8.0
ports:
- containerPort: 3306
name: mysql
env:
- name: MYSQL_ROOT_PASSWORD
valueFrom:
secretKeyRef:
name: mysql-secret
key: password
volumeMounts:
- name: data
mountPath: /var/lib/mysql
- name: xtrabackup
image: gcr.io/google-samples/xtrabackup:1.0
ports:
- containerPort: 3307
name: xtrabackup
volumeMounts:
- name: data
mountPath: /var/lib/mysql
- name: conf
mountPath: /etc/mysql/conf.d
volumes:
- name: conf
configMap:
name: mysql-config
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: fast-ssd
resources:
requests:
storage: 10Gi
# StatefulSet operations
kubectl apply -f statefulset.yaml
# Scale StatefulSet (scales sequentially)
kubectl scale statefulset mysql --replicas=5
# Ordered rollout
kubectl set image statefulset/mysql mysql=mysql:8.0.32
# Check rollout status
kubectl rollout status statefulset/mysql
# Delete StatefulSet (keeps PVCs)
kubectl delete statefulset mysql --cascade=orphan
# Force delete a stuck pod
kubectl delete pod mysql-2 --force --grace-period=0
DaemonSets
DaemonSets ensure a pod runs on all (or selected) nodes, typically for node-level services like monitoring, logging, or networking.
# daemonset.yaml - Node monitoring agent
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-exporter
namespace: monitoring
spec:
selector:
matchLabels:
app: node-exporter
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
template:
metadata:
labels:
app: node-exporter
spec:
hostNetwork: true
hostPID: true
tolerations:
# Run on all nodes including masters
- operator: Exists
effect: NoSchedule
containers:
- name: node-exporter
image: prom/node-exporter:v1.6.0
args:
- --path.sysfs=/host/sys
- --path.rootfs=/host/root
- --collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)
ports:
- containerPort: 9100
protocol: TCP
name: metrics
volumeMounts:
- name: sys
mountPath: /host/sys
readOnly: true
- name: root
mountPath: /host/root
readOnly: true
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 200m
memory: 256Mi
volumes:
- name: sys
hostPath:
path: /sys
- name: root
hostPath:
path: /
# DaemonSet operations
kubectl apply -f daemonset.yaml
# Check DaemonSet status (should match node count)
kubectl get daemonset -n monitoring
kubectl get pods -l app=node-exporter -o wide
# Update DaemonSet image
kubectl set image daemonset/node-exporter node-exporter=prom/node-exporter:v1.7.0 -n monitoring
# Run DaemonSet on specific nodes only
kubectl label nodes node1 monitoring=enabled
# Then add nodeSelector to DaemonSet spec
Cluster Autoscaling
Automatically adjust cluster capacity based on resource demands.
Key Concepts
- Cluster Autoscaler (CA): Adds/removes nodes based on pending pods
- Vertical Pod Autoscaler (VPA): Adjusts pod resource requests/limits
- Horizontal Pod Autoscaler (HPA): Scales pod replicas based on metrics
- Node Groups/Pools: Logical groupings of nodes for scaling
flowchart TD
A[Metrics Server] --> B[HPA]
A --> C[VPA]
B --> D{CPU/Memory<br/>Threshold?}
D -->|High| E[Scale Out Pods]
D -->|Low| F[Scale In Pods]
E --> G{Nodes<br/>Available?}
G -->|No| H[Cluster Autoscaler]
H --> I[Add Nodes]
F --> J{Nodes<br/>Underutilised?}
J -->|Yes| H
H --> K[Remove Nodes]
C --> L[Adjust Pod<br/>Requests/Limits]
Cluster Autoscaler Setup (AWS)
# cluster-autoscaler.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: cluster-autoscaler
namespace: kube-system
spec:
replicas: 1
selector:
matchLabels:
app: cluster-autoscaler
template:
metadata:
labels:
app: cluster-autoscaler
spec:
serviceAccountName: cluster-autoscaler
containers:
- image: registry.k8s.io/autoscaling/cluster-autoscaler:v1.28.0
name: cluster-autoscaler
command:
- ./cluster-autoscaler
- --v=4
- --stderrthreshold=info
- --cloud-provider=aws
- --skip-nodes-with-local-storage=false
- --expander=least-waste
- --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/my-cluster
- --balance-similar-node-groups
- --skip-nodes-with-system-pods=false
- --scale-down-enabled=true
- --scale-down-delay-after-add=10m
- --scale-down-unneeded-time=10m
resources:
requests:
cpu: 100m
memory: 300Mi
limits:
cpu: 100m
memory: 300Mi
# Check cluster autoscaler logs
kubectl logs -f deployment/cluster-autoscaler -n kube-system
# View autoscaler status
kubectl describe configmap cluster-autoscaler-status -n kube-system
# Trigger scale-up by creating resource-intensive pods
kubectl create deployment test --image=nginx --replicas=100
Vertical Pod Autoscaler (VPA)
# Install VPA
git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
# vpa.yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: nginx-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: nginx
updatePolicy:
updateMode: "Auto" # Off, Initial, Recreate, Auto
resourcePolicy:
containerPolicies:
- containerName: nginx
minAllowed:
cpu: 100m
memory: 128Mi
maxAllowed:
cpu: 2
memory: 2Gi
controlledResources: ["cpu", "memory"]
# Check VPA recommendations
kubectl describe vpa nginx-vpa
kubectl get vpa nginx-vpa -o jsonpath='{.status.recommendation}'
Custom Metrics and HPA
Scale applications based on custom or external metrics beyond CPU/memory.
Metrics Architecture
graph TB
subgraph "Metrics Pipeline"
APP[Application] -->|Expose| PROM[Prometheus]
PROM --> ADAPTER[Prometheus Adapter]
ADAPTER --> CUSTOM[Custom Metrics API]
end
subgraph "Autoscaling"
HPA[HorizontalPodAutoscaler]
HPA -->|Query| CUSTOM
HPA -->|Query| EXTERNAL[External Metrics API]
HPA -->|Scale| DEPLOY[Deployment]
end
METRICS[Metrics Server] -->|Resource Metrics| HPA
Standard HPA (CPU/Memory)
# hpa-basic.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: nginx-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: nginx
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 50
periodSeconds: 60
- type: Pods
value: 2
periodSeconds: 60
selectPolicy: Min
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 30
- type: Pods
value: 4
periodSeconds: 30
selectPolicy: Max
Custom Metrics HPA
# Install Prometheus Adapter first
# helm install prometheus-adapter prometheus-community/prometheus-adapter
# hpa-custom.yaml - Scale based on HTTP request rate
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-app
minReplicas: 3
maxReplicas: 20
metrics:
# Custom metric from Prometheus
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: "1000"
# Object metric
- type: Object
object:
metric:
name: queue_depth
describedObject:
apiVersion: v1
kind: Service
name: rabbitmq
target:
type: Value
value: "30"
# Check HPA status
kubectl get hpa
kubectl describe hpa web-hpa
# View HPA events
kubectl get events --field-selector involvedObject.name=web-hpa
# Test autoscaling with load
kubectl run -i --tty load-generator --rm --image=busybox --restart=Never -- /bin/sh -c "while sleep 0.01; do wget -q -O- http://web-app; done"
External Metrics Example
# hpa-external.yaml - Scale based on SQS queue length
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: sqs-consumer
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: message-processor
minReplicas: 1
maxReplicas: 50
metrics:
- type: External
external:
metric:
name: sqs_queue_messages_visible
selector:
matchLabels:
queue: "production-queue"
target:
type: AverageValue
averageValue: "30" # Target 30 messages per pod
Network Policies
Control network traffic between pods and external endpoints using firewall-like rules.
Key Concepts
- Network Policy: Defines allowed ingress/egress traffic for pods
- Pod Selector: Identifies pods the policy applies to
- Policy Types: Ingress, Egress, or both
- CNI Plugin: Must support NetworkPolicy (Calico, Cilium, Weave)
graph LR
subgraph "Namespace: default"
FE[Frontend Pods<br/>role=frontend]
BE[Backend Pods<br/>role=backend]
DB[Database Pods<br/>role=database]
end
EXT[External Client] -->|Port 80| FE
FE -->|Port 8080| BE
BE -->|Port 5432| DB
FE -.X.-> DB
EXT -.X.-> BE
EXT -.X.-> DB
style FE fill:#90EE90
style BE fill:#87CEEB
style DB fill:#FFB6C1
Default Deny All Traffic
# deny-all.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production
spec:
podSelector: {} # Applies to all pods in namespace
policyTypes:
- Ingress
- Egress
Allow Specific Traffic
# backend-policy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: backend-policy
namespace: production
spec:
podSelector:
matchLabels:
role: backend
policyTypes:
- Ingress
- Egress
ingress:
# Allow from frontend pods
- from:
- podSelector:
matchLabels:
role: frontend
ports:
- protocol: TCP
port: 8080
# Allow from monitoring namespace
- from:
- namespaceSelector:
matchLabels:
name: monitoring
ports:
- protocol: TCP
port: 9090
egress:
# Allow to database
- to:
- podSelector:
matchLabels:
role: database
ports:
- protocol: TCP
port: 5432
# Allow DNS
- to:
- namespaceSelector:
matchLabels:
name: kube-system
ports:
- protocol: UDP
port: 53
# Allow external API calls
- to:
- ipBlock:
cidr: 0.0.0.0/0
except:
- 10.0.0.0/8
- 172.16.0.0/12
- 192.168.0.0/16
ports:
- protocol: TCP
port: 443
Complex Network Policy Example
# microservices-policy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: api-gateway-policy
spec:
podSelector:
matchLabels:
app: api-gateway
policyTypes:
- Ingress
- Egress
ingress:
# Allow from ingress controller
- from:
- namespaceSelector:
matchLabels:
name: ingress-nginx
- podSelector:
matchLabels:
app: nginx-ingress
ports:
- protocol: TCP
port: 8000
egress:
# Allow to specific microservices
- to:
- podSelector:
matchExpressions:
- key: app
operator: In
values:
- user-service
- product-service
- order-service
ports:
- protocol: TCP
port: 8080
# Allow to external auth provider
- to:
- ipBlock:
cidr: 203.0.113.0/24
ports:
- protocol: TCP
port: 443
# Apply network policies
kubectl apply -f deny-all.yaml
kubectl apply -f backend-policy.yaml
# Test network policy
kubectl run test-pod --rm -it --image=busybox -- wget -O- http://backend-service:8080
# View network policies
kubectl get networkpolicies
kubectl describe networkpolicy backend-policy
# Check if CNI supports NetworkPolicy
kubectl get nodes -o jsonpath='{.items[*].status.nodeInfo.containerRuntimeVersion}'
Pod Security Policies and OPA/Gatekeeper
Enforce security standards and compliance policies across the cluster.
Pod Security Standards (PSS)
Pod Security Standards replaced Pod Security Policies (PSPs) in Kubernetes 1.25+.
graph TB
A[Pod Security Standards] --> B[Privileged]
A --> C[Baseline]
A --> D[Restricted]
B --> B1[Unrestricted<br/>For trusted workloads]
C --> C1[Minimal restrictions<br/>Prevents known privilege escalations]
D --> D1[Heavily restricted<br/>Hardened security best practices]
Pod Security Admission
# namespace with pod security labels
apiVersion: v1
kind: Namespace
metadata:
name: production
labels:
# Enforce restricted standard
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: latest
# Warn on baseline violations
pod-security.kubernetes.io/warn: baseline
pod-security.kubernetes.io/warn-version: latest
# Audit privileged violations
pod-security.kubernetes.io/audit: privileged
pod-security.kubernetes.io/audit-version: latest
# Example: Pod that meets restricted standard
apiVersion: v1
kind: Pod
metadata:
name: secure-pod
namespace: production
spec:
securityContext:
runAsNonRoot: true
runAsUser: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: app
image: nginx:1.25
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
readOnlyRootFilesystem: true
volumeMounts:
- name: cache
mountPath: /var/cache/nginx
- name: run
mountPath: /var/run
volumes:
- name: cache
emptyDir: {}
- name: run
emptyDir: {}
OPA Gatekeeper
Gatekeeper is a policy controller using Open Policy Agent (OPA) to enforce custom policies.
# Install Gatekeeper
kubectl apply -f https://raw.githubusercontent.com/open-policy-agent/gatekeeper/master/deploy/gatekeeper.yaml
# Verify installation
kubectl get pods -n gatekeeper-system
kubectl get crds | grep gatekeeper
Constraint Template (Policy Definition)
# constraint-template.yaml - Require labels
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8srequiredlabels
spec:
crd:
spec:
names:
kind: K8sRequiredLabels
validation:
openAPIV3Schema:
type: object
properties:
labels:
type: array
items:
type: string
targets:
- target: admission.k8s.gatekeeper.sh
rego: |
package k8srequiredlabels
violation[{"msg": msg, "details": {"missing_labels": missing}}] {
provided := {label | input.review.object.metadata.labels[label]}
required := {label | label := input.parameters.labels[_]}
missing := required - provided
count(missing) > 0
msg := sprintf("You must provide labels: %v", [missing])
}
Constraint (Policy Instance)
# constraint.yaml - Enforce labels on namespaces
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: namespace-must-have-owner
spec:
match:
kinds:
- apiGroups: [""]
kinds: ["Namespace"]
parameters:
labels:
- "owner"
- "environment"
Common Gatekeeper Policies
# container-limits.yaml - Require resource limits
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8scontainerlimits
spec:
crd:
spec:
names:
kind: K8sContainerLimits
targets:
- target: admission.k8s.gatekeeper.sh
rego: |
package k8scontainerlimits
violation[{"msg": msg}] {
container := input.review.object.spec.containers[_]
not container.resources.limits.cpu
msg := sprintf("Container %v has no CPU limit", [container.name])
}
violation[{"msg": msg}] {
container := input.review.object.spec.containers[_]
not container.resources.limits.memory
msg := sprintf("Container %v has no memory limit", [container.name])
}
---
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sContainerLimits
metadata:
name: must-have-limits
spec:
match:
kinds:
- apiGroups: ["apps"]
kinds: ["Deployment", "StatefulSet", "DaemonSet"]
# allowed-repos.yaml - Restrict image registries
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8sallowedrepos
spec:
crd:
spec:
names:
kind: K8sAllowedRepos
validation:
openAPIV3Schema:
type: object
properties:
repos:
type: array
items:
type: string
targets:
- target: admission.k8s.gatekeeper.sh
rego: |
package k8sallowedrepos
violation[{"msg": msg}] {
container := input.review.object.spec.containers[_]
satisfied := [good | repo = input.parameters.repos[_] ; good = startswith(container.image, repo)]
not any(satisfied)
msg := sprintf("Container %v uses disallowed registry: %v", [container.name, container.image])
}
---
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sAllowedRepos
metadata:
name: allowed-registries
spec:
match:
kinds:
- apiGroups: [""]
kinds: ["Pod"]
- apiGroups: ["apps"]
kinds: ["Deployment", "StatefulSet"]
parameters:
repos:
- "gcr.io/my-company/"
- "docker.io/library/"
- "registry.k8s.io/"
# Check policy violations
kubectl get constraints
kubectl describe k8srequiredlabels namespace-must-have-owner
# Test policy
kubectl create namespace test # Should fail without labels
# A namespace WITH the required labels is admitted:
kubectl create namespace test --dry-run=client -o yaml \
| kubectl label --local -f - owner=platform environment=dev -o yaml \
| kubectl apply -f -
# View audit violations
kubectl get constraints -o json | jq '.items[].status.violations'
Service Meshes (Istio, Linkerd)
Service meshes provide observability, traffic management, and security for microservices without changing application code.
Service Mesh Architecture
graph TB
subgraph "Control Plane"
CP[Control Plane<br/>istiod/linkerd-controller]
end
subgraph "Data Plane"
subgraph Pod1[Pod: Service A]
APP1[App Container]
PROXY1[Sidecar Proxy<br/>Envoy/Linkerd2-proxy]
end
subgraph Pod2[Pod: Service B]
APP2[App Container]
PROXY2[Sidecar Proxy]
end
subgraph Pod3[Pod: Service C]
APP3[App Container]
PROXY3[Sidecar Proxy]
end
end
CP -.Config.-> PROXY1
CP -.Config.-> PROXY2
CP -.Config.-> PROXY3
APP1 --> PROXY1
PROXY1 -->|mTLS| PROXY2
APP2 --> PROXY2
PROXY2 -->|mTLS| PROXY3
APP3 --> PROXY3
Istio Installation
# Download Istio
curl -L https://istio.io/downloadIstio | sh -
cd istio-1.20.0
export PATH=$PWD/bin:$PATH
# Install Istio with demo profile
istioctl install --set profile=demo -y
# Enable sidecar injection for namespace
kubectl label namespace default istio-injection=enabled
# Verify installation
kubectl get pods -n istio-system
istioctl verify-install
Istio Traffic Management
# virtual-service.yaml - Canary deployment
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: reviews
spec:
hosts:
- reviews
http:
- match:
- headers:
user-agent:
regex: ".*Mobile.*"
route:
- destination:
host: reviews
subset: v2
weight: 100
- route:
- destination:
host: reviews
subset: v1
weight: 90
- destination:
host: reviews
subset: v2
weight: 10
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: reviews
spec:
host: reviews
trafficPolicy:
connectionPool:
tcp:
maxConnections: 100
http:
http1MaxPendingRequests: 10
http2MaxRequests: 100
maxRequestsPerConnection: 2
outlierDetection:
consecutiveErrors: 5
interval: 30s
baseEjectionTime: 30s
maxEjectionPercent: 50
subsets:
- name: v1
labels:
version: v1
- name: v2
labels:
version: v2
trafficPolicy:
loadBalancer:
simple: ROUND_ROBIN
Istio Security (mTLS)
# peer-authentication.yaml - Enforce mTLS
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default
namespace: production
spec:
mtls:
mode: STRICT # PERMISSIVE, STRICT, DISABLE
---
# authorization-policy.yaml
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata:
name: frontend-policy
namespace: production
spec:
selector:
matchLabels:
app: frontend
action: ALLOW
rules:
- from:
- source:
principals:
- "cluster.local/ns/production/sa/api-gateway"
to:
- operation:
methods: ["GET", "POST"]
paths: ["/api/*"]
when:
- key: request.headers[x-api-key]
values: ["*"]
Istio Circuit Breaking and Retries
# circuit-breaker.yaml
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: backend-circuit-breaker
spec:
host: backend.production.svc.cluster.local
trafficPolicy:
connectionPool:
tcp:
maxConnections: 100
http:
http1MaxPendingRequests: 10
maxRequestsPerConnection: 2
outlierDetection:
consecutive5xxErrors: 5
interval: 5s
baseEjectionTime: 30s
maxEjectionPercent: 50
minHealthPercent: 40
---
# retry-policy.yaml
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: backend-retries
spec:
hosts:
- backend
http:
- route:
- destination:
host: backend
retries:
attempts: 3
perTryTimeout: 2s
retryOn: 5xx,reset,connect-failure,refused-stream
timeout: 10s
Linkerd Installation
# Install Linkerd CLI
curl -fsL https://run.linkerd.io/install | sh
export PATH=$PATH:$HOME/.linkerd2/bin
# Validate cluster
linkerd check --pre
# Install Linkerd
linkerd install --crds | kubectl apply -f -
linkerd install | kubectl apply -f -
# Verify installation
linkerd check
# Enable auto-injection for namespace
kubectl annotate namespace default linkerd.io/inject=enabled
# Manually inject sidecar
kubectl get deploy -o yaml | linkerd inject - | kubectl apply -f -
Linkerd Traffic Split (SMI)
# traffic-split.yaml - Canary with SMI
apiVersion: split.smi-spec.io/v1alpha2
kind: TrafficSplit
metadata:
name: backend-split
spec:
service: backend
backends:
- service: backend-v1
weight: 900
- service: backend-v2
weight: 100
---
# Service for each version
apiVersion: v1
kind: Service
metadata:
name: backend-v1
spec:
selector:
app: backend
version: v1
ports:
- port: 8080
---
apiVersion: v1
kind: Service
metadata:
name: backend-v2
spec:
selector:
app: backend
version: v2
ports:
- port: 8080
Service Mesh Observability
# Istio - Access Kiali dashboard
istioctl dashboard kiali
# Istio - View metrics in Prometheus
istioctl dashboard prometheus
# Istio - Distributed tracing with Jaeger
istioctl dashboard jaeger
# Linkerd - Dashboard
linkerd viz dashboard
# Linkerd - View service metrics
linkerd viz stat deployments
linkerd viz top deployments
linkerd viz tap deployment/backend
Hybrid and Multi-Cluster Setups
Manage applications across multiple clusters or hybrid cloud/on-premises environments.
Multi-Cluster Patterns
graph TB
subgraph "Region 1"
C1[Cluster 1<br/>Primary]
C2[Cluster 2<br/>DR]
end
subgraph "Region 2"
C3[Cluster 3<br/>Replica]
end
subgraph "Control Plane"
MC[Multi-Cluster<br/>Controller<br/>KubeFed/ArgoCD]
end
MC -->|Sync| C1
MC -->|Sync| C2
MC -->|Sync| C3
C1 -.Replicate.-> C2
C1 -.Replicate.-> C3
KubeFed (Kubernetes Federation)
KubeFed is archived/retired (moved to
kubernetes-retired/kubefed, read-only since 2023). For new multi-cluster work prefer Argo CD ApplicationSets, Cluster API, or a fleet manager (Rancher/Karmada). The commands below are kept for existing deployments only.
# Install KubeFed (archived project)
kubectl create ns kube-federation-system
helm repo add kubefed-charts https://raw.githubusercontent.com/kubernetes-sigs/kubefed/master/charts
helm install kubefed kubefed-charts/kubefed --namespace kube-federation-system
# Join clusters to federation
kubefedctl join cluster1 --cluster-context cluster1 --host-cluster-context cluster1
kubefedctl join cluster2 --cluster-context cluster2 --host-cluster-context cluster1
# Verify joined clusters
kubectl get kubefedclusters -n kube-federation-system
# federated-deployment.yaml
apiVersion: types.kubefed.io/v1beta1
kind: FederatedDeployment
metadata:
name: test-deployment
namespace: default
spec:
template:
metadata:
labels:
app: nginx
spec:
replicas: 3
selector:
matchLabels:
app: nginx
template:
metadata:
labels:
app: nginx
spec:
containers:
- name: nginx
image: nginx:1.25
placement:
clusters:
- name: cluster1
- name: cluster2
overrides:
- clusterName: cluster2
clusterOverrides:
- path: "/spec/replicas"
value: 5
Multi-Cluster Service Discovery (Istio)
# Install Istio on both clusters with multi-cluster support
# Cluster 1 (primary)
istioctl install --set profile=default \
--set values.global.meshID=mesh1 \
--set values.global.multiCluster.clusterName=cluster1 \
--set values.global.network=network1
# Create remote secret for cluster2
istioctl create-remote-secret --name=cluster2 | kubectl apply -f -
# Cluster 2 (remote)
istioctl install --set profile=default \
--set values.global.meshID=mesh1 \
--set values.global.multiCluster.clusterName=cluster2 \
--set values.global.network=network2
# Enable cross-cluster service discovery
kubectl apply -f - <<EOF
apiVersion: networking.istio.io/v1
kind: Gateway
metadata:
name: cross-network-gateway
namespace: istio-system
spec:
selector:
istio: eastwestgateway
servers:
- port:
number: 15443
name: tls
protocol: TLS
tls:
mode: AUTO_PASSTHROUGH
hosts:
- "*.local"
EOF
Cluster API (ClusterAPI)
# Install clusterctl
curl -L https://github.com/kubernetes-sigs/cluster-api/releases/download/v1.6.0/clusterctl-linux-amd64 -o clusterctl
chmod +x clusterctl
sudo mv clusterctl /usr/local/bin/
# Initialise management cluster
clusterctl init --infrastructure aws
# Create workload cluster
clusterctl generate cluster workload-cluster \
--kubernetes-version v1.28.0 \
--control-plane-machine-count=3 \
--worker-machine-count=5 \
> workload-cluster.yaml
kubectl apply -f workload-cluster.yaml
# Get kubeconfig for new cluster
clusterctl get kubeconfig workload-cluster > workload-cluster.kubeconfig
GitOps Multi-Cluster (ArgoCD)
# argocd-cluster-secret.yaml
apiVersion: v1
kind: Secret
metadata:
name: cluster2-secret
namespace: argocd
labels:
argocd.argoproj.io/secret-type: cluster
type: Opaque
stringData:
name: cluster2
server: https://cluster2.example.com
config: |
{
"bearerToken": "<token>",
"tlsClientConfig": {
"insecure": false,
"caData": "<base64-ca-cert>"
}
}
# multi-cluster-app.yaml
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: multi-cluster-app
namespace: argocd
spec:
generators:
- clusters:
selector:
matchLabels:
environment: production
template:
metadata:
name: '{{name}}-guestbook'
spec:
project: default
source:
repoURL: https://github.com/example/manifests
targetRevision: HEAD
path: guestbook
destination:
server: '{{server}}'
namespace: guestbook
syncPolicy:
automated:
prune: true
selfHeal: true
Quick Reference
CRD Operations
# List CRDs
kubectl get crds
# Describe CRD
kubectl explain crontabs.spec
# Get custom resources
kubectl get crontabs
kubectl get crontabs.stable.example.com
# Delete CRD (deletes all instances)
kubectl delete crd crontabs.stable.example.com
Scheduling Commands
# Label nodes
kubectl label nodes node1 disktype=ssd
kubectl label nodes node1 zone=us-east-1a
# Taint nodes
kubectl taint nodes node1 key=value:NoSchedule
kubectl taint nodes node1 key=value:NoExecute
kubectl taint nodes node1 key- # Remove taint
# View node details
kubectl describe node node1 | grep -A 5 Taints
kubectl get nodes --show-labels
Autoscaling Commands
# HPA
kubectl autoscale deployment nginx --cpu-percent=70 --min=2 --max=10
kubectl get hpa
kubectl describe hpa nginx
# VPA
kubectl get vpa
kubectl describe vpa nginx-vpa
# Cluster Autoscaler
kubectl logs -f deployment/cluster-autoscaler -n kube-system
Network Policy Commands
# List network policies
kubectl get networkpolicies
kubectl get netpol
# Describe policy
kubectl describe networkpolicy backend-policy
# Test connectivity
kubectl run test --rm -it --image=busybox -- wget -O- http://service:8080
Service Mesh Commands
# Istio
istioctl version
istioctl proxy-status
istioctl analyze
kubectl label namespace default istio-injection=enabled
# Linkerd
linkerd version
linkerd check
linkerd viz stat deployments
linkerd viz tap deployment/backend
kubectl annotate namespace default linkerd.io/inject=enabled
Common Issues and Solutions
CRD Issues
| Issue | Solution |
|---|---|
| CRD not found after creation | Wait for API discovery cache refresh (up to 60s) or restart API server |
| Schema validation failing | Check openAPIV3Schema matches CR fields exactly |
| Status subresource not updating | Ensure subresources.status: {} in CRD spec and use /status endpoint |
| Version conflict | Use conversion webhooks or set single storage: true version |
Scheduling Issues
| Issue | Solution |
|---|---|
| Pod pending despite resources | Check node selectors, taints, and PodTopologySpread constraints |
| Affinity rules too restrictive | Use preferredDuringScheduling instead of required for soft constraints |
| Taints preventing scheduling | Add tolerations to pod spec matching node taints |
| Priority preemption not working | Ensure PriorityClass exists and value differences are significant (>100) |
Autoscaling Issues
| Issue | Solution |
|---|---|
| HPA not scaling | Verify metrics-server is running and pod has resource requests defined |
| Cluster Autoscaler not adding nodes | Check node group min/max settings and cloud provider IAM permissions |
| VPA conflicts with HPA | Don't use VPA and HPA on same metric (use VPA for requests, HPA for replicas) |
| Flapping (rapid scale up/down) | Increase stabilizationWindowSeconds in HPA behavior config |
Network Policy Issues
| Issue | Solution |
|---|---|
| Policy not enforced | Verify CNI plugin supports NetworkPolicy (Calico, Cilium, Weave) |
| DNS resolution failing | Allow egress to kube-dns/coredns on port 53 UDP |
| Can't reach external IPs | Add egress rule with ipBlock for external CIDR ranges |
| Policy applied but traffic still allowed | Check for multiple policies - allow rules are additive |
Service Mesh Issues
| Issue | Solution |
|---|---|
| Sidecar not injected | Verify namespace label and check mutating webhook configuration |
| mTLS connection failures | Ensure PeerAuthentication mode matches (PERMISSIVE during migration) |
| High latency after mesh install | Adjust resource limits for sidecars, check circuit breaker settings |
| Certificate expiration | Istio auto-rotates certs, check istio-ca-secret and citadel logs |
Multi-Cluster Issues
| Issue | Solution |
|---|---|
| Cross-cluster service discovery fails | Verify east-west gateway and DNS configuration |
| Federation sync delays | Check KubeFed controller logs and network latency between clusters |
| Cluster join failing | Ensure correct kubeconfig context and network connectivity |
| Resource conflicts across clusters | Use cluster-specific overrides in FederatedResource |