Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Kubernetes Advanced Topics

Advanced Kubernetes patterns, resources, and architectures for production-grade cluster management.

Kubernetes Advanced Topics

Advanced Kubernetes patterns, resources, and architectures for production-grade cluster management.

Overview

This cheatsheet covers advanced Kubernetes concepts beyond basic pod and deployment management, including custom resources, operators, advanced scheduling, security policies, service meshes, and multi-cluster architectures. These topics are essential for building scalable, secure, and resilient production Kubernetes environments.

Core K8sAdvanced K8s EcosystemCustom Resources& CRDsOperatorsAdmissionControllersAdvancedSchedulingSecurityPoliciesService MeshMulti-ClusterAPI ServerCore K8sAdvanced K8s EcosystemCustom Resources& CRDsOperatorsAdmissionControllersAdvancedSchedulingSecurityPoliciesService MeshMulti-ClusterAPI Server

Custom Resource Definitions (CRDs)

CRDs extend the Kubernetes API to create custom resource types, allowing you to define domain-specific objects that behave like native Kubernetes resources.

Key Concepts

  • CustomResourceDefinition: Defines the schema and validation for a new resource type
  • Custom Resource (CR): Instance of a CRD
  • API Group/Version: Organises CRDs into logical groups with versioning
  • Structural Schema: OpenAPI v3 schema defining the resource structure
  • Subresources: Optional status and scale subresources
Define CRDApply CRD to ClusterCreate CR InstancesController WatchesCRsReconcile DesiredStateDefine CRDApply CRD to ClusterCreate CR InstancesController WatchesCRsReconcile DesiredState

Creating a CRD

# crontab-crd.yaml
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: crontabs.stable.example.com
spec:
  group: stable.example.com
  versions:
    - name: v1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                cronSpec:
                  type: string
                  pattern: '^(\d+|\*)(/\d+)?(\s+(\d+|\*)(/\d+)?){4}$'
                image:
                  type: string
                replicas:
                  type: integer
                  minimum: 1
                  maximum: 10
              required: ["cronSpec", "image"]
            status:
              type: object
              properties:
                lastScheduleTime:
                  type: string
                  format: date-time
      subresources:
        status: {}
  scope: Namespaced
  names:
    plural: crontabs
    singular: crontab
    kind: CronTab
    shortNames:
    - ct
# Apply the CRD
kubectl apply -f crontab-crd.yaml

# Verify CRD creation
kubectl get crds
kubectl describe crd crontabs.stable.example.com

# Create a custom resource instance
cat <<EOF | kubectl apply -f -
apiVersion: stable.example.com/v1
kind: CronTab
metadata:
  name: my-cron-job
spec:
  cronSpec: "*/5 * * * *"
  image: my-cron-image:v1
  replicas: 3
EOF

# Interact with custom resources like native resources
kubectl get crontabs
kubectl get ct
kubectl describe crontab my-cron-job
kubectl delete crontab my-cron-job

Advanced CRD Features

# CRD with validation, defaults, and printer columns
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: databases.db.example.com
spec:
  group: db.example.com
  versions:
    - name: v1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                engine:
                  type: string
                  enum: ["postgresql", "mysql", "mongodb"]
                  default: "postgresql"
                version:
                  type: string
                storageSize:
                  type: string
                  pattern: '^\d+Gi$'
                backupEnabled:
                  type: boolean
                  default: true
      additionalPrinterColumns:
      - name: Engine
        type: string
        jsonPath: .spec.engine
      - name: Version
        type: string
        jsonPath: .spec.version
      - name: Storage
        type: string
        jsonPath: .spec.storageSize
      - name: Age
        type: date
        jsonPath: .metadata.creationTimestamp
  scope: Namespaced
  names:
    plural: databases
    singular: database
    kind: Database

Operators and Operator Lifecycle Manager (OLM)

Operators are Kubernetes extensions that use custom resources to manage applications and their components, encoding operational knowledge in software.

Key Concepts

  • Operator Pattern: Controller + CRD that manages complex applications
  • Operator SDK: Framework for building operators (Go, Ansible, Helm)
  • Operator Lifecycle Manager (OLM): Manages operator installation, updates, and dependencies
  • OperatorHub: Repository of certified operators
  • Reconciliation Loop: Continuous process ensuring desired state matches actual state
ApplicationKubernetes APIOperatorCustom ResourceUserApplicationKubernetes APIOperatorCustom ResourceUserCreate/Update CRWatch for changesEvent triggeredReconcile resourcesCreate/Update pods, services, etc.Status updateUpdate CR statusApplicationKubernetes APIOperatorCustom ResourceUserApplicationKubernetes APIOperatorCustom ResourceUserCreate/Update CRWatch for changesEvent triggeredReconcile resourcesCreate/Update pods, services, etc.Status updateUpdate CR status

Installing OLM

# Install OLM on cluster
curl -sL https://github.com/operator-framework/operator-lifecycle-manager/releases/download/v0.25.0/install.sh | bash -s v0.25.0

# Verify OLM installation
kubectl get pods -n olm
kubectl get crds | grep operators

# List available operators in OperatorHub
kubectl get packagemanifests -n olm

Installing an Operator via OLM

# subscription.yaml - Install Prometheus Operator
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
  name: prometheus
  namespace: operators
spec:
  channel: beta
  name: prometheus
  source: operatorhubio-catalog
  sourceNamespace: olm
# Create namespace and operator group
kubectl create ns operators

cat <<EOF | kubectl apply -f -
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
  name: operator-group
  namespace: operators
spec:
  targetNamespaces:
  - operators
EOF

# Install operator via subscription
kubectl apply -f subscription.yaml

# Check operator status
kubectl get csv -n operators
kubectl get subscription -n operators
kubectl get installplan -n operators

Simple Operator Example (Conceptual)

// Simplified operator reconciliation logic
func (r *DatabaseReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
    // Fetch the Database instance
    var database dbv1.Database
    if err := r.Get(ctx, req.NamespacedName, &database); err != nil {
        return ctrl.Result{}, client.IgnoreNotFound(err)
    }

    // Define desired state (e.g., StatefulSet for database)
    desired := constructStatefulSet(&database)

    // Get current state
    var current appsv1.StatefulSet
    err := r.Get(ctx, types.NamespacedName{Name: desired.Name, Namespace: desired.Namespace}, &current)

    if err != nil && errors.IsNotFound(err) {
        // Create if doesn't exist
        return ctrl.Result{}, r.Create(ctx, desired)
    } else if err != nil {
        return ctrl.Result{}, err
    }

    // Update if current differs from desired
    if !equality.Semantic.DeepEqual(current.Spec, desired.Spec) {
        current.Spec = desired.Spec
        return ctrl.Result{}, r.Update(ctx, &current)
    }

    // Update status
    return ctrl.Result{}, r.updateStatus(ctx, &database)
}

Admission Controllers and Webhooks

Admission controllers intercept API requests before they're persisted to etcd, enabling validation, mutation, and policy enforcement.

Key Concepts

  • Validating Webhooks: Reject requests that violate policies
  • Mutating Webhooks: Modify requests before persistence (e.g., inject sidecars)
  • Admission Controller Chain: Built-in controllers that process requests
  • Webhook Configuration: Defines when and how webhooks are invoked
  • Failure Policy: Determines behaviour when webhook is unavailable
YesNokubectl applyAPI ServerAuthenticationAuthorisationMutating AdmissionObject SchemaValidationValidating AdmissionAll Checks Pass?Persist to etcdReject RequestYesNokubectl applyAPI ServerAuthenticationAuthorisationMutating AdmissionObject SchemaValidationValidating AdmissionAll Checks Pass?Persist to etcdReject Request

Enabling Admission Controllers

# Check enabled admission controllers
kubectl exec -it kube-apiserver-master -n kube-system -- kube-apiserver -h | grep enable-admission-plugins

# Edit kube-apiserver manifest to enable controllers
# /etc/kubernetes/manifests/kube-apiserver.yaml
--enable-admission-plugins=NodeRestriction,MutatingAdmissionWebhook,ValidatingAdmissionWebhook,PodSecurity

Mutating Webhook Example (Sidecar Injection)

# mutating-webhook.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: MutatingWebhookConfiguration
metadata:
  name: sidecar-injector
webhooks:
- name: sidecar-injector.example.com
  clientConfig:
    service:
      name: sidecar-injector
      namespace: default
      path: "/mutate"
    caBundle: LS0tLS1CRUdJTi... # Base64 encoded CA cert
  rules:
  - operations: ["CREATE"]
    apiGroups: [""]
    apiVersions: ["v1"]
    resources: ["pods"]
  admissionReviewVersions: ["v1", "v1beta1"]
  sideEffects: None
  timeoutSeconds: 5
  reinvocationPolicy: Never
  failurePolicy: Ignore  # or Fail
  namespaceSelector:
    matchLabels:
      sidecar-injection: enabled

Validating Webhook Example (Resource Quota)

# validating-webhook.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
  name: resource-validator
webhooks:
- name: validate.resources.example.com
  clientConfig:
    service:
      name: resource-validator
      namespace: validation
      path: "/validate"
    caBundle: LS0tLS1CRUdJTi...
  rules:
  - operations: ["CREATE", "UPDATE"]
    apiGroups: ["apps"]
    apiVersions: ["v1"]
    resources: ["deployments"]
  admissionReviewVersions: ["v1"]
  sideEffects: None
  failurePolicy: Fail
  matchPolicy: Equivalent
# Apply webhook configurations
kubectl apply -f mutating-webhook.yaml
kubectl apply -f validating-webhook.yaml

# Test webhook by creating a pod in labeled namespace
kubectl label namespace default sidecar-injection=enabled
kubectl run test --image=nginx

# Check webhook rejections
kubectl describe validatingwebhookconfiguration resource-validator

Advanced Scheduling

Control pod placement using node affinity, pod affinity/anti-affinity, taints, and tolerations.

Key Concepts

  • Node Affinity: Schedule pods to specific nodes based on labels
  • Pod Affinity: Co-locate pods based on labels
  • Pod Anti-Affinity: Spread pods across nodes/zones
  • Taints: Mark nodes to repel pods
  • Tolerations: Allow pods to schedule on tainted nodes
  • Priority Classes: Define pod scheduling priority
Node PoolScheduling DecisionPod to ScheduleFilter NodesApply Affinity RulesCheckTaints/TolerationsScore NodesSelect Best NodeNode 1zone=us-east-1atype=computeNode 2zone=us-east-1btype=memoryNode 3zone=us-east-1atype=gputaint=gpuNode PoolScheduling DecisionPod to ScheduleFilter NodesApply Affinity RulesCheckTaints/TolerationsScore NodesSelect Best NodeNode 1zone=us-east-1atype=computeNode 2zone=us-east-1btype=memoryNode 3zone=us-east-1atype=gputaint=gpu

Node Affinity

# Node affinity example
apiVersion: v1
kind: Pod
metadata:
  name: with-node-affinity
spec:
  affinity:
    nodeAffinity:
      # Hard requirement - must match
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:
        - matchExpressions:
          - key: kubernetes.io/arch
            operator: In
            values:
            - amd64
            - arm64
          - key: node-type
            operator: NotIn
            values:
            - spot
      # Soft preference - prefers but not required
      preferredDuringSchedulingIgnoredDuringExecution:
      - weight: 80
        preference:
          matchExpressions:
          - key: topology.kubernetes.io/zone
            operator: In
            values:
            - us-east-1a
      - weight: 20
        preference:
          matchExpressions:
          - key: instance-type
            operator: In
            values:
            - c5.xlarge
  containers:
  - name: app
    image: nginx

Pod Affinity and Anti-Affinity

# Pod affinity and anti-affinity
apiVersion: apps/v1
kind: Deployment
metadata:
  name: web-cache
spec:
  replicas: 3
  selector:
    matchLabels:
      app: web-cache
  template:
    metadata:
      labels:
        app: web-cache
    spec:
      affinity:
        # Co-locate with web servers
        podAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
          - labelSelector:
              matchExpressions:
              - key: app
                operator: In
                values:
                - web-server
            topologyKey: kubernetes.io/hostname
        # Spread across availability zones
        podAntiAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
          - weight: 100
            podAffinityTerm:
              labelSelector:
                matchExpressions:
                - key: app
                  operator: In
                  values:
                  - web-cache
              topologyKey: topology.kubernetes.io/zone
      containers:
      - name: redis
        image: redis:7

Taints and Tolerations

# Add taint to node
kubectl taint nodes node1 gpu=true:NoSchedule
kubectl taint nodes node2 maintenance=true:NoExecute
kubectl taint nodes node3 environment=production:PreferNoSchedule

# Remove taint
kubectl taint nodes node1 gpu=true:NoSchedule-

# View node taints
kubectl describe node node1 | grep Taints
# Pod with tolerations
apiVersion: v1
kind: Pod
metadata:
  name: gpu-workload
spec:
  tolerations:
  # Exact match toleration
  - key: "gpu"
    operator: "Equal"
    value: "true"
    effect: "NoSchedule"
  # Tolerate any value for key
  - key: "maintenance"
    operator: "Exists"
    effect: "NoExecute"
    tolerationSeconds: 3600  # Evict after 1 hour
  # Wildcard toleration (tolerates all taints)
  - operator: "Exists"
  containers:
  - name: tensorflow
    image: tensorflow/tensorflow:latest-gpu

Priority Classes

# priority-class.yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: high-priority
value: 1000000
globalDefault: false
description: "High priority for critical services"
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: low-priority
value: 100
globalDefault: false
description: "Low priority for batch jobs"
# Use priority class in pod
apiVersion: v1
kind: Pod
metadata:
  name: critical-app
spec:
  priorityClassName: high-priority
  containers:
  - name: app
    image: critical-app:v1
# List priority classes
kubectl get priorityclasses

# Preemption in action (high priority pod evicts low priority)
kubectl describe pod critical-app | grep -A 5 Events

StatefulSets and DaemonSets

Specialised workload controllers for stateful applications and node-level services.

StatefulSets

StatefulSets manage stateful applications requiring stable network identities, persistent storage, and ordered deployment/scaling.

Headless ServiceStatefulSetDNSDNSDNSPod: web-0PVC: data-web-0Pod: web-1PVC: data-web-1Pod: web-2PVC: data-web-2web.default.svcweb-0.web.default.svcweb-1.web.default.svcweb-2.web.default.svcHeadless ServiceStatefulSetDNSDNSDNSPod: web-0PVC: data-web-0Pod: web-1PVC: data-web-1Pod: web-2PVC: data-web-2web.default.svcweb-0.web.default.svcweb-1.web.default.svcweb-2.web.default.svc
# statefulset.yaml - MySQL cluster
apiVersion: v1
kind: Service
metadata:
  name: mysql
  labels:
    app: mysql
spec:
  ports:
  - port: 3306
    name: mysql
  clusterIP: None  # Headless service
  selector:
    app: mysql
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: mysql
spec:
  serviceName: mysql
  replicas: 3
  selector:
    matchLabels:
      app: mysql
  template:
    metadata:
      labels:
        app: mysql
    spec:
      containers:
      - name: mysql
        image: mysql:8.0
        ports:
        - containerPort: 3306
          name: mysql
        env:
        - name: MYSQL_ROOT_PASSWORD
          valueFrom:
            secretKeyRef:
              name: mysql-secret
              key: password
        volumeMounts:
        - name: data
          mountPath: /var/lib/mysql
      - name: xtrabackup
        image: gcr.io/google-samples/xtrabackup:1.0
        ports:
        - containerPort: 3307
          name: xtrabackup
        volumeMounts:
        - name: data
          mountPath: /var/lib/mysql
        - name: conf
          mountPath: /etc/mysql/conf.d
      volumes:
      - name: conf
        configMap:
          name: mysql-config
  volumeClaimTemplates:
  - metadata:
      name: data
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: fast-ssd
      resources:
        requests:
          storage: 10Gi
# StatefulSet operations
kubectl apply -f statefulset.yaml

# Scale StatefulSet (scales sequentially)
kubectl scale statefulset mysql --replicas=5

# Ordered rollout
kubectl set image statefulset/mysql mysql=mysql:8.0.32

# Check rollout status
kubectl rollout status statefulset/mysql

# Delete StatefulSet (keeps PVCs)
kubectl delete statefulset mysql --cascade=orphan

# Force delete a stuck pod
kubectl delete pod mysql-2 --force --grace-period=0

DaemonSets

DaemonSets ensure a pod runs on all (or selected) nodes, typically for node-level services like monitoring, logging, or networking.

# daemonset.yaml - Node monitoring agent
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-exporter
  namespace: monitoring
spec:
  selector:
    matchLabels:
      app: node-exporter
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1
  template:
    metadata:
      labels:
        app: node-exporter
    spec:
      hostNetwork: true
      hostPID: true
      tolerations:
      # Run on all nodes including masters
      - operator: Exists
        effect: NoSchedule
      containers:
      - name: node-exporter
        image: prom/node-exporter:v1.6.0
        args:
        - --path.sysfs=/host/sys
        - --path.rootfs=/host/root
        - --collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)
        ports:
        - containerPort: 9100
          protocol: TCP
          name: metrics
        volumeMounts:
        - name: sys
          mountPath: /host/sys
          readOnly: true
        - name: root
          mountPath: /host/root
          readOnly: true
        resources:
          requests:
            cpu: 100m
            memory: 128Mi
          limits:
            cpu: 200m
            memory: 256Mi
      volumes:
      - name: sys
        hostPath:
          path: /sys
      - name: root
        hostPath:
          path: /
# DaemonSet operations
kubectl apply -f daemonset.yaml

# Check DaemonSet status (should match node count)
kubectl get daemonset -n monitoring
kubectl get pods -l app=node-exporter -o wide

# Update DaemonSet image
kubectl set image daemonset/node-exporter node-exporter=prom/node-exporter:v1.7.0 -n monitoring

# Run DaemonSet on specific nodes only
kubectl label nodes node1 monitoring=enabled
# Then add nodeSelector to DaemonSet spec

Cluster Autoscaling

Automatically adjust cluster capacity based on resource demands.

Key Concepts

  • Cluster Autoscaler (CA): Adds/removes nodes based on pending pods
  • Vertical Pod Autoscaler (VPA): Adjusts pod resource requests/limits
  • Horizontal Pod Autoscaler (HPA): Scales pod replicas based on metrics
  • Node Groups/Pools: Logical groupings of nodes for scaling
HighLowNoYesMetrics ServerHPAVPACPU/MemoryThreshold?Scale Out PodsScale In PodsNodesAvailable?Cluster AutoscalerAdd NodesNodesUnderutilised?Remove NodesAdjust PodRequests/LimitsHighLowNoYesMetrics ServerHPAVPACPU/MemoryThreshold?Scale Out PodsScale In PodsNodesAvailable?Cluster AutoscalerAdd NodesNodesUnderutilised?Remove NodesAdjust PodRequests/Limits

Cluster Autoscaler Setup (AWS)

# cluster-autoscaler.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: cluster-autoscaler
  namespace: kube-system
spec:
  replicas: 1
  selector:
    matchLabels:
      app: cluster-autoscaler
  template:
    metadata:
      labels:
        app: cluster-autoscaler
    spec:
      serviceAccountName: cluster-autoscaler
      containers:
      - image: registry.k8s.io/autoscaling/cluster-autoscaler:v1.28.0
        name: cluster-autoscaler
        command:
        - ./cluster-autoscaler
        - --v=4
        - --stderrthreshold=info
        - --cloud-provider=aws
        - --skip-nodes-with-local-storage=false
        - --expander=least-waste
        - --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/my-cluster
        - --balance-similar-node-groups
        - --skip-nodes-with-system-pods=false
        - --scale-down-enabled=true
        - --scale-down-delay-after-add=10m
        - --scale-down-unneeded-time=10m
        resources:
          requests:
            cpu: 100m
            memory: 300Mi
          limits:
            cpu: 100m
            memory: 300Mi
# Check cluster autoscaler logs
kubectl logs -f deployment/cluster-autoscaler -n kube-system

# View autoscaler status
kubectl describe configmap cluster-autoscaler-status -n kube-system

# Trigger scale-up by creating resource-intensive pods
kubectl create deployment test --image=nginx --replicas=100

Vertical Pod Autoscaler (VPA)

# Install VPA
git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
# vpa.yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: nginx-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: nginx
  updatePolicy:
    updateMode: "Auto"  # Off, Initial, Recreate, Auto
  resourcePolicy:
    containerPolicies:
    - containerName: nginx
      minAllowed:
        cpu: 100m
        memory: 128Mi
      maxAllowed:
        cpu: 2
        memory: 2Gi
      controlledResources: ["cpu", "memory"]
# Check VPA recommendations
kubectl describe vpa nginx-vpa
kubectl get vpa nginx-vpa -o jsonpath='{.status.recommendation}'

Custom Metrics and HPA

Scale applications based on custom or external metrics beyond CPU/memory.

Metrics Architecture

AutoscalingMetrics PipelineExposeQueryQueryScaleResource MetricsApplicationPrometheusPrometheus AdapterCustom Metrics APIHorizontalPodAutoscalerExternal Metrics APIDeploymentMetrics ServerAutoscalingMetrics PipelineExposeQueryQueryScaleResource MetricsApplicationPrometheusPrometheus AdapterCustom Metrics APIHorizontalPodAutoscalerExternal Metrics APIDeploymentMetrics Server

Standard HPA (CPU/Memory)

# hpa-basic.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: nginx-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: nginx
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70
  - type: Resource
    resource:
      name: memory
      target:
        type: Utilization
        averageUtilization: 80
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300
      policies:
      - type: Percent
        value: 50
        periodSeconds: 60
      - type: Pods
        value: 2
        periodSeconds: 60
      selectPolicy: Min
    scaleUp:
      stabilizationWindowSeconds: 0
      policies:
      - type: Percent
        value: 100
        periodSeconds: 30
      - type: Pods
        value: 4
        periodSeconds: 30
      selectPolicy: Max

Custom Metrics HPA

# Install Prometheus Adapter first
# helm install prometheus-adapter prometheus-community/prometheus-adapter

# hpa-custom.yaml - Scale based on HTTP request rate
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: web-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web-app
  minReplicas: 3
  maxReplicas: 20
  metrics:
  # Custom metric from Prometheus
  - type: Pods
    pods:
      metric:
        name: http_requests_per_second
      target:
        type: AverageValue
        averageValue: "1000"
  # Object metric
  - type: Object
    object:
      metric:
        name: queue_depth
      describedObject:
        apiVersion: v1
        kind: Service
        name: rabbitmq
      target:
        type: Value
        value: "30"
# Check HPA status
kubectl get hpa
kubectl describe hpa web-hpa

# View HPA events
kubectl get events --field-selector involvedObject.name=web-hpa

# Test autoscaling with load
kubectl run -i --tty load-generator --rm --image=busybox --restart=Never -- /bin/sh -c "while sleep 0.01; do wget -q -O- http://web-app; done"

External Metrics Example

# hpa-external.yaml - Scale based on SQS queue length
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: sqs-consumer
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: message-processor
  minReplicas: 1
  maxReplicas: 50
  metrics:
  - type: External
    external:
      metric:
        name: sqs_queue_messages_visible
        selector:
          matchLabels:
            queue: "production-queue"
      target:
        type: AverageValue
        averageValue: "30"  # Target 30 messages per pod

Network Policies

Control network traffic between pods and external endpoints using firewall-like rules.

Key Concepts

  • Network Policy: Defines allowed ingress/egress traffic for pods
  • Pod Selector: Identifies pods the policy applies to
  • Policy Types: Ingress, Egress, or both
  • CNI Plugin: Must support NetworkPolicy (Calico, Cilium, Weave)
Namespace: defaultPort 80Port 8080Port 5432XXXFrontend Podsrole=frontendBackend Podsrole=backendDatabase Podsrole=databaseExternal ClientNamespace: defaultPort 80Port 8080Port 5432XXXFrontend Podsrole=frontendBackend Podsrole=backendDatabase Podsrole=databaseExternal Client

Default Deny All Traffic

# deny-all.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-all
  namespace: production
spec:
  podSelector: {}  # Applies to all pods in namespace
  policyTypes:
  - Ingress
  - Egress

Allow Specific Traffic

# backend-policy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: backend-policy
  namespace: production
spec:
  podSelector:
    matchLabels:
      role: backend
  policyTypes:
  - Ingress
  - Egress
  ingress:
  # Allow from frontend pods
  - from:
    - podSelector:
        matchLabels:
          role: frontend
    ports:
    - protocol: TCP
      port: 8080
  # Allow from monitoring namespace
  - from:
    - namespaceSelector:
        matchLabels:
          name: monitoring
    ports:
    - protocol: TCP
      port: 9090
  egress:
  # Allow to database
  - to:
    - podSelector:
        matchLabels:
          role: database
    ports:
    - protocol: TCP
      port: 5432
  # Allow DNS
  - to:
    - namespaceSelector:
        matchLabels:
          name: kube-system
    ports:
    - protocol: UDP
      port: 53
  # Allow external API calls
  - to:
    - ipBlock:
        cidr: 0.0.0.0/0
        except:
        - 10.0.0.0/8
        - 172.16.0.0/12
        - 192.168.0.0/16
    ports:
    - protocol: TCP
      port: 443

Complex Network Policy Example

# microservices-policy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: api-gateway-policy
spec:
  podSelector:
    matchLabels:
      app: api-gateway
  policyTypes:
  - Ingress
  - Egress
  ingress:
  # Allow from ingress controller
  - from:
    - namespaceSelector:
        matchLabels:
          name: ingress-nginx
    - podSelector:
        matchLabels:
          app: nginx-ingress
    ports:
    - protocol: TCP
      port: 8000
  egress:
  # Allow to specific microservices
  - to:
    - podSelector:
        matchExpressions:
        - key: app
          operator: In
          values:
          - user-service
          - product-service
          - order-service
    ports:
    - protocol: TCP
      port: 8080
  # Allow to external auth provider
  - to:
    - ipBlock:
        cidr: 203.0.113.0/24
    ports:
    - protocol: TCP
      port: 443
# Apply network policies
kubectl apply -f deny-all.yaml
kubectl apply -f backend-policy.yaml

# Test network policy
kubectl run test-pod --rm -it --image=busybox -- wget -O- http://backend-service:8080

# View network policies
kubectl get networkpolicies
kubectl describe networkpolicy backend-policy

# Check if CNI supports NetworkPolicy
kubectl get nodes -o jsonpath='{.items[*].status.nodeInfo.containerRuntimeVersion}'

Pod Security Policies and OPA/Gatekeeper

Enforce security standards and compliance policies across the cluster.

Pod Security Standards (PSS)

Pod Security Standards replaced Pod Security Policies (PSPs) in Kubernetes 1.25+.

Pod SecurityStandardsPrivilegedBaselineRestrictedUnrestrictedFor trustedworkloadsMinimal restrictionsPrevents knownprivilegeescalationsHeavily restrictedHardened securitybest practicesPod SecurityStandardsPrivilegedBaselineRestrictedUnrestrictedFor trustedworkloadsMinimal restrictionsPrevents knownprivilegeescalationsHeavily restrictedHardened securitybest practices

Pod Security Admission

# namespace with pod security labels
apiVersion: v1
kind: Namespace
metadata:
  name: production
  labels:
    # Enforce restricted standard
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/enforce-version: latest
    # Warn on baseline violations
    pod-security.kubernetes.io/warn: baseline
    pod-security.kubernetes.io/warn-version: latest
    # Audit privileged violations
    pod-security.kubernetes.io/audit: privileged
    pod-security.kubernetes.io/audit-version: latest
# Example: Pod that meets restricted standard
apiVersion: v1
kind: Pod
metadata:
  name: secure-pod
  namespace: production
spec:
  securityContext:
    runAsNonRoot: true
    runAsUser: 1000
    seccompProfile:
      type: RuntimeDefault
  containers:
  - name: app
    image: nginx:1.25
    securityContext:
      allowPrivilegeEscalation: false
      capabilities:
        drop:
        - ALL
      readOnlyRootFilesystem: true
    volumeMounts:
    - name: cache
      mountPath: /var/cache/nginx
    - name: run
      mountPath: /var/run
  volumes:
  - name: cache
    emptyDir: {}
  - name: run
    emptyDir: {}

OPA Gatekeeper

Gatekeeper is a policy controller using Open Policy Agent (OPA) to enforce custom policies.

# Install Gatekeeper
kubectl apply -f https://raw.githubusercontent.com/open-policy-agent/gatekeeper/master/deploy/gatekeeper.yaml

# Verify installation
kubectl get pods -n gatekeeper-system
kubectl get crds | grep gatekeeper

Constraint Template (Policy Definition)

# constraint-template.yaml - Require labels
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
  name: k8srequiredlabels
spec:
  crd:
    spec:
      names:
        kind: K8sRequiredLabels
      validation:
        openAPIV3Schema:
          type: object
          properties:
            labels:
              type: array
              items:
                type: string
  targets:
  - target: admission.k8s.gatekeeper.sh
    rego: |
      package k8srequiredlabels

      violation[{"msg": msg, "details": {"missing_labels": missing}}] {
        provided := {label | input.review.object.metadata.labels[label]}
        required := {label | label := input.parameters.labels[_]}
        missing := required - provided
        count(missing) > 0
        msg := sprintf("You must provide labels: %v", [missing])
      }

Constraint (Policy Instance)

# constraint.yaml - Enforce labels on namespaces
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
  name: namespace-must-have-owner
spec:
  match:
    kinds:
    - apiGroups: [""]
      kinds: ["Namespace"]
  parameters:
    labels:
    - "owner"
    - "environment"

Common Gatekeeper Policies

# container-limits.yaml - Require resource limits
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
  name: k8scontainerlimits
spec:
  crd:
    spec:
      names:
        kind: K8sContainerLimits
  targets:
  - target: admission.k8s.gatekeeper.sh
    rego: |
      package k8scontainerlimits

      violation[{"msg": msg}] {
        container := input.review.object.spec.containers[_]
        not container.resources.limits.cpu
        msg := sprintf("Container %v has no CPU limit", [container.name])
      }

      violation[{"msg": msg}] {
        container := input.review.object.spec.containers[_]
        not container.resources.limits.memory
        msg := sprintf("Container %v has no memory limit", [container.name])
      }
---
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sContainerLimits
metadata:
  name: must-have-limits
spec:
  match:
    kinds:
    - apiGroups: ["apps"]
      kinds: ["Deployment", "StatefulSet", "DaemonSet"]
# allowed-repos.yaml - Restrict image registries
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
  name: k8sallowedrepos
spec:
  crd:
    spec:
      names:
        kind: K8sAllowedRepos
      validation:
        openAPIV3Schema:
          type: object
          properties:
            repos:
              type: array
              items:
                type: string
  targets:
  - target: admission.k8s.gatekeeper.sh
    rego: |
      package k8sallowedrepos

      violation[{"msg": msg}] {
        container := input.review.object.spec.containers[_]
        satisfied := [good | repo = input.parameters.repos[_] ; good = startswith(container.image, repo)]
        not any(satisfied)
        msg := sprintf("Container %v uses disallowed registry: %v", [container.name, container.image])
      }
---
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sAllowedRepos
metadata:
  name: allowed-registries
spec:
  match:
    kinds:
    - apiGroups: [""]
      kinds: ["Pod"]
    - apiGroups: ["apps"]
      kinds: ["Deployment", "StatefulSet"]
  parameters:
    repos:
    - "gcr.io/my-company/"
    - "docker.io/library/"
    - "registry.k8s.io/"
# Check policy violations
kubectl get constraints
kubectl describe k8srequiredlabels namespace-must-have-owner

# Test policy
kubectl create namespace test  # Should fail without labels

# A namespace WITH the required labels is admitted:
kubectl create namespace test --dry-run=client -o yaml \
  | kubectl label --local -f - owner=platform environment=dev -o yaml \
  | kubectl apply -f -

# View audit violations
kubectl get constraints -o json | jq '.items[].status.violations'

Service Meshes (Istio, Linkerd)

Service meshes provide observability, traffic management, and security for microservices without changing application code.

Service Mesh Architecture

Data PlanePod: Service CPod: Service BPod: Service AControl PlaneConfigConfigConfigmTLSmTLSControl Planeistiod/linkerd-controllerApp ContainerSidecar ProxyEnvoy/Linkerd2-proxyApp ContainerSidecar ProxyApp ContainerSidecar ProxyData PlanePod: Service CPod: Service BPod: Service AControl PlaneConfigConfigConfigmTLSmTLSControl Planeistiod/linkerd-controllerApp ContainerSidecar ProxyEnvoy/Linkerd2-proxyApp ContainerSidecar ProxyApp ContainerSidecar Proxy

Istio Installation

# Download Istio
curl -L https://istio.io/downloadIstio | sh -
cd istio-1.20.0
export PATH=$PWD/bin:$PATH

# Install Istio with demo profile
istioctl install --set profile=demo -y

# Enable sidecar injection for namespace
kubectl label namespace default istio-injection=enabled

# Verify installation
kubectl get pods -n istio-system
istioctl verify-install

Istio Traffic Management

# virtual-service.yaml - Canary deployment
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
  name: reviews
spec:
  hosts:
  - reviews
  http:
  - match:
    - headers:
        user-agent:
          regex: ".*Mobile.*"
    route:
    - destination:
        host: reviews
        subset: v2
      weight: 100
  - route:
    - destination:
        host: reviews
        subset: v1
      weight: 90
    - destination:
        host: reviews
        subset: v2
      weight: 10
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
  name: reviews
spec:
  host: reviews
  trafficPolicy:
    connectionPool:
      tcp:
        maxConnections: 100
      http:
        http1MaxPendingRequests: 10
        http2MaxRequests: 100
        maxRequestsPerConnection: 2
    outlierDetection:
      consecutiveErrors: 5
      interval: 30s
      baseEjectionTime: 30s
      maxEjectionPercent: 50
  subsets:
  - name: v1
    labels:
      version: v1
  - name: v2
    labels:
      version: v2
    trafficPolicy:
      loadBalancer:
        simple: ROUND_ROBIN

Istio Security (mTLS)

# peer-authentication.yaml - Enforce mTLS
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
  name: default
  namespace: production
spec:
  mtls:
    mode: STRICT  # PERMISSIVE, STRICT, DISABLE
---
# authorization-policy.yaml
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata:
  name: frontend-policy
  namespace: production
spec:
  selector:
    matchLabels:
      app: frontend
  action: ALLOW
  rules:
  - from:
    - source:
        principals:
        - "cluster.local/ns/production/sa/api-gateway"
    to:
    - operation:
        methods: ["GET", "POST"]
        paths: ["/api/*"]
    when:
    - key: request.headers[x-api-key]
      values: ["*"]

Istio Circuit Breaking and Retries

# circuit-breaker.yaml
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
  name: backend-circuit-breaker
spec:
  host: backend.production.svc.cluster.local
  trafficPolicy:
    connectionPool:
      tcp:
        maxConnections: 100
      http:
        http1MaxPendingRequests: 10
        maxRequestsPerConnection: 2
    outlierDetection:
      consecutive5xxErrors: 5
      interval: 5s
      baseEjectionTime: 30s
      maxEjectionPercent: 50
      minHealthPercent: 40
---
# retry-policy.yaml
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
  name: backend-retries
spec:
  hosts:
  - backend
  http:
  - route:
    - destination:
        host: backend
    retries:
      attempts: 3
      perTryTimeout: 2s
      retryOn: 5xx,reset,connect-failure,refused-stream
    timeout: 10s

Linkerd Installation

# Install Linkerd CLI
curl -fsL https://run.linkerd.io/install | sh
export PATH=$PATH:$HOME/.linkerd2/bin

# Validate cluster
linkerd check --pre

# Install Linkerd
linkerd install --crds | kubectl apply -f -
linkerd install | kubectl apply -f -

# Verify installation
linkerd check

# Enable auto-injection for namespace
kubectl annotate namespace default linkerd.io/inject=enabled

# Manually inject sidecar
kubectl get deploy -o yaml | linkerd inject - | kubectl apply -f -

Linkerd Traffic Split (SMI)

# traffic-split.yaml - Canary with SMI
apiVersion: split.smi-spec.io/v1alpha2
kind: TrafficSplit
metadata:
  name: backend-split
spec:
  service: backend
  backends:
  - service: backend-v1
    weight: 900
  - service: backend-v2
    weight: 100
---
# Service for each version
apiVersion: v1
kind: Service
metadata:
  name: backend-v1
spec:
  selector:
    app: backend
    version: v1
  ports:
  - port: 8080
---
apiVersion: v1
kind: Service
metadata:
  name: backend-v2
spec:
  selector:
    app: backend
    version: v2
  ports:
  - port: 8080

Service Mesh Observability

# Istio - Access Kiali dashboard
istioctl dashboard kiali

# Istio - View metrics in Prometheus
istioctl dashboard prometheus

# Istio - Distributed tracing with Jaeger
istioctl dashboard jaeger

# Linkerd - Dashboard
linkerd viz dashboard

# Linkerd - View service metrics
linkerd viz stat deployments
linkerd viz top deployments
linkerd viz tap deployment/backend

Hybrid and Multi-Cluster Setups

Manage applications across multiple clusters or hybrid cloud/on-premises environments.

Multi-Cluster Patterns

Control PlaneRegion 2Region 1SyncSyncSyncReplicateReplicateCluster 1PrimaryCluster 2DRCluster 3ReplicaMulti-ClusterControllerKubeFed/ArgoCDControl PlaneRegion 2Region 1SyncSyncSyncReplicateReplicateCluster 1PrimaryCluster 2DRCluster 3ReplicaMulti-ClusterControllerKubeFed/ArgoCD

KubeFed (Kubernetes Federation)

KubeFed is archived/retired (moved to kubernetes-retired/kubefed, read-only since 2023). For new multi-cluster work prefer Argo CD ApplicationSets, Cluster API, or a fleet manager (Rancher/Karmada). The commands below are kept for existing deployments only.

# Install KubeFed (archived project)
kubectl create ns kube-federation-system
helm repo add kubefed-charts https://raw.githubusercontent.com/kubernetes-sigs/kubefed/master/charts
helm install kubefed kubefed-charts/kubefed --namespace kube-federation-system

# Join clusters to federation
kubefedctl join cluster1 --cluster-context cluster1 --host-cluster-context cluster1
kubefedctl join cluster2 --cluster-context cluster2 --host-cluster-context cluster1

# Verify joined clusters
kubectl get kubefedclusters -n kube-federation-system
# federated-deployment.yaml
apiVersion: types.kubefed.io/v1beta1
kind: FederatedDeployment
metadata:
  name: test-deployment
  namespace: default
spec:
  template:
    metadata:
      labels:
        app: nginx
    spec:
      replicas: 3
      selector:
        matchLabels:
          app: nginx
      template:
        metadata:
          labels:
            app: nginx
        spec:
          containers:
          - name: nginx
            image: nginx:1.25
  placement:
    clusters:
    - name: cluster1
    - name: cluster2
  overrides:
  - clusterName: cluster2
    clusterOverrides:
    - path: "/spec/replicas"
      value: 5

Multi-Cluster Service Discovery (Istio)

# Install Istio on both clusters with multi-cluster support
# Cluster 1 (primary)
istioctl install --set profile=default \
  --set values.global.meshID=mesh1 \
  --set values.global.multiCluster.clusterName=cluster1 \
  --set values.global.network=network1

# Create remote secret for cluster2
istioctl create-remote-secret --name=cluster2 | kubectl apply -f -

# Cluster 2 (remote)
istioctl install --set profile=default \
  --set values.global.meshID=mesh1 \
  --set values.global.multiCluster.clusterName=cluster2 \
  --set values.global.network=network2

# Enable cross-cluster service discovery
kubectl apply -f - <<EOF
apiVersion: networking.istio.io/v1
kind: Gateway
metadata:
  name: cross-network-gateway
  namespace: istio-system
spec:
  selector:
    istio: eastwestgateway
  servers:
  - port:
      number: 15443
      name: tls
      protocol: TLS
    tls:
      mode: AUTO_PASSTHROUGH
    hosts:
    - "*.local"
EOF

Cluster API (ClusterAPI)

# Install clusterctl
curl -L https://github.com/kubernetes-sigs/cluster-api/releases/download/v1.6.0/clusterctl-linux-amd64 -o clusterctl
chmod +x clusterctl
sudo mv clusterctl /usr/local/bin/

# Initialise management cluster
clusterctl init --infrastructure aws

# Create workload cluster
clusterctl generate cluster workload-cluster \
  --kubernetes-version v1.28.0 \
  --control-plane-machine-count=3 \
  --worker-machine-count=5 \
  > workload-cluster.yaml

kubectl apply -f workload-cluster.yaml

# Get kubeconfig for new cluster
clusterctl get kubeconfig workload-cluster > workload-cluster.kubeconfig

GitOps Multi-Cluster (ArgoCD)

# argocd-cluster-secret.yaml
apiVersion: v1
kind: Secret
metadata:
  name: cluster2-secret
  namespace: argocd
  labels:
    argocd.argoproj.io/secret-type: cluster
type: Opaque
stringData:
  name: cluster2
  server: https://cluster2.example.com
  config: |
    {
      "bearerToken": "<token>",
      "tlsClientConfig": {
        "insecure": false,
        "caData": "<base64-ca-cert>"
      }
    }
# multi-cluster-app.yaml
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: multi-cluster-app
  namespace: argocd
spec:
  generators:
  - clusters:
      selector:
        matchLabels:
          environment: production
  template:
    metadata:
      name: '{{name}}-guestbook'
    spec:
      project: default
      source:
        repoURL: https://github.com/example/manifests
        targetRevision: HEAD
        path: guestbook
      destination:
        server: '{{server}}'
        namespace: guestbook
      syncPolicy:
        automated:
          prune: true
          selfHeal: true

Quick Reference

CRD Operations

# List CRDs
kubectl get crds

# Describe CRD
kubectl explain crontabs.spec

# Get custom resources
kubectl get crontabs
kubectl get crontabs.stable.example.com

# Delete CRD (deletes all instances)
kubectl delete crd crontabs.stable.example.com

Scheduling Commands

# Label nodes
kubectl label nodes node1 disktype=ssd
kubectl label nodes node1 zone=us-east-1a

# Taint nodes
kubectl taint nodes node1 key=value:NoSchedule
kubectl taint nodes node1 key=value:NoExecute
kubectl taint nodes node1 key-  # Remove taint

# View node details
kubectl describe node node1 | grep -A 5 Taints
kubectl get nodes --show-labels

Autoscaling Commands

# HPA
kubectl autoscale deployment nginx --cpu-percent=70 --min=2 --max=10
kubectl get hpa
kubectl describe hpa nginx

# VPA
kubectl get vpa
kubectl describe vpa nginx-vpa

# Cluster Autoscaler
kubectl logs -f deployment/cluster-autoscaler -n kube-system

Network Policy Commands

# List network policies
kubectl get networkpolicies
kubectl get netpol

# Describe policy
kubectl describe networkpolicy backend-policy

# Test connectivity
kubectl run test --rm -it --image=busybox -- wget -O- http://service:8080

Service Mesh Commands

# Istio
istioctl version
istioctl proxy-status
istioctl analyze
kubectl label namespace default istio-injection=enabled

# Linkerd
linkerd version
linkerd check
linkerd viz stat deployments
linkerd viz tap deployment/backend
kubectl annotate namespace default linkerd.io/inject=enabled

Common Issues and Solutions

CRD Issues

Issue Solution
CRD not found after creation Wait for API discovery cache refresh (up to 60s) or restart API server
Schema validation failing Check openAPIV3Schema matches CR fields exactly
Status subresource not updating Ensure subresources.status: {} in CRD spec and use /status endpoint
Version conflict Use conversion webhooks or set single storage: true version

Scheduling Issues

Issue Solution
Pod pending despite resources Check node selectors, taints, and PodTopologySpread constraints
Affinity rules too restrictive Use preferredDuringScheduling instead of required for soft constraints
Taints preventing scheduling Add tolerations to pod spec matching node taints
Priority preemption not working Ensure PriorityClass exists and value differences are significant (>100)

Autoscaling Issues

Issue Solution
HPA not scaling Verify metrics-server is running and pod has resource requests defined
Cluster Autoscaler not adding nodes Check node group min/max settings and cloud provider IAM permissions
VPA conflicts with HPA Don't use VPA and HPA on same metric (use VPA for requests, HPA for replicas)
Flapping (rapid scale up/down) Increase stabilizationWindowSeconds in HPA behavior config

Network Policy Issues

Issue Solution
Policy not enforced Verify CNI plugin supports NetworkPolicy (Calico, Cilium, Weave)
DNS resolution failing Allow egress to kube-dns/coredns on port 53 UDP
Can't reach external IPs Add egress rule with ipBlock for external CIDR ranges
Policy applied but traffic still allowed Check for multiple policies - allow rules are additive

Service Mesh Issues

Issue Solution
Sidecar not injected Verify namespace label and check mutating webhook configuration
mTLS connection failures Ensure PeerAuthentication mode matches (PERMISSIVE during migration)
High latency after mesh install Adjust resource limits for sidecars, check circuit breaker settings
Certificate expiration Istio auto-rotates certs, check istio-ca-secret and citadel logs

Multi-Cluster Issues

Issue Solution
Cross-cluster service discovery fails Verify east-west gateway and DNS configuration
Federation sync delays Check KubeFed controller logs and network latency between clusters
Cluster join failing Ensure correct kubeconfig context and network connectivity
Resource conflicts across clusters Use cluster-specific overrides in FederatedResource