DEV Community
Follow
I Deliberately Destroyed My Kubernetes Cluster at 2 AM. Here's What Died First.
Chaos engineering, even on a homelab, is crucial for uncovering hidden vulnerabilities. A homelab Kubernetes cluster, despite appearing robust on paper with various cloud-native tools, was discovered to be fragile under stress. Testing revealed that Kubernetes' rapid pod recreation doesn't extend to stateful workloads, with database failovers taking significant time. Network partitions caused distributed deadlocks, as nodes remained "Ready" long enough to break services but not to allow clean workload migration. Physical node failure exposed excessively long default Kubernetes timeouts for bare metal, causing prolonged service unavailability. Stressing the control plane demonstrated its susceptibility to a single noisy process, leading to cascading failures across critical services. Finally, DNS failures crippled the entire cluster simultaneously, highlighting its fundamental importance. These tests effectively exposed the cluster's true mean time to recovery, prompting necessary configuration adjustments.