research

How Kubernetes behaves when its components disagree

Fault-injection experiments on an 11-VM Kubernetes lab, asking which guarantees still hold when the control plane and the nodes report different things about the same cluster. Three papers are in preparation.

The question

During a fault, parts of a Kubernetes cluster can report different things about the same object. The control plane can mark a node as failed while its containers keep running, and two API servers can disagree about whether the same node is Ready. I want to know which of the guarantees Kubernetes gives still hold when that happens.

The lab

The experiments run on 11 VMs on VMware vSphere: 3 control-plane nodes, 3 workers and a separate etcd cluster of 5 members. etcd sits outside the control-plane nodes so a scenario can slow it down, compact it or partition it on its own. It has five members so that a partition can split it 3 to 2 and still leave a majority.

I build the Ubuntu 24.04 template with Packer, clone the VMs from it with Terraform and install the cluster with Ansible: Kubernetes 1.36, Calico, and a keepalived address in front of the three API servers.

Between runs all 11 VMs go back to one cold snapshot, taken with the fleet powered down. vCenter cannot snapshot several running VMs at the same instant, and etcd members restored from different moments can come back at different points in their logs.

Diagram of the lab on VMware vSphere: three control-plane VMs behind a keepalived API address, three worker VMs below them, and a separate five-member etcd cluster beside them. A workstation outside the cluster runs Packer, Terraform and Ansible.
fig. 1 The 11 VMs. Workers reach the API servers through the keepalived address, and the API servers store cluster state in the external etcd cluster. Source: Drawn by Matej Bašić from the lab's Terraform configuration.

Fault scenarios

The suite has 23 scenarios. Fourteen are crash-stop and timing faults, such as network partitions, slow etcd writes, etcd compaction and clock skew. The other nine are gray failures, where every node stays Ready while something underneath has stopped: the pod network cut with the management network still up, containerd frozen with SIGSTOP, or etcd refusing writes after a NOSPACE alarm.

Each scenario checks that the cluster starts clean, injects the fault, watches for a set window, removes the fault and records the state once the cluster settles.

How a run is judged

There are nine invariants. Several of them go around the API server, which cannot see a partitioned node. The check for two copies of one pod reads the container runtime on every node over SSH and fires when the same pod name runs on two nodes at the same moment. At each snapshot, twelve sources are asked the same questions about worker1, a stateful test pod and a test service, and each API server is asked at its own address. Eleven of them can answer. The in-cluster informer is never created, so a snapshot records it as unreadable instead of a wrong answer.

Only PRESERVED and VIOLATED count towards a rate. INSUFFICIENT_DATA means the run lacks the evidence an invariant needs. NOT_APPLICABLE comes from a declared table of fault classes that cannot reach what an invariant reads. DETECTOR_INVALID means the invariant was already broken before the fault, so the detector is wrong and the verdict is dropped. A whole run is marked FAULT_INEFFECTIVE when the fault left no observable change, and none of its verdicts count.

Where the work stands

Three papers are in preparation, and I am first author on the first. It covers the 23 scenarios above. All of them have run on the lab, and I will re-run the suite with the current checker before the paper quotes a figure. Results go on this page once that paper is out.

The second paper uses the same 11 VMs and asks whether the rules the API server records are the rules the agents on each node apply, for example Calico’s policy agent, kube-proxy and CoreDNS. Three of its ten scenarios are built and have run on the lab, each with its fault and null arms, and a two-arm feasibility check ran before them. The other seven are not written yet. What those runs show goes on this page when the paper is out.

The third starts from the first paper’s runs and asks what happens to stored data when a fault leaves two copies of one workload running, each acting as the only one. It adds six scenarios, and none of them is built yet.

The code for each paper goes public with that paper.