Kubernetes overhead on bare metal
The same CPU, memory and disk benchmarks, run on three x86 servers and two Raspberry Pis both directly on the host and as Kubernetes CronJobs, plus a prototype scheduler that places pods by past benchmark results. The data fed a 2024 paper on workload placement in Kubernetes-based HPC data centres.
The question
Jurica Slovinac and I measured how much throughput a workload loses in a Kubernetes pod compared with a plain process on the same machine. I set up the clusters and the benchmark CronJobs and wrote the write-up; he wrote the log parsers that turn the runs into the workbook, and the scheduler described below.
The benchmark data went into Evolving High-Performance Computing Data Centers with Kubernetes, Performance Analysis, and Dynamic Workload Placement Based on Machine Learning Scheduling, by Vedran Dakić, Mario Kovač and Jurica Slovinac (Electronics 13(13), 2651, 2024), where the same kind of baseline runs feed a workload scheduler. I am not an author of that paper. This page looks at the workbook itself, which the paper does not report host by host.
Setup
doktor1, doktor2 and doktor3 are HP ProLiant Gen8 servers, each with two Xeon E5 processors and 64 to 102 GB of DDR3. pi1 and pi2 are Raspberry Pis with a four-core Broadcom BCM2711 and 4 GB of LPDDR4. doktor1 runs CentOS 7, the other two servers Ubuntu 22.04. The servers form one Kubernetes cluster and the Pis another.
Each benchmark runs as a CronJob pinned to its node, from one multi-architecture image, and as the same command in root’s crontab on the host. The schedules keep tests from overlapping, and the runs span two weeks.
sysbench tests CPU (10,000 events, prime limit 100,000) and memory writes, and stress-ng runs its CPU stressor. Each of the three runs with one thread and with every hardware thread. hdparm times a cached read and a buffered disk read. That makes eight tests, with 13 to 305 recorded runs per test, node and environment.
Results
Kubernetes average against OS-level average, in percent. Positive means the pod was faster.
| Test | doktor1 | doktor2 | doktor3 | pi1 | pi2 |
|---|---|---|---|---|---|
| sysbench CPU, 1 thread | +6.5 | −0.2 | −0.3 | −1.0 | −1.0 |
| sysbench CPU, all threads | −3.0 | −2.1 | −3.1 | −2.2 | −2.3 |
| sysbench memory, 1 thread | +13.5 | +0.1 | +0.3 | +11.6 | +11.1 |
| sysbench memory, all threads | +7.1 | −0.3 | −0.4 | +9.1 | +9.6 |
| stress-ng CPU, 1 thread | +392.8 | −2.8 | −8.2 | +115.1 | +110.6 |
| stress-ng CPU, all threads | +385.1 | +0.2 | +0.2 | +164.8 | +154.4 |
| hdparm cached read | −0.3 | +1.9 | +1.2 | +0.9 | +5.0 |
| hdparm disk read | +1.2 | +27.6 | −0.2 | −0.1 | −0.1 |
Multi-threaded sysbench CPU is slower in a pod on all five nodes, by 2.1 to 3.1%.
doktor2 and doktor3 give the cleanest comparison, since both sides there ran the same tool builds. Memory and multi-threaded stress-ng stay within 0.4%, while single-threaded stress-ng loses 2.8 and 8.2%. doktor2’s disk read looks 27.6% faster in Kubernetes because 77 of its 160 averaged OS-level runs read 107 to 155 MB/s. The other 83 read 224 to 252 MB/s, inside the 222 to 254 band of its Kubernetes runs.
On doktor1, pi1 and pi2 the pods are 7.1 to 13.5% faster on memory and 2.1 to 4.9 times faster on stress-ng. The raw logs show a confound: the OS-level runs on these hosts used older sysbench and stress-ng builds from the host’s packages, while every pod ran the same Ubuntu 22.04 image.
The write-up drops doktor1 from the single-threaded stress-ng result, citing its older operating system, and calls that result “present only on hp-1”, its name for doktor1. It is not: pi1 and pi2 gain 115.1 and 110.6% on the same test, and the write-up discusses the Pis elsewhere, crediting their LPDDR4 memory for the memory results. The 2 to 8% CPU cost in its conclusion matches multi-threaded sysbench and single-threaded stress-ng on doktor2 and doktor3. Its other conclusion, a memory gain of about 5%, matches no host in the workbook: the three with mixed builds gain 7.1 to 13.5%, and doktor2 and doktor3 move by less than 0.4%.
The benchmark-aware scheduler
Jurica Slovinac wrote a prototype scheduler in Python, in the same repository. A pod opts in with schedulerName: benchmark-scheduler and a test label naming a benchmark. At start-up the scheduler adds up each node’s recorded results per benchmark and keeps the node with the highest total. It then binds each opted-in pod in the default namespace to the node stored for its label. Resource requests, current load and node readiness play no part. It does ask the API at start-up for the Ready, schedulable nodes labelled node=worker, and then never uses that list.
Adding up favours nodes with more runs. For multi-threaded memory in Kubernetes it picks doktor1 (160 runs, average 7.78 million operations per second) over pi2 (112 runs, 8.65 million). The candidates also span both clusters, while one scheduler process talks to one.
It was built and not evaluated. There are example Jobs that opt in, but no recorded placements, no comparison with the default scheduler, and no mention in the write-up.
Limits
The workbook holds one two-week campaign, and the write-up itself was never published. Sample sizes are uneven: in Kubernetes, doktor2 has 13 single-threaded stress-ng runs against 96 on doktor1, and the workbook’s hdparm averages cover 159 to 162 values from columns that hold 183 to 305. Only doktor2 and doktor3 compare like with like, a thin base for a general figure. Only throughput was measured, with no latency, network or disk-write tests.