Kubernetes Networking on Multi-Homed ConnectX-5 Nodes: Node Address Path Selection, CNI Performance Over IPoIB, and Pod-Level RoCE
Matej Bašić, Vedran Dakić, Mario Kovač and Josip Knezović. Kubernetes Networking on Multi-Homed ConnectX-5 Nodes: Node Address Path Selection, CNI Performance Over IPoIB, and Pod-Level RoCE. Electronics 15(20), 4604, 2026.
Abstract
On multi-homed Kubernetes nodes, the interface carrying node addresses can determine the physical path used by default CNI deployments, potentially obscuring the performance available from a high-speed fabric. We evaluate Antrea, Cilium, Calico, and Flannel across nine Linux tuning profiles on 100 GbE and 100 Gb/s InfiniBand platforms, including pod-level RoCE and IPoIB measurements. On Ethernet, chart-default CNIs shared a 0.908 Gbit/s TCP ceiling, matching the payload rate of the 1 GbE management link carrying node addresses. On InfiniBand, relocating node addresses from 10 GbE management to IPoIB increased peak CNI throughput from 9.1 to 44.4 Gbit/s, approaching the 46 Gbit/s host baseline. Over IPoIB, median throughput was 42.4 Gbit/s for Flannel, 32.9 for Cilium, 32.2 for Antrea, and 7.0 for default IP-in-IP Calico. Flannel minimized pod-to-host TCP latency, although tuning nearly tripled CNI latency. On Ethernet, pod RoCEv2 reached 96.9–97.8 Gbit/s at MTU 9000 with 0.835–2.035 µs of latency. Host and pod RoCE performance was comparable for the measured profile. These results validate node address/data path placement as a primary determinant of observed CNI throughput and support explicitly directing CNI traffic onto the intended high-speed interface while exposing RDMA-capable devices to performance-critical workloads.
My contribution
I designed the lab environment and handled all of the work at the operating-system level. I also ran every test the paper reports.
CC BY 4.0.