eBPF Networking is Fast—But at What Cost?

Every vendor pitch deck this year has eBPF on it. Cilium benchmarks show 28.5 Gbps versus iptables' 22.1 Gbps. Faster service mesh. Faster observability. Faster everything. And it's all true — until it isn't.
The trade-off nobody wants to talk about is predictability. We've moved from "the kernel does networking, applications don't crash it" to "anyone with CAP_BPF can attach a program to your packet path". That's a different operational model, even if the marketing pretends otherwise.
Programmability has a tax
Every eBPF program attached to a hook runs on every packet that hits that hook. Most are fine. Some aren't. I've seen a perfectly innocent-looking observability agent add 40µs of p99 latency to a tail-sensitive workload because it was hashing every flow into a map that occasionally needed rebalancing. The verifier passed it. The benchmarks looked great. Production hated it.
The bigger problem is composition. Run Cilium for CNI, Pixie for observability, a security agent for runtime detection, and maybe Tetragon for policy. Each one is "lightweight" in isolation. Together you've got half a dozen programs on tc ingress alone, and nobody owns the cumulative budget. There's no top for eBPF. bpftool prog show tells you what's loaded, not what it's costing you.
What breaks, and how
When eBPF goes wrong, it goes wrong in ways traditional networking doesn't. Map exhaustion silently drops events. A kernel upgrade changes a struct layout and your CO-RE relocations start behaving oddly on a subset of nodes. Verifier limits force someone to split a program into tail calls, and now your debugging story involves reading instruction counts.
A 2025 paper made the point that many networked applications can't actually benefit from eBPF anyway — the bottleneck is in userspace, or in the application protocol, not in the packet path. We're paying the complexity cost for performance gains that, for plenty of workloads, don't materialise.
So what
Use eBPF. It's genuinely the best tool we have for several jobs. But treat each program as production code with an SLO budget, not a free upgrade. Audit what's loaded. Measure tail latency, not throughput averages. And stop pretending kernel-level programmability is the same operational risk profile as iptables. It isn't. It's better in most ways and worse in a few that matter a lot when something goes sideways at 3am.