You Rate-Limited the Internet — But Not Yourself

There's a pattern that shows up, reliably, in post-mortems for incidents that shouldn't have happened. The external API is rate-limited. The CDN has WAF rules. The ingress has throttling policies carefully tuned over years of production traffic. And then one internal service — a batch job, a newly deployed consumer, a reporting query that someone decided should run every thirty seconds instead of every five minutes — hammers a downstream service into the ground, and half the platform goes with it.
The irony is rarely appreciated in the moment. You spent months hardening the perimeter, and you got taken out by your own infrastructure.
The Trust Assumption
The underlying problem is a trust model that was outdated before most of us started writing production code. The implicit assumption — that traffic originating inside the network is safe by default — comes from an era when "internal" meant something. A VLAN with controlled membership. A data centre where you knew every machine. It was never a great assumption, but it was at least coherent.
It isn't coherent any more. A Kubernetes cluster is not a trust boundary. A VPC is not a trust boundary. A service mesh with permissive default policies is not a trust boundary. What you have, in most modern microservices deployments, is a flat network with good intentions and the faint hope that nothing will go wrong.
The security implications of this are well understood — lateral movement, compromised workloads, supply chain entry points. The reliability implications get less airtime, which is odd given that they're considerably more likely to bite you in any given week. A misbehaving internal service doesn't need to be compromised to cause damage. It just needs to retry too aggressively, or get deployed with a misconfigured polling interval, or hit a code path that wasn't tested under load. The intent is entirely benign. The outcome is indistinguishable from an attack.
This is what Google's SRE book calls "friendly-fire denial of service" — resource starvation caused not by malicious actors but by errors in configuration or software elsewhere in the system. The name is slightly whimsical. The failure mode is not.
How the Cascade Actually Works
Cascading failures follow a consistent script, which is part of why they're so demoralising: you've seen this before, you know what's happening, and you're still watching it happen.
A service starts returning elevated latency. Its clients, which are also your services, don't notice immediately — their timeouts haven't been hit yet, but their connection pools are filling up. Requests start queuing. The threads or goroutines handling those requests are now sitting idle waiting for responses that are slow to arrive, which means fewer workers are available for new incoming requests. Latency climbs further. Eventually something trips — a timeout, a circuit breaker if you're lucky enough to have one, a hard connection limit — and you start seeing errors. Those errors propagate upstream. Upstream services begin retrying, because that's what they're configured to do. The retry load hits the already-struggling downstream service. The feedback loop is now running, and it's not running in your favour.
The specific failure mode depends on where the load concentrates. If a large number of service instances all fan out to a small number of downstream hosts — the classic database or shared cache scenario — circuit breaking per instance becomes almost impossible to tune correctly. The limits that protect you during steady-state traffic may be too loose to prevent cascade when the system is already degraded, because each individual instance sees only its slice of the total load. Envoy's global rate limiting exists precisely for this reason: when downstream capacity is small relative to upstream fan-out, you need coordinated admission control, not per-instance circuit breaking.
The part that catches teams off guard is the speed. You're not watching a slow degradation over hours. A cluster can go from normal operation to service-wide failure in two minutes, because the load balancer is doing exactly what it was designed to do — redistributing traffic away from struggling instances, which just moves the problem onto instances that were previously fine.
External Rate Limiting Is Not the Same Problem
When you rate-limit external traffic, you're primarily protecting against clients you don't control, who may be acting in bad faith, and whose retry behaviour you can't influence. The strategy is largely static: per-key limits, IP-based limits, aggregate limits per tier. The client gets a 429, hopefully backs off, hopefully reads your documentation. Some of them don't, which is why you have the limits in the first place.
Internal rate limiting has a different character entirely. The clients are your own services, which means you can design both sides of the interaction. You can propagate backpressure signals rather than just returning errors and hoping for the best. You can make the rate limiting adaptive rather than static, because you have visibility into system state that you'd never have for an external caller.
Static limits on internal traffic are better than nothing, but they're a blunt instrument. You set the limit at deployment time based on some estimate of normal load, and you discover the estimate was wrong during an incident. Adaptive approaches tie the limit to observed system behaviour: latency percentiles, error rates, queue depth, in-flight request counts. The AIMD algorithm — additive increase, multiplicative decrease — is essentially TCP congestion control applied to request throughput. When conditions are good, you slowly increase the allowed rate. When you see signs of overload, you cut sharply. The asymmetry is deliberate; you want recovery to be cautious, not optimistic.
Stripe's published architecture for their internal load shedding takes a related approach: rather than a single rate limit applied uniformly, they stratify traffic by criticality and shed lowest-priority categories first. Test-mode traffic goes before live-mode GET requests go before live-mode POST requests go before the methods that actually move money. This isn't just rate limiting — it's prioritised load shedding with explicit capacity tiers. The system degrades gracefully under pressure rather than collapsing uniformly.
The client side of this matters too. A server-side rate limit that produces a flood of fast error responses can itself become a capacity problem — processing the errors and returning 429s costs something, and if you have enough clients hammering simultaneously, that cost isn't negligible. Client-side throttling, where a service monitors its own rejection rate and begins dropping requests before sending them, reduces the load on both sides. It's not a substitute for server-side limits, but it changes the failure mode from "server drowns in rejected requests" to "client self-regulates before it becomes the problem."
The Identity Problem
One reason internal rate limiting is harder to implement than it looks is that most of it gets bolted on after the fact, which means the services weren't designed with admission control in mind. But there's a subtler problem underneath that: inside the perimeter, callers frequently aren't identified at all.
External API rate limiting relies on identity — API keys, JWT claims, IP addresses, OAuth tokens. You know who's making the request, which is why you can apply per-caller limits and track usage over time. Internal services often pass no caller identity whatsoever. The request arrives at a service, and that service has no idea whether it's coming from the nightly batch job, the real-time event consumer, the data science team's new notebook that's been accidentally pointed at production, or a bug that's causing some other service to call this one in a tight loop.
Without caller identity, your rate limiting is forced to be aggregate — total request rate regardless of source — which means a single misbehaving caller can consume the headroom that you'd need for everyone else. mTLS with SPIFFE/SPIRE workload identities, or RBAC in a service mesh like Istio or Linkerd, gives you cryptographic caller identity on every request without requiring your application code to do anything about it. Once you have that, you can rate limit per workload rather than per endpoint, which is a substantially more useful primitive.
The Kubernetes Gateway API's GAMMA initiative is pushing this direction for east-west traffic specifically — standardised policy attachment that applies rate limiting or traffic shaping at the namespace or workload level, with identity sourced from service accounts rather than network addresses. The security model is correct: trust is based on cryptographic identity, not on which CIDR range the packet came from.
What Observability Actually Needs to Catch
You can have rate limiting in place and still miss the problem, because the signals don't always surface where you're looking.
The most common gap is aggregate versus per-caller metrics. If your internal service endpoint is emitting request rate and error rate at the aggregate level, a single caller consuming 80% of your capacity will look fine until it doesn't — the aggregate numbers may be well within normal bounds right up until they aren't. Cardinality matters here: you need request rate and latency broken down by calling service, not just by endpoint. Prometheus with appropriate labels does this well; the cost is higher cardinality in your time series, which is a real trade-off but usually the right one.
The second gap is leading versus lagging indicators. By the time error rates are climbing, you're already in the cascade. The leading indicators are latency percentiles — specifically p99 and p999, not just mean — and in-flight request count relative to expected concurrency. If your p99 latency on a service doubles while your p50 is stable, that's a sign the tail is getting worse, which is often the earliest observable sign of the system approaching a phase transition. Mean latency will look fine until it suddenly isn't, because the distribution is changing shape before the centre moves.
Queue depth is worth instrumenting on any async path. A RabbitMQ queue that's normally processing with near-zero depth will start accumulating messages before consumers show any distress; that's your early warning. On synchronous HTTP paths, connection pool saturation is the equivalent signal — if your connection pool to a downstream service is consistently at 80%+ utilisation, you have very little headroom before backpressure becomes errors.
The third gap is the absence of attribution on shed load. When your service drops a request due to rate limiting or load shedding, you should be recording why and, if you have caller identity, who. Aggregate drop counts tell you the rate limiter is working; attributed drop counts tell you which caller is causing you to hit it, which is the information you need to have a productive conversation with the team that owns that service. "Your batch job is consuming 60% of our internal API capacity between 02:00 and 04:00 every night" is a useful observation. "We're seeing elevated rejection rates" is not.
Treating Internal Traffic as Hostile by Default
The practical upshot of all this is that "trust but verify" needs to become "verify, then trust within defined limits." That's not a security posture, or not only a security posture — it's an availability posture. Services that are designed to handle known callers with known workloads are substantially easier to operate than services that are designed to handle whatever arrives.
Concretely, that means: mTLS on service-to-service traffic so you have cryptographic caller identity; per-caller admission control, not just aggregate rate limits; adaptive limits that respond to system state rather than static thresholds set at deploy time; load shedding with priority tiers for any service that sees mixed-criticality traffic; and instrumentation that attributes shed load to specific callers rather than reporting it as undifferentiated noise.
None of this is architecturally novel. The Google SRE book covers cascading failure in detail; the Amazon Builders' Library has a clear-eyed treatment of admission control and fairness in multi-tenant systems. The Envoy docs explain global rate limiting for exactly the fan-out-to-small-cluster scenario that accounts for a large fraction of production incidents. The patterns exist. The tooling exists.
The missing piece is usually the assumption — the quiet, inherited belief that the traffic coming from your own services doesn't need to be treated with the same suspicion you'd apply to an anonymous request from the public internet. It does. The call is coming from inside the cluster.
Tagged: distributed systems, SRE, reliability, security, microservices