Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Your Gateway is Both Shield and Single Point of Failure

A lone fortified stone gateway on a plain with every road converging through its single arch and a hairline crack running up one pillar — both shield and chokepoint, painted as a 1960s gouache magazine cover.

There is a particular kind of architectural decision that looks unambiguously correct right up until the moment it isn't. The API gateway is one of them. You introduce it to centralise control — authentication, rate limiting, routing, observability all in one place — and for a while it does exactly what it says on the tin. Then, at some point, you notice that the thing you put in front of everything is now the thing everything depends on. The abstraction has become a chokepoint, and the blast radius of any failure inside it is, by design, total.

This is not an argument against API gateways. It is an argument for thinking more carefully about what you are trading when you deploy one.


What You're Actually Buying

The canonical pitch for an API gateway covers auth, rate limiting, and observability. These are real benefits, and it is worth being specific about why they matter rather than treating them as self-evident.

Authentication and authorisation at the gateway gives you a single enforcement point. Rather than every upstream service implementing its own token validation logic — with varying degrees of correctness, freshness, and error handling — you centralise that concern. A JWT checked once, consistently, at the ingress point is substantially easier to audit than the same check performed twelve different ways across twelve different services. Mutual TLS between internal services is a different problem, but edge authentication is a natural gateway responsibility and one it generally handles well.

Rate limiting at scale is harder than it looks. A naive per-instance rate limiter is wrong the moment you run more than one gateway instance, because the counters are local. Shared state — typically Redis — solves the accuracy problem but introduces its own availability coupling: if your rate limit store becomes unavailable, you either fail open (and remove the protection) or fail closed (and take down your API). Neither option is comfortable at three in the morning. Still, centralised rate limiting is far more tractable than pushing that responsibility to every service individually, and it gives you a coherent surface for adjusting limits when a client starts misbehaving.

Observability is where the gateway genuinely earns its keep. A single request traversal log, latency histogram, and error rate dashboard — all produced from one place, in a consistent format, without requiring every upstream service to instrument itself identically — is enormously valuable. It is also, incidentally, much easier to trust than telemetry aggregated from N services with N slightly different interpretations of what counts as an error.

So the benefits are real. They are also, to a first approximation, the benefits of centralisation in general: consistency, reduced duplication, a single place to look. The costs follow from exactly the same source.


The Failure Taxonomy

API gateway failure modes split roughly into three categories: availability, configuration, and amplification. The first is obvious; the second is underestimated; the third is the one that causes post-mortems.

Availability is the straightforward case. If the gateway is the only path to your services and it goes away, everything goes away. This is manageable — multiple instances across availability zones, a load balancer in front, health checks removing unhealthy nodes — but it requires active effort and ongoing discipline to maintain. It is not free, and the discipline has a habit of eroding between incidents. The target of 99.99% uptime sounds achievable until you start counting the ways a deployment pipeline, a configuration change, or a noisy neighbour can eat into that budget.

The June 2023 AWS us-east-1 incident is instructive. A latent defect in Lambda's capacity management subsystem activated when usage crossed a threshold that had never been reached in production before. API Gateway, Lambda, and over a hundred other services degraded simultaneously, because the shared internal dependency failed under novel load. Your own gateway has analogous structural dependencies: a backing Redis cluster for distributed state, an external identity provider for token validation, a service registry for upstream discovery. Any of them can become your capacity management subsystem.

Configuration failures are subtler and arguably more dangerous, because they do not look like failures immediately. A misconfigured timeout hierarchy — where the gateway timeout exceeds the client timeout, so the gateway continues processing requests the client has already abandoned — wastes resources at exactly the moment you can least afford to. A bad routing rule sends traffic to the wrong upstream; if the wrong upstream happens to accept the request format, you may not notice until a consumer raises a ticket. A certificate renewal handled correctly by cert-manager but not signalled correctly to the gateway controller results in TLS errors that look, from the outside, like the downstream service is broken.

NGINX-based ingress controllers are particularly prone to configuration-churn problems. Each configuration reload is fast — under 100ms — but at 50 to 100 reloads an hour, driven by aggressive autoscaling or a busy deployment pipeline, the cumulative tail latency impact on persistent connections is real and surprisingly hard to attribute. Envoy-based controllers avoid the reload problem via xDS dynamic configuration, but they add their own operational complexity and their own failure surfaces. There is no version of this that is entirely free.

Amplification is the failure mode that tends to generate the interesting post-mortems. The gateway's job is to centralise traffic. When something goes wrong downstream — a service slows down, a database starts timing out, a dependency becomes flaky — the gateway retries on your behalf. Multiple retries. Across all clients. Simultaneously. A slow backend that would, in a direct-call model, only affect the clients calling it directly instead receives a wave of retried requests from every client in the system, mediated by a gateway that is trying to be helpful. The circuit breaker pattern exists precisely to interrupt this feedback loop, but it requires correct configuration of thresholds, half-open states, and what constitutes a recoverable vs. permanent failure — configuration that is easy to get subtly wrong and hard to test until the failure mode actually presents.

Timeout ordering is worth stating explicitly because it is frequently wrong in real deployments. If the client timeout is 10 seconds, the gateway timeout should be 8, and the backend timeout 5. If the gateway timeout is longer than the client timeout, the gateway is doing work for requests the client has already given up on. This sounds obvious; it is not obvious enough, given how often you encounter the reverse configuration in practice.


The Latency Budget Problem

Every hop costs money. Not money-money — though at sufficient scale, yes, also that — but latency budget. A gateway processing request at 10,000 RPS with 10ms per-request overhead is not a trivial infrastructure item. At P99, that overhead compounds with backend latency, network variance, and any shared-state operations the gateway needs to perform (rate limit checks, auth token validation, plugin pipelines). The difference between a gateway that adds 5ms and one that adds 15ms is invisible in a smoke test and significant at scale.

The architecture of what runs your traffic matters more than any feature comparison matrix suggests. Plugin pipelines that execute sequentially for every request — authentication, transformation, logging, rate limiting, routing — add latency in series. The order of those plugins is not neutral. An auth check that fails early is cheaper than one that runs after a response transformation. A logging plugin that buffers asynchronously is cheaper than one that writes synchronously to a remote store. These are not gotchas; they are design decisions that affect your tail latency at scale, and they tend not to appear in the vendor's benchmark figures.

Caching at the gateway can recover some of this budget, but it is architecturally awkward. A cache at the edge is easy to reason about for idempotent reads; it is considerably harder to reason about for mutation-heavy APIs, versioned contracts, or responses that vary by caller identity. Stale data served from a gateway cache is a category of bug that is genuinely difficult to reproduce in staging.


Centralised vs. Decentralised Models

The centralism of the traditional API gateway is a design choice, not a law of physics. The service mesh model distributes the cross-cutting concerns differently: instead of a central proxy through which all traffic passes, a sidecar proxy runs alongside every service instance, intercepting and managing communication locally. Istio and Linkerd are the obvious exemplars; Envoy does the heavy lifting in both.

The difference matters in a failure scenario. A sidecar failure takes down the service instance it is colocated with. A gateway failure takes down everything. The blast radius of a central gateway misconfiguration is the whole platform; the blast radius of a misconfigured sidecar policy is the service that carries it. This is a meaningful distinction if your reliability goal is "most things keep working when something goes wrong" rather than "nothing goes wrong".

The sidecar model has its own costs. You now have proxies at every service instance, which means CPU and memory overhead that scales with pod count rather than traffic volume. Kubernetes 1.28 brought sidecars in as first-class citizens, which removed some of the lifecycle awkwardness, but the resource maths still adds up uncomfortably in large clusters. Istio's ambient mode — generally available since 1.24 — attempts a middle path: a lightweight L4 node-level proxy (ztunnel) handling mTLS and basic traffic, with optional L7 waypoint proxies for services that need it. The headline claim is up to 92% lower resource cost compared to the full sidecar deployment, which is the kind of figure that deserves scrutiny in your specific workload rather than acceptance at face value.

The honest answer is that the centralised gateway and the service mesh solve partially overlapping but genuinely distinct problems. The gateway is good at the edge: managing external consumer identity, enforcing API contracts, providing a stable public surface over a volatile internal topology. The mesh is good at the interior: mTLS everywhere, consistent retry and circuit-breaker behaviour across service calls, deep telemetry without application-level instrumentation. Running both is not redundancy; it is recognising that north-south and east-west traffic have different enough characteristics to warrant different tooling. This conclusion annoys people who would prefer a single answer, but the annoyance does not make it less correct.


Designing for Graceful Degradation

The gateway can fail in ways that are complete (everything is down), partial (some routes fail), or degraded (everything is slow). The third mode is the most insidious because it often evades alerting until users are already suffering. Designing for graceful degradation means deciding in advance how each of those modes should behave, rather than discovering the decisions retrospectively in an incident.

Circuit breakers are the obvious starting point. When an upstream service starts returning errors or timing out above a threshold, the circuit breaker opens, and the gateway stops forwarding requests to that upstream for a period. Clients get a fast failure rather than a slow one; the upstream gets relief; and the gateway avoids holding threads open on connections that are not going anywhere. The pattern is well understood and well supported by most gateway implementations. The implementation detail that matters is what happens at the half-open state, when the circuit tentatively closes to test recovery. A half-open probe that sends too much traffic can re-trigger the failure. A probe that sends too little may take too long to confirm recovery. Hystrix, which popularised the pattern, demonstrated roughly 70% reduction in outage duration from correct circuit-breaker usage; the catch is that the configuration is nowhere near as self-tuning as the marketing suggests.

Fallback responses — returning a cached result, a degraded payload, or a sensible default when the upstream is unavailable — are valuable but easy to implement incorrectly. Fallback code paths are infrequently exercised in testing, which means latent bugs in them tend to surface under the exact conditions — high load, partial failures, unusual request patterns — where you least want a surprise. If you have fallback logic, exercise it regularly in production, even if only against a small fraction of traffic.

Rate limiting and load shedding at the gateway are related but distinct. Rate limiting controls the rate at which individual clients can consume resources; load shedding controls the total load the system will accept, regardless of client identity. A gateway that rate-limits correctly but does not shed load under overload will still collapse if aggregate demand exceeds capacity. Priority queuing — accepting some request classes over others when capacity is constrained — is operationally complex but substantially better than the uniform degradation that results from treating all requests as equally important when the system is on its knees.

Configuration deployment deserves separate attention. Bad routing configuration — sending traffic to the wrong upstream, or creating a routing loop — is a self-inflicted failure mode that has ended many an on-call rotation badly. Configuration validation before deployment, canary rollouts for configuration changes, and automatic rollback on error-rate spikes are the obvious mitigations. Version control and mandatory code review for gateway configuration is less obvious but equally important; it is remarkable how many organisations treat their gateway config as operational state rather than infrastructure-as-code, right up until they need to explain to an incident commander how a route ended up pointing at a service that was deprecated six months ago.


The Trade-Off, Honestly Stated

The API gateway centralises control and concentrates risk. These are the same thing stated differently. The control gives you consistent auth, coherent observability, and a stable external surface. The risk means that when things go wrong inside the gateway — or inside the shared infrastructure it depends on — they tend to go wrong everywhere at once.

The response to this is not to abandon the gateway model. It is to be clear-eyed about what you are buying, to invest in the resilience patterns that reduce the blast radius of gateway failures, and to avoid adding capability to the gateway beyond what it needs to do its job. Every plugin you add is another thing that can fail. Every piece of shared state it acquires is another dependency that can go down at an inopportune moment. The gateway is not the right place to solve problems that belong to the services behind it, and it tends to accumulate exactly that kind of scope over time if you let it.

The question to ask periodically is not "what can the gateway do?" but "what happens when the gateway doesn't?" If the answer is "everything stops", you have not finished the architecture yet.