Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Scaling Feature Switches

A hand lifts old illustrated files – a server rack, a laptop, a network diagram and a rising chart – from the open drawer of a battered green filing cabinet in a sunlit study, 1960s gouache.

Originally published on my old blog on 17 November 2013. Lightly edited; notes added in 2026 are marked.

I was recently helping a friend's company to scale out. This covers some of the issues we hit when scaling the code behind their feature switches.

Once your applications get big, being able to toggle feature switches on and off – to enable and disable features, or change behaviour in real time – is a great advantage.

If you are getting 10 page requests per second on each of two servers and you have 10 feature blocks on each page, then 200 lookups per second can easily be served by Redis or Memcache directly. The per-page penalty shouldn't be massive if the key/value server isn't too far removed from your application servers – typically 10 ms, assuming my benchmarks are in the right ballpark (60,000 GETs per second and around 1 ms round-trip time).

This becomes a bigger issue when you have many more servers, many more feature blocks and many more pages per second. The Redis server becomes a bottleneck, and the round-trip time means you need many more threads to do the same work.

To this end, I redesigned the standard "gatekeeper" code they used so that it would scale more effectively.

They used to use APC for opcode caching and now use XCache. Both support storing variables in memory for blazingly fast access. That gives great speed, but it is a pain to set those values from outside the web server – from the CLI, for instance.

So what we ended up doing was using the local cache to store the feature flags for a short time after fetching them from the key/value store (Redis, in this case). With a 10-second cache time and 100 requests per second, this cut the number of requests hitting Redis by three orders of magnitude.

The second issue was that, as the application was scaled out across multiple data centres, we had multiple Redis servers for the feature flags – one per cluster. Distributing changes to the flags was a bit of a pain. We settled on a RabbitMQ fanout exchange, with a worker sitting on each Redis server listening for global updates and pushing the changes into Redis. This was the most fragile part of the setup, but it worked well for us.

RabbitMQ supports server-side keep-alives, but our worker code was a pain to change to support client-side keep-alives. This meant a network partition would see RabbitMQ close the connection without the client being aware of it.

A second worker was used to check on the first (and on other processes), and restarted it if it hadn't responded for a while, either to a heartbeat or to a real message.

The result was a very scalable solution without a massive amount of overhead.

2026 note: The shape here – evaluate flags locally against a short-lived cached copy, and push changes out rather than looking each one up per request – is how the commercial and open-source flag services (LaunchDarkly, Unleash, flagd behind OpenFeature) work today. APC's user cache lives on as APCu; XCache is long gone.

The half-open connection problem is much rarer now: AMQP heartbeats are negotiated by default (RabbitMQ proposes 60 seconds), and modern clients honour them. A worker blocked in long-running work can still miss its heartbeats, though, so the watchdog still earns its keep. More on running RabbitMQ from Python in RabbitMQ Without Celery.