Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Case Study: Optimising a Cloud Application

A hand lifts old illustrated files – a server rack, a laptop, a network diagram and a rising chart – from the open drawer of a battered green filing cabinet in a sunlit study, 1960s gouache.

Originally published on my old blog on 24 November 2011. Lightly edited; notes added in 2026 are marked.

I was recently brought in to examine the infrastructure of a small startup. That in itself wasn't anything special – I do it quite often, for various reasons. What was different was that they didn't have a problem scaling out; they had that working well with their shared-nothing web application and MongoDB backend. What they had a problem with was their infrastructure costs.

I normally work through a six-step process that I've built up over time:

  1. Monitor and gather stats; measure the problem.
  2. Standardise on a reference architecture.
  3. Add configuration management and version control.
  4. Start to define a playbook for how to do things (scaling up and down, provisioning new machines and clusters) and start to automate them.
  5. Bring everything to the reference architecture; consolidate underutilised servers and eliminate unused infrastructure.
  6. Consider architecture changes to make it more efficient.
  7. …and repeat.

I'll take you through a case study showing how this process was used to lower their monthly costs. Names and details have been changed in places to protect the guilty. ;)

Monitor and gather stats; measure the problem

Unless you know what the servers are doing, it is hard to know what to adjust or fix when things go wrong. Proper instrumentation is essential for any non-trivial application.

The architecture roughly consisted of:

  • 13 web/PHP servers
  • 3 MongoDB servers (a replica set)
  • 2 MySQL servers in a master-master configuration
  • A log/Nagios server
  • An email server
  • 2 Varnish load balancers

The biggest problem was that they just didn't know what their infrastructure was doing. They had server and service monitoring with Nagios, but beyond knowing whether a service was up or down, they had no metrics or stats from the components in their architecture. We installed Munin to correct this. It isn't ideal, but it is a pretty good starting point for looking at historical data for trends and patterns of usage.

What we found was what I'd suspected at the start: they could scale out really well, but much of their infrastructure was running at an average of only 10–15% utilisation. Some of this was the normal slack after the application's peak times. It was aimed at business users, so from around 10am to 12pm and 2pm to 4pm usage would peak at hundreds of times the load it saw at 2am, their quietest time. That's fairly typical for modern SaaS apps and wasn't a massive surprise.

As I said, the team was pretty savvy and had largely nailed scaling out by choosing software (Varnish, PHP and MongoDB in this case) that could scale horizontally without too much trouble. So I was surprised, given how well they'd implemented the rest of the architecture, to find no obvious cache. Memcache is usual, and found in probably 90% of large-scale cloud apps.

Talking it through with one of the developers, we found that Memcache had originally been specified, but the team manager at the time shot it down in flames because of a previous bad experience. Instead, the developers had worked around it by using APC (the PHP opcode cache extension) as a per-node cache. That wasn't particularly efficient given how the app was designed, so the developers had also added session affinity to the Varnish configuration, to make sure a user's data was only stored on a single node, rather than updating the database on every request (which is what they had ended up doing with their sessions). Storing sessions in the database had also brought MySQL into the mix when it really wasn't needed.

Personally, I feel you should go through all the steps before making architecture changes, but one of the developers set up two Memcache clusters (one for sessions, one for general caching) and rewrote the caching section of the app in a few hours, so we rolled with it and started again at step 1…

Monitor and gather stats; measure the problem (again)

The disadvantage of storing sessions in Memcache is that if one of the Memcache servers goes offline, or starts to run out of memory, it will log users out. This is a BadThing™.

The architecture now roughly consisted of:

  • 13 web/PHP servers
  • 3 MongoDB servers (a replica set)
  • 4 Memcache servers (2 pairs)
  • A log/Nagios server
  • An email server
  • 2 Varnish load balancers

Adding Memcache actually made the utilisation far worse: the next day it was down to 7%. The load from keeping sessions in the database had, it seems, been roughly halving the overall efficiency of the servers.

We added Memcache monitoring to Munin, set up some alerts in Nagios for the new servers, and moved on to the next stage.

Standardise on a reference architecture

Having dozens of separate hand-rolled configurations is a maintenance nightmare. Make all boxes as similar as possible, except where they need to differ. Take a standard image and install only the software needed for that function – no more, no less.

Many times I've arrived at a company and found that while all the servers might be running the same OS (Ubuntu, for example), they are at various patch levels, and sometimes one or two versions behind. This time it wasn't quite that bad, but it wasn't far off.

All the servers were running Debian Squeeze – nothing wrong with that – and most had similar configurations… but not all. Many of the boxes had been built when they started writing the application and had excess packages and configuration-file backups all over the place. Two of the systems had also been set up with more memory than the others, from when someone was testing something. It's a pretty common situation, and not too bad.

We took a clone of each type of system and pared it down to the essentials, then started to document exactly what was needed for each role: which repos, how much memory, and which configuration files needed editing from the defaults. We added this to the playbook – a document that tells you exactly what to do, and in what order, to carry out a given task; in this case, provisioning a new machine from the default image.

Add configuration management and version control

In this case the developers didn't feel it was the right time to add a configuration management system to their architecture. I argued against that, but the client is always right… (supposedly). I did get them to agree to use version control (git, in this case) to keep the changing versions of each machine's config files. That lets the developers roll back changes when they need to, and see what has changed between now and some point in the past. Not ideal, but you take the victories you can get.

Start to define a playbook and automate it

The playbook isn't actually a stage; it is a document that runs alongside the day-to-day running of the systems. It exists so everyone knows the correct way to do the tasks needed to create, maintain and repair your systems. It can be a wiki (as it was in this case), a file share with text documents on it, or even an actual book (a grimoire).

We'd already made notes on how to take a new machine from bare image to working server, and we kept adding notes throughout the day. We ended up with about 20 separate tasks that are needed on a frequent basis to keep the systems running properly.

Bring everything to the reference architecture; consolidate and eliminate

One by one, we added new servers to replace the few that had been built with too much memory (the equivalent of going from large to small on Amazon – not that they were on Amazon) and swapped them out. For availability reasons we couldn't simply remove them outright. As a rule of thumb, you want enough capacity for peak times plus one spare server, and that's what they had at this point. Even without increasing average utilisation much, we cut their monthly costs by around $190 per server for those two servers.

We also noticed that some servers really were running nothing most of the time. They were simply eating money for no real return. The few tasks they did run were moved to a single small instance and the machines were terminated. In total around eight machines were removed, which together cut monthly costs by around $500.

Total monthly savings from all the changes so far: around $700… and as they grow, the savings will grow too.

Consider architecture changes to make it more efficient

They had already added Memcache to reduce the number of queries hitting the database, but peak load was still the main problem. Some companies follow best practice and move all their static images to a fast image server, or onto a CDN. Neither had been done with this application.

The other thing you can do with static files is give them a nice long expiry time, so that browsers and intermediate proxies keep them around for a while and don't have to fetch them every time someone visits the page. The default Apache install they were using at the time does a reasonable job here: while it doesn't set an Expires header unless you tell it to, it does send an entity tag, which the browser can use to ask whether the file has changed since it last fetched it. That saves bandwidth, but the browser still has to ask the server, which on slow links can take a few hundred milliseconds.

They already had analytics on their pages, and together with their server logs it was easy to see that image and script loading accounted for around 30% of their bandwidth costs. Their pages were importing six or seven scripts on top of any static images. Four of those were available on the Google CDN, which saved around 6% of their bandwidth. The nice side effect was that pages also loaded faster, since most browsers limit the number of parallel requests they make to one host. With the scripts coming from Google's servers (and being cacheable), they were often already in the browser's cache, or could be fetched as soon as the page loaded, alongside any scripts or images from the application's own servers.

2026 note: Don't do this now. Since around 2020 the major browsers partition their caches by site, so a script fetched from a shared public CDN on someone else's site is no longer in the cache when your page asks for it. The shared-cache benefit has gone, and what's left is a third-party dependency, a privacy leak and a supply-chain risk. Serve your own assets with content-hashed file names and long cache lifetimes; if you must load from a third party, use Subresource Integrity.

Most of the other static files were moved to a separate pair of static servers running nginx. Once the Expires header was set to max, most static files were only loaded once across the whole site. With the static images gone from the PHP servers, much of the peak load went with them, and the number of PHP servers came down from 13 (12 plus a spare) to 8 – which made up for the two extra servers they'd added to the architecture. They could just as easily have pushed all the static files to a service such as S3 or Rackspace Cloud Files, but they believed that was too much of a risk. It isn't.

One of their problems at peak times was that a request would often fire off a function that consumed a lot of processing power for a couple of tenths of a second. When many of those requests arrived at a single server in a short space of time, the server would slow down appreciably. These processes didn't need to run at request time – often the data wouldn't be used for a few minutes, and sometimes never. With that in mind, I suggested adding Gearman to the mix. Gearman is a priority-based job queue that, in this situation, would let the request drop the data on a queue to be picked up by a worker later. That would have flattened many of the peaks and reduced the number of servers needed to cover peak periods. Alas, this was vetoed too, and we couldn't reduce the number of servers any further.

2026 note: Gearman has largely faded; today I'd reach for RabbitMQ or even a Postgres-backed queue. The principle hasn't changed – see Not Everything Needs to Be Real-Time.

If they hadn't rejected configuration management outright, bringing a new PHP server up would have been pretty simple. Add it to the config database and turn it on; the server installs the required packages and pulls its config from the configuration system; Nagios notices the server is up and runs a second check from the notification script; on success it is added to the Varnish config, Varnish is triggered to pull its config, and load starts hitting the server. That means once the peak has passed you can simply turn a server off to save money, and conversely, when load approaches a threshold, you start another one up. I personally dislike the idea of autoscaling databases and similar things, since there is more possibility of data loss if something goes wrong.

2026 note: Configuration management, immutable images and autoscaling groups are table stakes now, so the tooling in this piece has dated. The order of the steps hasn't: measure first, standardise, write it down, and only then start changing the architecture.