Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

The Cloud Bill Is a Design Document: Reading Your Architecture in the Invoice

A mid-century engineer unrolling a long printed invoice across the floor, where its ruled rows become the architectural floor plan of the building he is standing in — 1960s gouache.

Your organisation has one document that accurately describes the system you are running. It isn't the architecture diagram, which describes the version that got approved. It isn't the wiki page, which describes the version somebody meant to build in 2023. It's the invoice.

Usage is metered into hourly buckets, resource by resource, and priced by the hour and the gigabyte. The invoice itself arrives monthly, and the detailed data behind it refreshes through the day. Nobody drafts any of it. Nobody can quietly leave out the bit they're embarrassed about. It never lags a refactor, because it is produced by metering rather than by memory. It is, in the most literal sense available, a description of your architecture — and in most companies the only person reading it closely works in finance and has no way to tell a NAT gateway from a hole in the ground.

Corey Quinn has spent years making a version of this point, usually with the observation that most of an AWS bill is data transfer and mistakes. The part that interests me is narrower: every surprising line item is evidence of a design decision, and quite often a decision nobody remembers taking.

This is the first of four pieces on that. The other three are about doing something with what you find: stopping the bill growing, cutting what's already there, and the architectural moves that change the shape of it. This one is about reading.

Cost is telemetry that nobody alerts on

Latency gets percentiles, dashboards, and a pager. Error rates get an SLO. Spend gets a monthly email to someone who cannot read it architecturally, and a quarterly meeting where an arbitrary percentage is demanded back.

That's an odd gap, because cost has properties most signals would envy. It is complete: everything that ran is in there, including the things you forgot. It is attributable down to the resource and the hour for most services, with a residue that stubbornly isn't. It is retained for a year or more without anyone having to fund a retention policy. And it is causally dense — a line item that moves without a corresponding deploy is telling you about behaviour that changed underneath you.

The reason it doesn't get treated as telemetry is that its native vocabulary is procurement rather than engineering. "EC2 — Other" is not a sentence about your system. Underneath that label, though, the data is exact, and it is saying something specific.

Where the meter runs

The single most useful mental shift is to stop reading the bill as a list of services and start reading it as a map of where traffic crosses a boundary someone is paid to operate. Most surprises live at those boundaries.

internet egress, perGBcross-zone, chargedboth directionsNAT hourly + per GBprocessedper request, permetricidle outside thepeakUserLoad balancerApp instance, zone ADatabase replica,zone BObject storageMetrics and logsProvisioned capacityinternet egress, perGBcross-zone, chargedboth directionsNAT hourly + per GBprocessedper request, permetricidle outside thepeakUserLoad balancerApp instance, zone ADatabase replica,zone BObject storageMetrics and logsProvisioned capacity

Nothing in that diagram is exotic. It is a perfectly ordinary three-tier application, and it has five separate meters running on a single request path. The architecture diagram for the same system has four boxes and no meters at all, which is the difference between the two documents.

Once you know the boundaries, the line items translate:

The line item What it's telling you Where to look first
NAT gateway data processing Private subnets reaching AWS services through the gateway, so same-region object-store traffic that is itself free still attracts a per-gigabyte processing charge Missing gateway endpoints for S3 and DynamoDB
Cross-zone data transfer A chatty path that crosses an availability zone, charged in both directions Replicas, brokers, and service-to-service calls spread across zones for HA
Internet egress Responses larger than anyone needs, or an integration partner polling you Payload sizes, missing CDN, a partner on a one-second timer
Idle provisioned capacity Capacity bought for a peak that has since moved, or autoscaling that scales out and never back in Utilisation at 04:00 versus the scaling policy's cooldown
Per-request charges on storage and metrics Chatty clients, absent caching, and metric cardinality The top ten callers by request count, not by bytes
Public IPv4 addresses A census of every network interface you forgot you had, now billed per hour since AWS started charging for all of them in 2024 Orphaned interfaces, per-task public IPs, old load balancers
Log and trace retention Defaults nobody chose, retained because nobody set a policy Whether anyone has queried data older than a fortnight

Each row is a design decision wearing an accountant's clothes. Putting the database replica in another zone is a deliberate availability choice; the cross-zone line is its running cost, stated honestly, in a way the design review never did.

Read usage types, not service names

Practically: the Cost and Usage Report is the data, and everything else is a view of it. Configure it for hourly granularity with resource IDs included, because neither is the default and both are what make this analysis possible, and expect a residue of line items that carry no resource ID at all. What you get is usage types that name the mechanism rather than the product — NatGateway-Bytes, DataTransfer-Regional-Bytes, and the rest. Group by usage type before you group by anything else. The service name tells you which team at AWS invoices you; the usage type tells you what your system did.

The first pass takes an afternoon and is almost always the same shape: a handful of usage types make up most of the bill, and at least one of them will be something nobody in the room can immediately explain. That one is the interesting one. It's the part of your architecture that has drifted furthest from the version in everyone's head.

Attribution is the boring prerequisite

You cannot act on any of this without knowing whose it is. The bill arrives as one number for the whole estate, and turning it into per-team, per-feature, or per-tenant numbers depends on tagging that was either done at creation or never done at all.

Two things about tags catch people out, and the first is a trap that has partly closed. Activating a cost allocation tag key used to apply only from the moment you activated it; since 2024 you can request a backfill of up to twelve months instead. What a backfill cannot do is invent a tag that was never applied — it resurfaces tags that were already sitting on the resources, unactivated, at the time.

Which leads to the second: untagged is the default state of anything created in a hurry, so untagged spend correlates precisely with the incidents and experiments you most want to account for, and no amount of retroactive activation rescues it. The tag has to be applied at creation. The activation can wait.

The alternative to attribution is an across-the-board cut, where every team gives back the same percentage regardless of what they're spending it on. That reliably removes the cheapest useful thing and leaves the expensive pointless thing untouched, because the expensive pointless thing belongs to whoever argued hardest.

Unit economics, or the number that actually matters

Absolute spend is close to meaningless on its own. The number worth tracking is cost per unit of whatever your business counts: per thousand requests, per tenant, per job, per gigabyte ingested.

It reframes the conversation immediately. Spend rising with flat unit cost is a sales success and nobody's problem. Spend flat with rising unit cost is decay, and you have a few months' warning before it becomes visible. Unit cost rising as you scale is the loudest signal of the three: superlinear cost means coordination overhead you designed in — chatter that crosses zones, fan-out that multiplies, retries that amplify.

That last one has a useful property. A retry storm shows up on the bill before it shows up in your SLO, because the extra attempts are billable long before they are numerous enough to hurt latency. The signal is early; its delivery is not, so expect to read it a day late rather than a minute late. If you've wired retry budgets, the bill is still the cheapest instrument you have for telling whether they're holding.

When this becomes a false economy

Cost work goes wrong in three predictable ways, and it's worth naming them before the series talks about cutting anything.

The first is spending three weeks of engineering time to save two hundred pounds a month, then declaring victory. The maths on that takes about a year to break even, assuming the saving survives the next architecture change, which it usually doesn't.

The second is buying a reliability problem. Collapsing to a single availability zone genuinely removes the cross-zone line item, and it also removes the property that line item was paying for. Anything that reduces the bill by reducing redundancy is a trade, and should be argued as a trade rather than presented as an optimisation.

The third is locking in today's architecture. Reserved capacity and savings plans are real money, and also a one- or three-year bet that you'll still be running roughly this shape of thing. Take the longer term for the deeper discount and the refactor that would have halved the bill becomes the change nobody can afford to make.

The rule I use: optimise where the fix also simplifies the system, and be suspicious of any fix that adds a constraint. Deleting an unused NAT gateway makes the diagram smaller. Adding a caching layer to dodge per-request charges makes it larger, and now you own a cache.

What the rest of the series covers

Stopping the growth. The controls that make next quarter's bill predictable: budgets that alert on the derivative rather than the total, anomaly detection that fires on the day rather than the month, and the guardrails that stop an experiment becoming a standing charge.

Cutting what's there, and right-sizing honestly. Instance and container sizing derived from measured behaviour rather than from the number in the tutorial, which means testing real workloads rather than guessing from averages.

Architecture that changes the shape. The structural moves: what to push to the edge, what to cache, and the one most teams overlook — with thousands of clients and a handful of servers, work moved to the client is capacity you don't buy. That comes with caveats sharp enough to need their own section, starting with the fact that you can never trust what the client tells you afterwards.

For now, the exercise is smaller. Pull last month's usage types, sort by cost, and find the first line nobody can explain. That line is a design decision you've been paying for and haven't reviewed. Read the bill like a log.