Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Terraform State Is a Database — Start Treating It Like One

A single glowing manila file folder on a marble pedestal at the centre of a 1960s-style control room, fine luminous threads radiating out to a miniature city of servers and load balancers that hang from it, one thread frayed and close to snapping — painted as a mid-century gouache magazine cover.

There is a file in most infrastructure repositories that can delete the company. Not metaphorically. A single corrupted, misedited, or maliciously truncated terraform.tfstate can, on the next apply, instruct your provider to tear down the production database, the load balancers in front of it, and the DNS records that would have let anyone notice before the pager went off. The file is usually a few hundred kilobytes of JSON. Most teams treat it with roughly the care they'd extend to a node_modules directory.

This is a category error, and it's worth naming precisely. Terraform state is not a build artifact. It's a database — a stateful, authoritative, mutable record of what exists in the world, against which every future operation is diffed. Once you make that mental swap, the entire discipline around it changes. You stop asking "where do I put this file" and start asking the questions you'd ask of any production datastore: who can write to it, how is it locked, how is it backed up, and what happens at three in the morning when it's wrong.

The mental-model swap: artefact vs database

A build artifact is the output of a deterministic process. It's reproducible by definition — lose it, and you rebuild it from source. The source is the truth; the artifact is a convenience. This is why we treat artifacts casually, and why that casualness is mostly correct. You can throw a Docker image away.

State is the opposite shape. It is not derived from your .tf files — it's derived from reality, mediated by every apply you've ever run, in order, including the ones run by the colleague who left in 2023 and the one you ran by accident against the wrong workspace. The .tf files describe intent. State records outcome. Lose the state and you have not lost a convenience; you have lost the only record of which of your intentions actually took, which resource IDs the provider handed back, and which of the forty security groups in the account belong to this configuration versus the three other configurations that also touch it.

You cannot rebuild state from source the way you rebuild an artifact. You can import, painstakingly, resource by resource, hoping you catch them all — which is recovery, not reproduction, and the distinction is the whole point.

Here is the mapping made explicit, because the parallels are exact rather than cute:

Database concept Terraform state equivalent
The database itself terraform.tfstate — authoritative record of real resources
Schema / table definitions Your .tf configuration (intent, not record)
A transaction A single terraform apply
Row-level / table locking State lock (DynamoDB, Postgres advisory lock, blob lease)
Point-in-time recovery Versioned state in S3 / equivalent object versioning
Backups terraform state pull snapshots, bucket replication
ALTER TABLE / migrations terraform state mv, import, rm — schema surgery
Credentials at rest Secrets stored in plaintext inside the state JSON
DROP DATABASE terraform destroy against the wrong workspace

Nobody runs a production database without locking, backups, access control, and a migration process they've rehearsed. The argument of this piece is simply that state earns the same treatment, and that most of the famous Terraform disasters are what you get when it doesn't.

Split panel: disposable build artifact on a conveyor belt versus a bank vault labelled STATE with one engineer holding the only key

Remote state, and local state as a loaded gun

Local state — a terraform.tfstate file sitting in your working directory, possibly committed to git, possibly not — is the configuration most tutorials start with and no team should finish with. It fails in two directions at once.

If you commit it, you've put your infrastructure's authoritative record, including any secrets it captured, into version history forever, distributed to every clone. If you don't commit it, the file lives on exactly one laptop, which means the truth of your production estate is one spilled coffee away from gone. There is no configuration of local state that is safe for more than one engineer, and barely safe for one.

Remote state — an S3 bucket, a GCS bucket, Terraform Cloud, a Postgres backend — solves the location problem, but the reason it matters runs deeper than "a shared place to put the file". Remote backends are where locking lives. They're where versioning lives. They're the difference between treating state as a file you happen to share and treating it as a service with semantics. The backend is not storage; it's the transaction coordinator. Pick it on that basis, not on which bucket was easiest to spin up.

The one rule I'll state without qualification: the bucket holding your state should not be managed by the state it holds. Bootstrapping the state backend with the same Terraform that uses it produces a circular dependency that is fine until the day you need to recover, at which point you discover that the thing you need to recover with is the thing you're trying to recover. Bootstrap the backend separately, by hand or with a tiny dedicated configuration, and never speak of it again.

Locking: the concurrency bug waiting to happen

Two engineers, one workspace, no lock. Engineer A runs apply against a plan generated thirty seconds ago. Engineer B, who hasn't pulled, runs apply against a plan generated from stale state. Both writes land. The state file now reflects neither engineer's intent — it reflects an interleaving, and Terraform's own consistency model assumes that never happens. This is a classic lost-update race, the same one any database undergraduate learns about, except the rows are subnets and the anomaly is a half-built VPC.

State locking is the fix, and it's the one piece of hygiene the major backends give you nearly for free. S3 with DynamoDB (or, on newer versions, S3-native conditional-write locking), GCS with its object generation semantics, Postgres with advisory locks — all of them serialise writes so that the second apply blocks until the first finishes, then re-reads fresh state. The mechanism varies; the guarantee is the same. Turn it on. There is no team size at which you don't need it, because "team of one" includes "you, plus the CI pipeline you forgot was also allowed to apply".

Here is the race, and the lock that prevents it:

Engineer BState (S3)Lock (DynamoDB)Engineer AEngineer BState (S3)Lock (DynamoDB)Engineer Ablocks, retries with backoffacquire lockgranted (LockID held by A)read stateacquire lockDENIED — held by Awrite new staterelease lockacquire lockgrantedread FRESH statewrite new staterelease lockEngineer BState (S3)Lock (DynamoDB)Engineer AEngineer BState (S3)Lock (DynamoDB)Engineer Ablocks, retries with backoffacquire lockgranted (LockID held by A)read stateacquire lockDENIED — held by Awrite new staterelease lockacquire lockgrantedread FRESH statewrite new staterelease lock

The failure mode to watch for is the orphaned lock — Engineer A's machine dies mid-apply, the lock is never released, and everyone else is blocked. terraform force-unlock <LOCK_ID> exists for exactly this, and it is precisely as dangerous as it sounds, because the reason the lock is still held might be that an apply is genuinely still in flight somewhere you can't see. Force-unlock is a decision, not a reflex. Confirm the holding process is actually dead before you break its lock, the same way you'd confirm a database session was dead before you killed it.

Backups, versioning, and the three-in-the-morning restore

Object versioning on the state bucket is the cheapest insurance in infrastructure and one of the most commonly skipped. With S3 versioning on, every write to state retains the prior version. A bad apply, a fat-fingered state rm, a corrupted write — all recoverable by promoting the previous object version back into place. Without it, the bad write is the only write, and you're into import-everything-by-hand territory at the worst possible hour.

Versioning is necessary but not sufficient, because it protects against bad contents, not against losing the bucket. Bucket-level disasters — accidental deletion, a region event, an IAM change that locks you out of your own backend — need a second copy somewhere the first blast radius doesn't reach. Cross-region replication, or a scheduled terraform state pull dumped to separate cold storage, gives you that. One copy is not a backup; it's a single point of failure you've chosen to feel good about.

Test the restore. I cannot stress this enough, and I say it knowing full well most readers will nod and not do it. A backup you have never restored is a hypothesis, not a backup. Pick a non-production workspace, corrupt its state deliberately, and walk the recovery — pull the prior version, validate it with terraform plan showing no spurious changes, promote it. Do this before the night you need it, because the night you need it is not when you want to discover that your versioning was enabled on the wrong bucket. Ask me how I know.

State versions as sediment layers with a corrupted layer in red and an arrow promoting the prior clean version

Secrets in state: the plaintext problem nobody mentions

This is the part most teams find out about the hard way. Terraform state stores the full attributes of every resource it manages, including the sensitive ones. An RDS instance's master password, the contents of a random_password, a generated TLS private key, the initial secret of a service account — if the provider returns it, it lands in state. In plaintext. The sensitive = true flag suppresses the value in CLI output and nothing else; it is a UI feature, not encryption, and treating it as the latter has burned people who genuinely believed they were protected.

So the state file is not merely operationally critical — it's a credential store, and usually an undeclared one. Which means everyone who can read the bucket can read your database passwords, and "everyone who can read the bucket" is, in too many accounts, a much larger set than anyone intended.

Three controls, in order of how much they help. First, encrypt the backend at rest with a key you control — SSE-KMS on the bucket, not the default SSE-S3, so that bucket read access alone isn't sufficient and the KMS grant becomes a second gate you can audit and revoke. Second, lock the IAM down to the precise principals that need state access and no others; state-bucket access is production-credential access and should be governed as such. Third, and most durably, stop putting the secrets in state in the first place — generate them out of band in a real secrets manager (Vault, OpenBao, AWS Secrets Manager) and have Terraform reference them rather than birth them. You can't leak from state what was never written to state, which is the only category of this problem that actually goes away rather than merely shrinking.

Migrations: state mv, import, and surgery without downtime

State drifts from configuration constantly. You rename a module and Terraform, seeing the old address gone and a new one arrived, plans to destroy and recreate — which for a database is not a refactor, it's an outage with extra steps. You adopt a resource someone created in the console and need to bring it under management without recreating it. You split one monolithic configuration into three. Every one of these is a schema migration, and Terraform gives you the surgical tools to do them without touching the underlying resources.

terraform state mv re-addresses a resource inside state — it tells Terraform "the thing you knew as aws_instance.web is now module.frontend.aws_instance.web", changing the record without changing the world. terraform import (or import blocks, on recent versions) pulls an existing real resource into state so Terraform stops wanting to create a duplicate of something that already exists. terraform state rm forgets a resource without destroying it, which is how you hand ownership of something to another configuration. These are ALTER TABLE for infrastructure: they change the record of what is, deliberately and reversibly, while production keeps serving traffic.

The discipline that makes this safe is the same discipline that makes database migrations safe. Take a state backup immediately before any surgery — terraform state pull > backup.tfstate takes two seconds and has saved me more times than I'll admit. Run terraform plan after every move and read it like a hostile witness: the only acceptable diff after a clean state mv is no diff. If the plan wants to create or destroy anything, you've addressed something wrong, and the right move is to restore the backup and think again rather than to apply and hope. Surgery on a live patient demands you check the patient is still breathing after each cut. Same principle, fewer lawsuits.

A state-hardening checklist

If you do nothing else from this piece, do these. They're ordered roughly by return on effort, and none of them is exotic — they're the boring controls you already apply to databases, pointed at the file you've been treating as disposable.

  • Use a remote backend. No local state, no state in git, no exceptions for "just this one small project" that will outlive your tenure.
  • Enable locking on that backend, and confirm it actually engages — run two applies at once in a test workspace and watch the second block.
  • Turn on object versioning on the state bucket. Every prior version is a recovery point you'll eventually want.
  • Replicate or back up off-region. One copy is a liability. Cross-region replication or scheduled state pull to cold storage.
  • Encrypt at rest with a customer-managed key. SSE-KMS, not the default. Bucket access should not equal plaintext-secret access.
  • Restrict IAM to named principals. Treat state access as production-credential access, because it is.
  • Keep secrets out of state where you can — generate in a secrets manager, reference rather than create.
  • Separate state per environment. Production state and dev state in the same file is one destroy in the wrong terminal away from a very bad afternoon.
  • Bootstrap the backend independently of the configuration that uses it. No circular recovery dependencies.
  • Back up before every state operation. terraform state pull before any mv, rm, or import. It costs nothing.
  • Rehearse the restore. A backup you've never restored is a guess. Walk it in anger, once, before you need it.
  • Gate destroy. Require explicit approval, scope down the IAM that can run it, and never let it fire unreviewed from CI.

Treat it like prod, because it is

The through-line is a single reframe with a long tail of consequences. The moment you stop seeing terraform.tfstate as a file and start seeing it as the production database it has been all along, the controls stop feeling like overhead and start feeling obvious — you lock a database, you back up a database, you don't leave a database's credentials readable by the whole estate, and you certainly don't let one stray command drop it.

OpenTofu inherits all of this unchanged; the fork didn't alter the shape of the problem, only the licence on the tool that creates it. State is still the most dangerous file in the repository, and it will remain so regardless of which binary you run. The teams that sleep through their on-call rotations are not the ones with the cleverest modules. They're the ones who worked out, usually after one genuinely frightening night, that the JSON was production all along.