Key Management: The Bit That Actually Bites You

If the previous article was about the mathematics of keeping secrets, this one is about the part where humans get involved and everything gets complicated.
The algorithms are fine. AES hasn’t been broken. RSA, used correctly, is sound. ECC is solid. The mathematics underlying modern cryptography is the most reliable part of the entire stack, which is a slightly unsettling thing to say given what comes next.
Because the moment you move from “the algorithm” to “the system,” you’re dealing with institutions, humans, operational processes, legacy infrastructure, and incentive structures that don’t always point in the direction of security. This is where things go wrong. This is where they’ve gone wrong, repeatedly, with consequences ranging from embarrassing to catastrophic.
Let’s start at the foundations.
Public Key Infrastructure: Trusting the Keys
Cast your mind back to the previous article. Asymmetric cryptography lets two strangers exchange a shared secret over an open channel. But it has a silent assumption baked in: how do you know the public key you received actually belongs to who you think it does?
If I want to send you a secret message and I ask for your public key, what stops a third party from intercepting that request, substituting their public key, and reading everything I send? I encrypt with their key, thinking it’s yours. They decrypt, read it, re-encrypt with your actual key, and forward it on. Neither of us is any the wiser. This is a man-in-the-middle attack, and it’s the reason public keys need to be authenticated, not just distributed.
The dominant answer to this on the public internet is Public Key Infrastructure (PKI): a hierarchy of trusted parties who vouch for the relationship between a public key and an identity.
Certificate Authorities and the Chain of Trust
A Certificate Authority** (CA)** is an organisation that issues digital certificates. A certificate is essentially a signed statement: “I, the CA, attest that this public key belongs to this domain/organisation.” Your browser receives the certificate, checks the CA’s signature, and if it trusts the CA, it trusts the binding.
Simple enough in theory. In practice, you don’t want a flat list of a few root CAs signing everything. The root CA’s private key is the crown jewel — if it’s compromised, every certificate it’s ever signed is suspect. So in practice, root CAs almost never sign end-entity certificates directly. Instead, they sign intermediate CA certificates, which sign the actual certificates. The root key can then be kept offline, in a vault, on hardware that has never touched a network.
What you end up with is a certificate chain:
Root CA (offline, extremely well guarded) └── Intermediate CA (operational, signs end-entity certs) └── Your certificate (what the server presents)
When a browser connects to your server, it receives your certificate and needs to walk this chain back up to a root it trusts. This is why servers should send their intermediate certificate along with their own — a misconfigured server that only sends the leaf certificate will fail for clients that don’t already have the intermediate cached. This is one of the most common TLS misconfiguration errors in the wild, and it tends to surface at 2am after a cert renewal.
The chain can have multiple intermediate levels. Large CAs often have several, each with a different purpose or geographic scope.
The Root Store: The Most Important List You’ve Never Heard Of
Here’s a question worth sitting with: who decides which CAs your browser trusts?
Microsoft. Apple. Mozilla. Google. Each of them maintains a root store — a list of CA certificates that their software ships with and implicitly trusts. If your CA is in the root store, browsers trust your certificates. If it isn’t, users get a scary warning.
This is a significant concentration of power. A handful of private companies effectively control which entities can issue trusted certificates for the entire public internet. There’s a CA/Browser Forum (CA/B Forum) that sets baseline requirements for CAs — audit requirements, practices, what they’re allowed to issue — but the real enforcement mechanism is that Microsoft or Mozilla can remove a CA from their root store, which is functionally a death sentence for that CA’s business.
This has happened. We’ll come back to that.
Which raises an uncomfortable question that a Mozilla community member once made concrete in the most pointed way possible: in 2015, an application was submitted to Mozilla's root inclusion programme on behalf of "Honest Achmed's Used Cars and Certificates." It was satirical, but it was a real submission, and it made a real point and is worth reading at least once. The requirements for root inclusion are documented and auditable — but they don't ask whether you should be in the business of vouching for the identity of every domain on the internet. If Honest Achmed's could meet the baseline requirements and pass the required audits, there's no formal mechanism that would stop them being trusted by every browser to issue certificates for any domain on earth. The root store is a policy list maintained by humans, not a mathematical guarantee.
DigiNotar: A Cautionary Tale
In 2011, a Dutch Certificate Authority called DigiNotar was compromised. The attacker — widely attributed to Iranian state actors — used the access to issue fraudulent certificates for Google, Mozilla, and a number of other high-value domains. These certificates were valid. Browsers trusted them. They could be used to intercept traffic from users in Iran, which is exactly what appears to have happened.
The certificate for google.com, signed by a trusted CA, belonged to someone who was not Google.
Once the breach was discovered, the response was swift and terminal. Microsoft, Mozilla, and Apple all removed DigiNotar from their root stores. Every certificate DigiNotar had ever issued became untrusted simultaneously. DigiNotar — which also handled certificates for the Dutch government — was bankrupt within weeks.
The DigiNotar incident crystallised several things the security community had worried about for years: the CA system is only as strong as its weakest member, a rogue or compromised CA can issue valid-looking certificates for any domain, and the consequences when this goes wrong are severe.
It also prompted the development of countermeasures, which we’ll get to shortly.The DigiNotar failure was a dramatic external compromise.
The WoSign/StartCom case is more insidious — an inside job hidden in plain sight.
WoSign, a Chinese CA, quietly acquired StartCom (the company behind the well-regarded StartSSL service) and attempted to conceal the acquisition from the browser vendors — a direct violation of CA/B Forum requirements. When Mozilla investigated, they found worse: WoSign had been backdating certificate issuance timestamps to before the industry’s SHA-1 deprecation deadline, fraudulently extending the life of certificates that should have been rejected. Domain validation had also been bypassed in a number of cases.
The concealed acquisition is the detail worth dwelling on. StartCom had a reasonable reputation — I even used to use them. The acquisition was a mechanism for laundering that trust — inheriting the root store position while operating under different ownership with different standards. Mozilla removed both WoSign and StartCom from their root store in 2016. The lesson: the CA system’s trust is only as durable as its weakest institutional link, and institutional change can happen invisibly.
CAA Records: Telling the Internet Who’s Allowed to Sign for You
Certification Authority Authorisation (CAA) DNS records is an attempt to patch some of this CA-trust issue. It lets you specify which CAs are permitted to issue certificates for your domain. It looks something like this:
example.com. CAA 0 issue "letsencrypt.org" example.com. CAA 0 issuewild ";"
The first record says: only Let’s Encrypt may issue certificates for example.com. The second says: nobody may issue wildcard certificates at all.
Before issuing a certificate, CAs are required by CA/B Forum baseline requirements to check the CAA record and refuse if they’re not listed. If DigiNotar had checked CAA records (and if everyone had deployed them), the fraudulent Google certificate would have been blocked — Google’s CAA record wouldn’t have listed DigiNotar.
The catch, as ever, is the humans. CAA checking is mandatory in principle; enforcement relies on CAs actually implementing it correctly, and adoption of CAA records by domain owners remains lower than it should be. It’s a control that costs almost nothing to deploy and is largely invisible when working correctly, which perhaps explains why it flies under the radar.
If you run a domain and you haven’t got CAA records in your DNS, add them today. It takes ten minutes and meaningfully reduces your attack surface.
Certificate Transparency: An Audit Trail for the CA System
After DigiNotar, the industry needed a way to detect rogue certificate issuance before it caused harm, rather than after. Google’s answer was Certificate Transparency (CT).
The idea is straightforward: every publicly trusted certificate must be logged in one or more append-only, publicly auditable CT logs before browsers will trust it. CAs submit certificates to logs; logs return a signed timestamp as proof of inclusion; that timestamp is embedded in the certificate. Browsers check for it.
This means that if a CA issues a fraudulent certificate for your domain, it appears in the public logs. You can monitor those logs — there are services that will alert you whenever a new certificate is issued for your domain. The attacker can’t hide the certificate’s existence; the best they can do is hope nobody’s watching.
The Symantec CA mess in 2015-2017 was largely surfaced through Certificate Transparency. Google found that Symantec had issued certificates improperly — including test certificates for domains they didn’t own — and the CT logs provided the receipts. Symantec’s CA business was eventually wound down and transferred to DigiCert.
CT doesn’t prevent fraudulent issuance, but it makes it visible. That’s a meaningful improvement. Visibility is however, a double-edged sword; it makes it easier for attackers to see what domains you have generated certificates for, and use this list to inform their reconnaissance of your networks.
The Revocation Problem
Here’s an uncomfortable truth that the PKI ecosystem has been quietly papering over for years: certificate revocation is largely broken in practice.
The problem is conceptually simple. If a private key is compromised, you need a way to tell the world “stop trusting this certificate.” Two mechanisms exist:
Certificate Revocation Lists (CRLs) are lists of revoked certificate serial numbers, published by CAs. Clients download the list and check against it. CRLs can be enormous, they’re updated on a schedule rather than in real time, and requiring clients to download them for every connection is impractical at scale.
Online Certificate Status Protocol (OCSP) improves on this: instead of downloading a whole list, the client queries the CA’s OCSP responder in real time to ask “is this certificate still valid?” Better, but it introduces latency, creates a privacy problem (the CA now knows which sites you’re visiting), and creates a reliability dependency (what if the OCSP responder is down or being attacked?).
The reliability problem is where things fall apart. Browsers decided long ago that if the OCSP check fails — if the responder is unreachable, too slow, or returns an error — they would soft-fail: accept the certificate anyway. The reasoning was that a downed OCSP responder shouldn’t break the internet. This is understandable.
The consequence is that revocation is largely advisory. An attacker with a stolen private key can count on browsers accepting the certificate for the remainder of its validity period regardless of revocation status, because they can block or slow the OCSP check and rely on soft-fail behaviour.
This is widely known. Nobody has a great solution. The industry has responded in two ways.
OCSP Stapling: The Partial Fix
OCSP stapling moves the OCSP query from the client to the server. The server periodically fetches a signed OCSP response from the CA and “staples” it to the TLS handshake. The client gets the freshness information without making its own request — no latency, no privacy leak, no single point of failure.
This is genuinely better. It’s widely supported. And yet the revocation problem isn’t fully solved, because stapling is optional, and clients can’t distinguish “server is stapling correctly” from “server isn’t stapling because it’s misconfigured” from “attacker is stripping the staple.” Unless...
OCSP Must-Staple was the extension that could have completed the picture. A flag in the certificate that says: “this certificate must come with a valid stapled OCSP response; if it doesn’t, reject it.” Clients that respect must-staple would hard-fail if the staple was missing or invalid, closing the soft-fail loophole.
Must-Staple was standardised. It was implemented. It was a reasonable solution. And it went basically nowhere, because the failure mode for a misconfigured server — one with must-staple in its certificate but no working stapling setup — is complete inaccessibility. Web operators burned by it once tended not to go back. Adoption never reached the threshold where browsers could enforce it meaningfully, and it has quietly faded from the conversation.
Short-Lived Certificates: Engineering Around the Problem
If revocation doesn’t reliably work, and the operational cost of fixing it is too high, there’s another approach: make certificates expire so quickly that revocation barely matters.
This has been the direction of travel for years. The CA/B Forum progressively reduced maximum certificate validity: from five years, to three years, to two years, to the current 398-day limit (Apple pushed this in 2020 by simply refusing to trust longer certificates, bypassing the forum process). There’s a clear industry direction towards 90 days and shorter.
Let’s Encrypt issues 90-day certificates by default. Their rationale is explicit: shorter lifetimes force automation, reduce the window for a compromised certificate to do damage, and keep the ecosystem healthier. If your certificate is valid for 90 days and is compromised on day 1, you have a 90-day problem. If it’s valid for two years, you have a two-year problem.
The next frontier is 6-day certificates, which Let’s Encrypt has been working towards. At that validity window, revocation becomes essentially irrelevant — by the time you’ve detected a compromise and gone through a revocation process, the certificate is nearly expired anyway. The security property you want falls out of operational necessity.
The catch is that short-lived certificates require robust automation. You cannot manually renew a 6-day certificate. This is, arguably, by design — it forces organisations to build the automation they should have built anyway.
ACME and Let’s Encrypt: Making Automation the Default
Short-lived certificates only work if the renewal process is reliable and automated. This is what ACME (Automatic Certificate Management Environment) is for, and Let’s Encrypt is why it became ubiquitous.
Before Let’s Encrypt — which launched publicly in 2016 — getting a TLS certificate required: choosing a CA, going through a manual verification process, paying a fee, downloading the certificate, installing it, setting a calendar reminder for renewal, and doing it all again in a year. Many organisations didn’t bother. HTTP was the default, and HTTPS was something you only bothered with if you were handling payments.
Let’s Encrypt provided free certificates and, more importantly, the ACME protocol: a standard way for a client to automatically prove domain ownership and receive a certificate without any human involvement. The Certbot client could be set up in an afternoon. Certificates would renew themselves. The financial and operational cost of HTTPS dropped to near zero.
The proportion of web traffic served over HTTPS went from roughly 40% in 2016 to well over 95% today. Let’s Encrypt is substantially responsible for that. Arguably it’s one of the most impactful infrastructure projects of the last decade.
ACME is now an IETF standard (RFC 8555) and is supported by a large number of CAs beyond Let’s Encrypt. If you’re renewing certificates manually in 2026, you’re doing it wrong — and your CA almost certainly supports ACME.
The Web of Trust: A Different Model
Everything described so far is hierarchical: you trust the root CAs, they delegate to intermediates, intermediates sign your certificate. Trust flows top-down.
PGP (and its open implementations, GPG) took a different approach: the Web of Trust. There are no central authorities. Instead, individuals sign each other’s keys. If Alice has signed Bob’s key, and you trust Alice’s judgement, you can transitively trust Bob’s key. The more people who have signed a key, the more paths of trust exist to it.
In theory this is elegant and decentralised. In practice it ran into several problems. Building a meaningful web of trust requires key-signing parties — actual physical meetings where people check IDs and sign keys. The user experience of PGP has historically been catastrophic, to the point where security researchers write papers about it. Key discovery is a mess, with many approaches being suggested and deployed such as WKD, DANE, with none of them getting much traction — to the point where even github sidestepped the issues and now supports SSH keys as an alternative for signing commits. Revocation is even harder than in PKI. And the social graph embedded in the keyserver network is a privacy problem — it publicly reveals who knows whom, and it also has its own weaknesses.
The Web of Trust works reasonably well in small technical communities where everyone goes to the same conferences. It has never scaled to general use, and modern attempts like Keybase tried to bridge the gap with social proof (link your GitHub, Twitter, and domain to verify your key) with mixed results. For most practical purposes, hierarchical PKI won. The Web of Trust lives on in specific contexts — software package signing, some email security deployments — but it’s not the foundation of anything at scale.
mTLS: When Both Ends Need to Prove Themselves
Standard TLS is one-sided. The server presents a certificate; the client verifies it. The server has no cryptographic guarantee of who the client is. Authentication of the client, if it happens at all, is done at the application layer — a username and password, a session token, an API key.
Mutual TLS (mTLS) extends the handshake so both parties present certificates. The server still proves it’s the server; but now the client must also present a certificate, signed by a CA the server trusts. The server rejects connections from clients that can’t present a valid certificate.
The practical implications are significant. With mTLS, access control happens at the transport layer, before any application code runs, with cryptographic guarantees rather than shared secrets. An API key can be stolen and reused from anywhere; an mTLS client certificate requires the corresponding private key, which should never leave the client.
mTLS is foundational to zero-trust networking and is the mechanism underlying most service mesh implementations (Istio, Linkerd). In a mesh, every service has its own certificate, rotated automatically, and every service-to-service connection is mutually authenticated. No implicit trust based on network location; every call proves its identity. This is particularly powerful in Kubernetes environments where the network boundary between services has always been porous.
The operational cost is non-trivial: you need a PKI infrastructure to issue and rotate client certificates, you need tooling to distribute them, and you need to think carefully about how you handle certificate rotation without dropping connections. Service meshes handle most of this automatically, which is part of why they’ve become popular beyond just traffic management, but the failure modes are real, and can certainly cause downtime.
HSMs: Where the Important Keys Actually Live
Throughout all of this, there’s a question we’ve deferred: where does the private key actually live?
On a well-configured web server, the private key is a file on disk, ideally with tight permissions, ideally encrypted at rest. This is fine for most purposes. But for high-value keys — CA keys, signing keys, keys that protect significant amounts of data — a file on disk is inadequate. The server’s operating system, any privileged process, any sufficiently capable attacker who has compromised the host, can read that file.
A Hardware Security Module (HSM) is a dedicated piece of hardware designed to store cryptographic keys and perform cryptographic operations, with the fundamental property that the key never leaves the device in plaintext. You ask the HSM to sign something; it signs it internally and gives you the result. You cannot extract the raw key material. Physical tamper resistance means that attempting to open the device destroys the keys.
Root CA keys are almost universally stored in HSMs. The ceremony around generating and storing a root CA key — often literally a ceremony, with multiple witnesses, multiple physical keys required to unlock the HSM, recorded on video — is one of the more theatrical aspects of the PKI world.
For operational workloads at smaller scale, cloud KMS services (AWS KMS, Google Cloud KMS, Azure Key Vault) provide HSM-backed key storage as a service. Your application never handles the raw key; it calls the KMS API. The key material lives in hardware you don’t own, can’t extract, and can only use through a logged, audited API.
There are also local options for key storage in hardware — physical HSM USB keys, TPM modules, and even secure parts of microcontrollers called ‘enclaves’ — but none of these are perfect and I will go into these in more depth in a future article.
Secret Management: Beyond Certificates
Certificates and asymmetric keys are one category of secret. But production infrastructure is drowning in other secrets: database credentials, API keys, service account tokens, encryption keys for application data, third-party service credentials. Managing these is its own discipline, and it’s where a huge proportion of real-world breaches originate.
The pathology is well-documented: secrets end up in environment variables, which end up in log files. Secrets end up in configuration files, which end up in version control. Secrets are shared over Slack because the proper channel is too slow. Secrets are never rotated because rotation is manual and painful. Secrets have no audit trail.
HashiCorp Vault and its OpenBao fork became the dominant answer to this for a reason: it provides a unified interface for secret storage and access, with dynamic secrets (credentials generated on-demand and automatically revoked), fine-grained access policies, full audit logging, and automatic rotation. The operational model shifts from “distribute static secrets to services” to “services authenticate to Vault and receive temporary credentials.”
Dynamic secrets are particularly powerful. Instead of giving your application a database password, Vault generates a unique set of credentials for each application instance, valid for a defined period. When the period expires, the credentials are revoked. An attacker who exfiltrates those credentials has a ticking clock rather than indefinite access.
Cloud platforms have their own equivalents — AWS Secrets Manager, Google Secret Manager, Azure Key Vault — with varying feature sets and native integrations. The principle is the same: centralised, audited, access-controlled secret storage with lifecycle management.
The thing that none of these solve automatically is the bootstrap problem: your application needs credentials to authenticate to Vault, so where do those come from? This is turtles all the way down, and the answer is usually platform-native identity (IAM roles for EC2 instances, Kubernetes service accounts with projected tokens, etc.) combined with a careful threat model about what you’re actually protecting against.
Key Rotation: The Operational Side
Cryptographic keys should be rotated. This is widely agreed upon and inconsistently implemented.
The reasons for rotation are: reducing the window of exposure if a key is compromised without your knowledge, limiting the amount of data encrypted under any single key (important for some algorithms that weaken under high volumes), and ensuring that your rotation process actually works when you need it urgently, because the worst time to discover your rotation procedure is broken is when you have an incident.
The challenge is that rotation touches everything that uses the key. For TLS certificates, ACME makes this nearly painless. For database encryption keys, it may mean re-encrypting significant amounts of data. For signing keys, it means ensuring verifiers have the new public key before signers switch to the new private key, with an overlap period.
Key rotation should be automated wherever possible and tested regularly. The test that matters is “can I rotate this key, without downtime, under time pressure?” If you’ve never done a rotation drill, you don’t know the answer.
Certificate Pinning: A Well-Intentioned Own Goal
Worth a mention as a cautionary tale: certificate pinning is the practice of hardcoding an expected certificate (or public key hash) into a client, so it rejects any certificate for that endpoint that doesn’t match, even if it’s signed by a trusted CA.
The threat model is: a rogue CA issues a fraudulent certificate for your domain, and you want to be protected even against that. With pinning, even a rogue-CA-signed cert won’t be accepted by your client.
HTTP Public Key Pinning (HPKP) was the browser-based implementation: a response header that told browsers “only accept certs matching these hashes for this domain for the next N days.” It was standardised, shipped in browsers, and was a spectacular operational disaster.
The problem: if you lose your private key, or simply rotate it in a way that’s incompatible with your pinned hash, your site is inaccessible to every browser that received the pin, for the full duration of the pin — potentially months. No override, no escape hatch. Several high-profile sites accidentally pinned themselves into inaccessibility. HPKP was deprecated and removed from all major browsers by 2019.
Mobile app pinning has similar failure modes — an app pinned to a certificate that’s been rotated stops working until the app is updated, which you cannot guarantee for all users simultaneously. Some vendors have even started to enforce that their app versions are no more than 3 months old, great in theory, but what if the vendor goes out of business? Even with all the other infrastructure available and working, you are suddenly left with a countdown timer and the hope that the company will give you a new version before it explodes.
The lesson is not that pinning is always wrong, but that the operational complexity it introduces is significant and the failure modes are severe. Certificate Transparency largely addresses the rogue CA problem it was designed to solve, without the operational landmines.
The Trust We’ve Actually Built
Stepping back: what the internet has built is a system of institutional trust dressed up as mathematical trust. The cryptography is sound, but the foundation it rests on is a set of agreements between organisations, enforced by the threat of root store removal, audited by third parties whose audits are sometimes insufficient, and dependent on operational hygiene across thousands of organisations.
It works better than it has any right to, given how it’s constructed. It’s failed in specific, documented, serious ways. DigiNotar died. Symantec’s CA business was wound down. Several other CAs have been sanctioned, constrained, or removed. The ecosystem has correction mechanisms.
But the mental model of “the padlock means it’s safe” is a substantial simplification. The padlock means the connection to that server is encrypted and authenticated, within the limits of the CA system. It says nothing about what the server does with your data, whether the certificate was obtained by the legitimate owner of the domain, or whether the institution behind the certificate is trustworthy.
The maths is the easy part. Everything else is trust management, all the way down.
The algorithm is almost never the weakest link. The implementation, the key management, the operational processes, and the humans are where the bodies are buried.