Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Your Monitoring Stack Knows Too Much

A towering surveillance gallery of small screens and listening apparatus recording every street of a tiny model city, overflowing with intercepted notes, as one uneasy citizen notices — 1960s gouache.

Think about what your observability platform sees. Every HTTP request. Every database query. Every error stack trace with the full request context attached. Every header your application bothered to log because someone, three years ago, was debugging an auth issue and never turned it off.

Now ask yourself: who has access to that?

In most shops I've worked with, the answer is "anyone with a Grafana login." Which is to say, considerably more people than have access to production. The monitoring stack has quietly become the largest, least-governed copy of your production data, and we treat it like it's harmless because it's "just logs."

It isn't.

Authorisation tokens end up in trace attributes when someone instruments a client library a bit too enthusiastically. Email addresses get logged as part of "user not found" errors. Stripe webhook payloads land in request body logs because debug-level logging got promoted to staging and nobody remembered to demote it. Session cookies appear in HTTP header dumps. I've seen full credit card PANs in error logs from a payment integration where the exception handler helpfully serialised the entire request object.

The structural problem is that observability systems are designed for the opposite of confidentiality. They're built to be queryable, low-friction, and broadly accessible — because that's what makes them useful during an incident at 3am. Locking them down defeats the purpose. So we don't, and then we forward all of it to a SaaS vendor who indexes it for thirteen months.

A few things worth doing this quarter:

  • Audit who can read your logs and traces. Compare that list to who can read production. The delta is your problem.
  • Scrub at the collector, not the application. Application-side redaction will be forgotten by the next service. An OpenTelemetry Collector processor with regex-based PII scrubbing catches what developers miss.
  • Set retention deliberately. If you don't need 90 days of full-fidelity traces, don't keep them. Every day of retention is a day of breach exposure.
  • Treat your observability vendor as in-scope for compliance. Because they are, whether your auditor has noticed yet or not.

The dashboard that helps you find a bug in five minutes will help an attacker find your customers' data in three. Same index. Same query language. Different motive.