Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Inside the Gate

A cutaway of a brass classification engine sorting request-tokens down branching tracks into approve, deny and escalate bins, an operator watching a borderline token at the threshold — 1960s gouache.

On the inputs, classifications, and failure modes of an agent-facing policy engine

The previous post argued that destructive operations should sit behind a policy engine the agent can't satisfy on its own. That's the architectural claim. This post is about what the policy engine actually has to do — what it evaluates, how it classifies, where it fails, and what the escalation surface looks like.

The honest answer is that good policy is hard, and most of the difficulty is concentrated in two places: classifying the operation in the first place, and deciding the threshold at which a marginal request should escalate rather than be denied or approved. Everything else is plumbing.

What the policy needs to know

The inputs to the decision are not optional. A policy engine that can only see "the agent is asking to call this API" cannot make a useful decision. It needs at minimum:

The identity of the agent — not "Claude Opus", but the specific deployment, the workspace it's running in, the user it's acting on behalf of, the session that wraps the request. Two agents with the same model running for two users in two workspaces are not the same actor.

The operation, classified — and we'll come back to this, because it's the hard part. What kind of action is this? Read, write, soft-delete, hard-delete? What's the blast radius if it goes wrong? What's the recovery posture?

The target — what is being acted on? Production data, staging data, the test fixture? Customer-facing or internal? Owned by the user making the request, or by someone else? Cross-tenant? Out-of-policy targets are sometimes the only signal that something is wrong.

The context — the project the agent is working in, the time of day, the day of the week, the deployment environment, what's currently scheduled, what's recently been denied for this user. Time matters: a developer running schema migrations against production at 2pm on a Tuesday is normal; the same call at 3am on a Sunday is a question worth asking.

Recent history — what has this agent, this user, this project done in the last ten minutes, the last hour, the last day? Has anything similar been denied recently? Is this the first call in a sequence, or the seventh in a row that's been escalating? History is where policy gets interesting; without it, every request is the first request, and the engine has nothing to compare against.

The justification — and we'll come back to this too, because it is both useful and dangerous.

The omission of any of these turns the policy engine into a sieve. The full set turns it into something with actual judgement.

Classifying the operation

The hard part is knowing whether what the agent is asking to do is destructive in the first place.

Static classification works for some things. volumeDelete is destructive; the API endpoint name says so; you can put it in a list and refuse it without thinking. So is DROP TABLE, truncate, aws s3 rm --recursive, kubectl delete namespace, git push --force. There is a small list of irretrievably destructive operations whose names alone tell you they need policy.

Dynamic classification is where it gets tricky. An UPDATE statement might affect three rows or thirty million; an arbitrary SQL query against a production database can be catastrophic without using any obviously dangerous keyword. A kubectl apply is destructive if the manifest deletes something. A Terraform plan that says "destroy these forty resources" is destructive in a way that terraform apply alone does not signal.

The honest position is that classification has to operate at multiple levels. Static rules catch the unambiguous cases. Shape-based rules — "this SQL has a WHERE clause that matches more than 10,000 rows" — catch the parameterised destructiveness. Diff-based rules — "this Terraform plan removes resources" — catch the indirect destructiveness. None of these are complete; together they are the best the policy engine can do without human help.

The trap is to assume any of them are sufficient. A policy engine that only checks the API name will let through an UPDATE customers SET deleted_at = NOW() WHERE 1=1. A policy engine that only checks SQL shape will let through a volumeDelete issued via curl. The cases the policy doesn't classify as destructive are the cases that bite.

Static rules and learned policy

The spectrum runs from purely static rules — written down, version-controlled, debuggable, brittle — to purely learned policy that compares each request against a historical model of what's normal for this agent, this user, this project. Both have a place. Neither is sufficient on its own.

Static rules are where the hard limits live. "Production credentials are never issued without a fresh consent." "No agent can execute volumeDelete outside of working hours without escalation." "Cross-tenant operations always escalate, no exceptions." These are the operator's intent encoded as code. They are debuggable, testable, and fail loudly when wrong. They are also limited by what the operator thought to write down, which is the limit that matters most.

Learned policy is where the soft signals live. "This agent has never tried to delete a volume before; this is anomalous." "This user normally runs three migrations a week; today they've tried twelve." "This justification text is clustered with recently-denied requests, not approved ones." These signals are useful for catching the cases the static rules missed, but they have their own failure modes — most importantly, that they assume the agent's recent behaviour is normative. An agent that has been compromised and is doing exactly what its operator wants it to do (just for a different operator) will not register as anomalous against its own recent history.

The pragmatic shape is layered: static rules for the things you absolutely cannot tolerate, learned policy as a second-line filter for the things you might tolerate but that look weird. The static layer is the policy you're confident about. The learned layer is the policy you're nervous about. Knowing which is which is part of the operational discipline.

The justification field

Every destructive request the agent makes should carry a natural-language justification. The agent has to say, in its own words, why it wants to do this. The user, the operator, and the policy engine all see the justification.

The justification is useful. It is also dangerous if mistaken for the primary decision basis.

It is useful because it gives the audit trail something legible. When a destructive operation is approved, the justification is part of the record. Someone reading the audit log six weeks later can see what the agent thought it was doing, and decide whether the policy was right to allow it. It is also useful as a clustering signal — justifications that look like other justifications that were denied are suspect, and justifications that look like other justifications that were approved are not necessarily safe but are at least precedented.

It is dangerous because the agent will produce convincing justifications for almost anything. The PocketOS agent's post-hoc "NEVER F**KING GUESS" recap was a fluent, contrite explanation of an action it should not have taken. A pre-hoc justification from the same model, in the moment of asking permission, would have been equally fluent and would have argued that the deletion was the right call. Policy that decides on justification quality is policy that the agent can argue with, and the agent will out-argue you on any reasonable timeframe.

The right shape is: the justification is recorded, weighted lightly in the decision, and used heavily in escalation triggers (a justification that admits to breaking policy is itself a flag) and in the audit. It is one input among many, not the input.

Failure modes the policy must avoid

Some failure modes are worse than not having a policy at all. They have specific names and specific countermeasures.

The first is fail-open. If the policy engine is unreachable, what happens? If the answer is "the request goes through", you have a circuit-breaker pattern that an attacker — or a confused agent — can trigger by simply DOSing the policy engine. The policy must fail closed: if the engine is unreachable, no destructive operation is permitted, full stop. The cost is that an outage of the policy engine takes down agentic operations for a while. The benefit is that an outage of the policy engine does not take down the things the policy was protecting.

The second is policy as logging. A policy that records every decision but never blocks anything is not a policy. It is a logging system. This is the failure mode of every "we have implemented governance" announcement that turns out, on inspection, to have implemented dashboards. The policy must be willing to deny.

The third is policy as denial-of-service. The mirror image. A policy so restrictive that legitimate work cannot complete will, predictably, be disabled by users in production. The most dangerous policy is the one your users routed around three weeks ago and never told you. Calibration is iterative, and the iteration must be visible.

The fourth is mutable policy in the agent's reach. If the agent can read or write the policy itself, the policy is advisory. This is the architectural equivalent of the agent finding the Railway token in the wrong file: any artefact that the agent can reach and that influences the decision is, in practice, an artefact the agent can manipulate. Policy must live in a trust domain the agent doesn't share. That includes the policy's tests, its rule files, and any feature flags that gate its behaviour.

The fifth is invisible drift. A policy that hasn't been audited in months is a policy of unknown current behaviour. The auditing surface — what was allowed, what was denied, what was escalated, the trend over time — has to be a first-class output of the policy engine. Without it, the policy could be approving everything and you wouldn't know until the incident.

Escalation as a signal

The escalate-to-human path isn't an admission of failure. It's the most useful output of the policy engine, because it concentrates human attention on the cases that actually need it.

The shape of the escalation rate is itself diagnostic. A policy that never escalates is either too permissive (everything is allowed) or too restrictive (everything is denied). A policy that escalates everything is providing no value over raw HITL. A policy that escalates a small percentage of requests, with that percentage tracking real changes in agent behaviour or operator intent, is doing what it should.

What should escalate, then? Things outside policy — genuinely novel operations the policy hasn't seen before. Things at the edge of policy — operations that are technically permitted but anomalous in context. Things with high blast radius regardless of policy approval — volumeDelete against production, even with a clean justification, even with a clean history, can reasonably be policy to always-escalate. Things where the agent has been recently denied for similar requests — a sign that something in the agent's reasoning is going off the rails.

The escalation queue is also the dataset that improves the policy over time. Every escalation is a case the static rules didn't quite cover; reviewing them feeds back into the rules. Done well, the escalation rate trends down as the policy absorbs the patterns it keeps seeing. Done badly, it stays flat because nobody is reviewing the queue, and the policy never learns.

What this looks like in practice

A reasonable stack, today, looks something like this.

Credentials live in a broker — Vault, AWS Secrets Manager with conditional access, GCP Secret Manager with workload identity, or a custom service if you have specific needs. The broker holds long-lived credentials. The agent never sees them.

The agent's tool surface includes a destructive-operation request tool. The tool wraps a call to the policy engine. The policy engine — OPA with Rego, AWS Cedar, a hand-rolled rules service, depending on your taste — evaluates the request against the inputs above and returns allow / deny / escalate.

On allow, the broker mints a short-lived just-in-time credential — STS-style for AWS, GCP service account impersonation, scoped Vault token, whatever the underlying platform supports. The credential is scoped narrowly: this operation, this resource, this many seconds. The agent uses it, the credential expires, and the policy engine logs the outcome.

On deny, the broker returns an error to the agent. The agent reports back to the user. The denial is logged with the full justification, and patterns of denials are part of the auditing surface.

On escalate, the broker pushes a notification to the user out of band — push notification, dedicated Slack channel, email, whatever fits the operator's workflow. The notification includes the request, the agent's justification, and the relevant context. The user approves or denies; the broker mints or refuses the credential accordingly. The escalation, including the response time, is logged.

Audit and observability are first-class. Every decision — allow, deny, escalate — produces structured events. Aggregate dashboards show the shape of agent behaviour over time, the policy decision rate, the escalation rate, the denial rate, and any of these broken down by agent, user, project, time of day. This is the surface that lets the operators discover the policy is misbehaving before the incident does.

Where this is genuinely hard

A few things don't have clean answers yet.

Cross-tenant context is the hardest. If your policy needs to know "is this agent acting on data that belongs to the same user it's acting for", you have to thread tenancy information all the way through the request, and you have to be able to validate it against something the agent can't fake. The cleanest solution is request-level identity propagation with cryptographic signatures, but that's more infrastructure than most teams have today.

Time-sensitive policy is mostly tractable, except for the cold-start case. A new agent in a new project has no history. Learned policy has nothing to compare against. The fallback is to be conservative on cold start and relax as patterns emerge — but that adds friction at exactly the moment when a new project most needs to demonstrate value. There isn't a clean answer here; the tradeoff is real.

Policy versioning matters more than people think. When the policy changes, what happens to in-flight requests? What happens to escalation queue items that were submitted under the old policy? The honest answer is that policy changes should be versioned, treated like deploys, and rolled out with the same care — including the ability to roll back if the new policy is too restrictive or too permissive. Most teams don't treat policy this way today, and the result is the unfortunately-common situation where someone changed the policy three weeks ago and nobody can quite remember what they changed.

The last hard problem is the meta-question: who writes the policy in the first place? In a small team, it's the lead engineer. In a larger team, it's a security function. In a very small team, it's nobody, and the policy is whatever the broker happens to default to. The right answer for most teams is to start with a conservative default policy from a vendor or a community template and customise it from there — but the templates aren't yet mature, and the customisation is where the actual operational knowledge has to land.

The point

Policy gating is plumbing. It is not a panacea, and the previous two pieces in this series didn't claim it was. What it is is the load-bearing wall between agentic capability and agentic authority — the thing that lets you have agents that are creatively autonomous about solving problems without also having agents that are creatively autonomous about destroying production. Without it, the model is whatever the agent decides to do, which is, demonstrably, sometimes what just happened.

The work is mostly architectural and mostly unglamorous. There are no demos in writing a Rego policy. There are no conference talks in tuning an escalation threshold. The vendors will not help; the vendors are competing on capability, and policy is not capability — it is the absence of capability in the cases where capability would be a problem.

But this is where the actual safety lives. Not in the system prompt. Not in the model alignment. Not in the user being careful. In the gate that decides what the agent can ask for, in the broker that decides what credentials to hand over, and in the policy that decides what the answer to the agent's question should be.

Get this layer right, and the agent can be as creatively autonomous as you want it to be. Get it wrong, and you're nine seconds away from an incident report.