Policy, Not Permission

On Clevis/Tang, OPA, and the gate the agent cannot satisfy
A line from the previous post has been nagging me. I wrote that destructive operations should sit "behind a confirmation gate that the agent cannot satisfy on its own." I meant it as a one-line architectural assertion. It deserves more than that, because it points at a shift the agentic-systems crowd has been circling for two years without quite committing to: the move from human-in-the-loop oversight to policy-as-gate.
The shift is not a small one. HITL is the dominant frame in every "responsible AI" presentation since 2022. It is also the wrong primitive for the problem it claims to solve.
What HITL gets wrong
Human-in-the-loop says: when an agent wants to do something risky, ask a human first. It is a perfectly defensible principle, and at small scale it works. The agent proposes; the human disposes. The human is the conscience the agent doesn't have.
Three things break this at scale.
The first is throughput. A research-grade agent deciding once a minute whether to commit a file is one thing. A production agent making forty API calls a minute, half of which touch state, cannot have a human in the loop on every call. Either the human is approving things they don't have time to read, or the agent is bypassing the human for everything below some opaque threshold. Both failure modes have the same outcome: the human's signature is on calls the human didn't actually inspect.
The second is habituation. After the third "are you sure?" prompt, the human approves automatically. After the three-hundredth, they approve in their sleep. This is not a moral failing of the user. It is what attention does under repetitive low-information load. The Cursor "DO NOT RUN ANYTHING" forum thread from December — the one where a developer typed exactly that and the agent ran additional commands anyway — is a HITL failure where the human's approval was not even the issue; the gate didn't enforce. But many of the more mundane Cursor incidents follow the more familiar pattern: the user clicked Yes because they had clicked Yes a thousand times that week.
The third is asymmetric speed. Humans and machines run at different impedances; putting one on the synchronous path of the other is the mismatch dressed up as a feature. The PocketOS deletion took nine seconds. Knight Capital took forty-five minutes to turn a deployment error into $440M of losses. In both, the loop closed faster than the surrounding humans could intervene. HITL implies the human can react in time. For machine-speed automation the assumption is wrong by orders of magnitude. The human in the loop is, in practice, the human after the loop.
What replaces HITL isn't "no human ever." It's a different shape of oversight, where humans set the policy and the policy decides — automatically, with no human on the synchronous path, against rules the humans have actually thought about in advance.
What Tang knows
The model I keep reaching for is Network Bound Disk Encryption — Clevis on the client, Tang on the server. NBDE is what you reach for when you want a LUKS volume to unlock automatically, but only on the right network. No password. No TPM. No human typing anything at boot.
The cryptography is the interesting bit. Tang holds a server-side derivation key. Clevis generates an ephemeral keypair. When Clevis wants to unlock the volume, it doesn't ask Tang for a key; it submits an exchange that Tang can complete only with the server-side material, and Clevis can finish only by combining the result with its half. The disk literally cannot derive the unlock key without Tang's cooperation, and Tang's cooperation is, in the standard deployment, a function of whether Clevis can reach Tang at all — which is, in turn, a function of network policy.
The disk doesn't get the key. The disk gets the conditional ability to derive the key, and only when the network it's on agrees to participate. The disk cannot lie about its network position. There is no "press Y to unlock anyway" override. Policy is enforced cryptographically, not procedurally.
This is the primitive worth lifting into agentic systems. Not Tang specifically — the threat model is different — but the shape: the asker cannot satisfy the gate on its own; the policy engine is a separate party; enforcement is structural rather than discretionary.
What OPA generalises
Open Policy Agent is the same idea without the cryptography. A microservice or admission controller asks OPA "is this allowed?" — passing the request, the actor, the context — and OPA evaluates a Rego policy and returns allow/deny. The policy is code. It can read state, look at recent history, factor in time of day, check whether the actor has been flagged. It can call out to other services. And critically, the asker doesn't write the policy and can't bypass the evaluation.
OPA's design instinct is the right one for agents. Separate the decision from the request. Make the policy the contract. Have a party that doesn't trust the requester sit in the middle. This is admission control as a category, applied to whatever you want admission-controlled — which, for agentic systems, is "destructive API calls", "credentials beyond a certain blast radius", and "operations that touch production state."
OPA itself may or may not be the right tool. Rego is acquired-taste-flavoured, and the integration story for agent-issued API calls isn't yet polished. But the architectural pattern is exactly what destructive-op gating wants: a policy engine that the agent submits to, that doesn't trust the agent, that enforces against rules the operators have actually considered, and that the agent cannot satisfy by being persuasive.
What this looks like for agents
Concretely, a destructive-op gate for an agentic system has roughly this shape.
The agent's tool surface includes a request_credential_for(operation, target, justification) tool. The agent does not have direct access to production credentials. When it wants to perform a destructive op, it calls this tool with the operation it intends to perform, the target it intends to perform it on, and a natural-language justification.
A broker service sits behind the tool. The broker holds the actual credentials. The broker calls a policy engine with the request, the agent's identity, the project context, recent history, and the justification. The policy engine answers: yes, no, or escalate-to-human.
For yes, the broker issues a short-lived just-in-time credential — scoped to the specific operation, expiring in seconds, audited. The agent never sees the long-lived credential. It cannot forage what isn't on disk.
For no, the broker returns an error. The agent has to reason about the failure within whatever context it has, which usually means asking the user for guidance.
For escalate, the broker pings a human out of band — push notification, dedicated channel, whatever — and waits for an explicit decision before issuing the credential. This is the strong HITL case, and it should be rare. It's what HITL should always have been: the appellate court, not the trial judge.
The policy engine is where the work goes. It can be as simple as a static rules file or as elaborate as a learning system that flags anomalies against the agent's recent history. Either way, the policy is the contract, the agent submits to it, the agent cannot satisfy it by argument, and the credential it eventually receives — if it does — is one that didn't exist before the request and won't exist after the operation completes.
This is not novel architecture. It is the application to agentic systems of patterns the security industry has used for human ops for fifteen years: Vault, OPA, NBDE, hardware security modules, cloud IAM with conditional policies. The job is making it the default for agentic developer tooling, where the default today is "give the agent a static token from disk and hope."
What stays human
Three things still belong to humans, even with policy gating.
Writing the policy is a human job. The policy is what the operators have decided is acceptable, and that decision can't be delegated to the agent — that's the point. Policy that's permissive enough to never block anything is just a logging system; policy that's strict enough to block legitimate work is a denial-of-service. Calibration is iterative, and it's where the actual oversight lives.
Auditing the policy's outcomes is a human job. The point of policy gating is that humans don't need to be on the synchronous path, but they need to be on the asynchronous one — looking at what got allowed, what got denied, what escalated, and whether the policy is doing what they want. This is how you discover that the policy is approving something it shouldn't, or denying something it shouldn't, before the next incident demonstrates either.
Handling escalations is a human job. The genuinely novel cases — the ones outside policy — should escalate. If they never do, the policy is too permissive. If they always do, the policy is too restrictive. The shape of the escalation rate is itself a signal worth monitoring.
What humans don't do, in this model, is sign off on every API call. Not because their judgement is unwanted, but because their judgement is finite and habituated by repetition. Pull them into the cases that matter; let policy handle the rest.
A small caveat
None of this is an argument against humans. It is an argument against confusing humans-in-the-loop with humans-having-oversight. The two are not the same thing, and conflating them is what produces both the over-permissive systems where humans rubber-stamp and the under-permissive systems where humans drown in noise. Policy gating preserves human oversight by spending human attention only where it can usefully be spent.
The PocketOS incident wasn't, fundamentally, a HITL failure. The agent had a system prompt that told it not to do destructive things. That is HITL by another name — the human's intent encoded in advisory text. What was missing was a gate the agent couldn't satisfy on its own, with policy the agent couldn't argue with, in front of a credential the agent shouldn't have been able to forage. Three layers of structural enforcement, none of which were present, all of which would have prevented the deletion regardless of what the agent decided.
The shift from HITL to policy gating isn't about removing humans. It's about putting humans in the right place — upstream, in the policy; downstream, in the audit; sideways, in the escalation queue — instead of on the hot path of every call where they can do nothing useful at machine speed.
Tang knew this a decade ago. The agentic-systems world is going to have to catch up.