The Future Needs Fewer Agents and More Bureaucracy

Every agentic architecture that survives contact with production grows a filing system.
It never starts that way. It starts with one agent doing an afternoon's work in ninety seconds, and a room full of people deciding that this changes everything. Eighteen months later the same system has an approval queue, scoped credentials with expiry dates, a policy engine that classifies operations before they run, a log nobody wanted to pay for, and a person whose Monday morning is spent reading what the agents did on Friday. The usual reading of that arc is disappointment: the technology promised autonomy and delivered paperwork.
That reading is wrong. The paperwork is not what went wrong with the deployment. The paperwork is the deployment.
I've now watched enough of these proposals to notice they converge on the same furniture, independently, without anyone citing anyone. Authority that flows downward from a coordinator. Credentials scoped per task and expired on completion. Certain operations reserved for a human signature. A record of what was decided and why. A list of which agents exist and who owns them. Nobody sets out to design a government department. They arrive at one, because they are solving the problem government departments were invented to solve: how to let a large number of semi-autonomous actors do work on your behalf without any one of them being able to do something catastrophic, and without losing the ability to find out afterwards what happened.
What bureaucracy is actually for
Max Weber was a German sociologist who died in 1920, and his account of administration in Economy and Society, published after his death, is where most of what management theory believes about the subject started. He called it rational-legal authority, which is the phrase to search for; the surname on its own will mostly get you barbecues. His account is not the one the word carries now. He was describing something that had recently won, and explaining why it won. Written rules that apply regardless of who is asking. Defined jurisdictions, so every office knows the boundary of its own authority. Files, so the institution remembers what individuals forget. The separation of the office from the person holding it, so authority can be handed over without being handed to someone personally. Appointment by demonstrated competence rather than patronage.
Set against what it replaced — decisions made by whoever had the ear of whoever was in charge that week — the case was technical rather than moral. It was predictable, reviewable, and continuous across the people who came and went.
We use the word as an insult now because we only notice bureaucracy when it fails. Nobody writes a sitcom about the year the register was accurate. The queue, the form returned for a missing field, the department that cannot say yes: those are the visible moments, and they are pathologies rather than the design. The design is boring and works, which is why it disappeared into the background and took its reputation with it.
An agent fleet has Weber's problem exactly. Many actors, acting on someone else's behalf, whose individual actions are small and whose aggregate is consequential, none of whom can be handed ambient authority — not because they are malicious, but because ambient authority is indefensible no matter who holds it. We have a century of applied research on that problem. It is filed under a word engineers use to mean "waste".
Authority flows down, or it leaks
The supervision tree argument is the same argument as delegated authority, in different vocabulary. A coordinator decomposes work, scopes each task narrowly, and hands workers credentials that expire when the task does. Workers hold nothing worth stealing between tasks. Authority flows down from a place you can audit, rather than accumulating in N places you cannot.
No permanent secretary hands every clerk the departmental chequebook and trusts the induction pack to cover the rest. The unit is the jurisdiction: a defined area in which an office may act, and outside which it must ask. Least privilege and ministerial authority are the same concept reached from different directions. Both attach authority to a role and a task rather than to an individual, and both exist so that authority can be withdrawn without anybody having to be trusted less.
The part most designs skip is that the coordinator is bound by the same rules. A minister is not above the statute they administer. If the coordinator can bypass policy on the grounds that it is acting for the user, the architecture has moved the problem up a level and called it a solution.
The log is a minute, not an alibi
Civil servants write minutes so that a successor can reconstruct why, not to prove that nobody did anything wrong. That distinction decides what your audit log is for, and therefore what goes in it.
An agent remembers nothing beyond what fits in its context, and context is a budget you are already overspending. The log is where the institution keeps what the actors cannot. It needs the decision and its inputs, not just the effect: the plan the coordinator proposed, the classification the policy engine gave it, the justification offered, what was denied, who signed. A log that records only successful actions tells you what happened and nothing about the system's judgement, which is the thing you actually need when an incident review asks whether this was foreseeable.
It also has to be evidence rather than assertion. An audit table that anyone with database access can quietly edit is a claim about the past, not a record of it, and everyone in the room during an incident knows the difference.
This is the same gap I've written about in bus factor terms: a system can reconstruct what was done far more easily than why. Rationale decays first, and it decays fastest in systems where the actor doing the reasoning is discarded at the end of the task.
Some decisions are reserved
The policy gate argument was that destructive operations belong behind something the agent cannot satisfy on its own — not a confirmation prompt it can answer itself, but a gate whose key it does not hold. Administratively, that's a reserved decision: one that a junior office cannot take, not because the junior is untrusted, but because the consequence exceeds the office.
The test isn't "is this risky?" — almost everything is, and a gate on everything is a gate on nothing. The test is: if this goes wrong, who is answerable? Where the answer is a named person, that person signs. Where the answer is "the system", you have found a hole in your authority model rather than a candidate for automation.
Gates cost latency, and latency is what gets them routed around. A gate that is routinely bypassed is worse than no gate, because it leaves you with documentation of a control you don't have. So: few of them, fast, and approval latency measured as a first-class number. If nobody is watching how long the queue is, the queue is how the control dies.
You cannot govern what you cannot enumerate
Every administration keeps a register of its own offices, because the alternative is discovering an office by its mistakes.
Most organisations already run more service accounts than staff, and cannot say with confidence which of them can reach production. Agents make that worse quickly, because spawning one is cheap and retiring one is nobody's job. The register is the unglamorous prerequisite: which agents exist, what each may do, who owns it, when its authority lapses, and what it did most recently. Every control above depends on it. Least privilege is not a policy you can write about principals you cannot list.
graph TD
O["Owner: a named human"] -->|delegates a bounded remit| C[Coordinator]
C -->|scoped task credentials| W1[Worker]
C -->|scoped task credentials| W2[Worker]
W1 --> P{Policy engine}
W2 --> P
C -->|plans go through the same gate| P
P -->|routine, in jurisdiction| X[Execute]
P -->|reserved| H[Human signature]
P -->|out of policy| D[Deny]
R[(Register of principals)] -.->|who is asking, and may they| P
X -.-> L[(Audit log)]
H -.-> L
D -.-> L
The shape of that is not novel, which is the point. Replace the labels with permanent secretary, officials, standing orders, reserved matters, and the minute book, and you have an organisation chart that would be legible in 1955.
Change management is the layer nobody built
Here is the one that is genuinely missing rather than merely unfashionable.
Your agents' behaviour is a function of at least four things: a prompt, a model version, a set of tool definitions, and whatever the retrieval layer handed over this morning. Most teams have change control for one of them. The prompt lives in a repository if you are lucky, the model version is whatever the vendor is serving behind an alias, the tool schemas move when another team ships, and the retrieval index changes continuously by design.
Then something goes wrong, and the incident review asks the only question that reliably matters: what changed? For a normal service you answer it from the deploy log. For an agent fleet, four things can change underneath you and only one of them leaves a trace.
The fix is not exotic. Treat the prompt, the tool definitions, and the model version as versioned artefacts, and pin them, so an upgrade is a decision you made rather than a Tuesday the vendor had. Stage the rollout: canary a behaviour change across a few agents the way you would canary a service, because a prompt edit shipped to two hundred agents is a fleet-wide behaviour change with no compiler to catch it. Keep a change register, so the review has a column to look in. And have something to test against before the change lands, which for non-deterministic systems means an eval set rather than a unit test.
None of that is novel either. It is a change advisory board with the ceremony stripped out and the record kept.
| Dismissed as bureaucracy | Its agentic equivalent | What it actually buys |
|---|---|---|
| Delegated authority, defined jurisdiction | Coordinator scoping, per-task credentials | A blast radius you can state in advance |
| Reserved matters, ministerial sign-off | Human-in-the-loop gate on destructive operations | A decision with a name attached to it |
| The minute book | Tamper-evident audit log of decisions and denials | The ability to reconstruct why, months later |
| Register of offices | Registry of agent principals and owners | Knowing what exists before it surprises you |
| Standing orders | Policy engine rules that apply regardless of caller | Rules that don't bend for whoever is asking |
| Change advisory board | Versioned prompts, pinned models, staged rollout | A "what changed" column in the incident review |
| Separation of office from officeholder | Authority attached to the role, not the agent | Replacing an agent without re-granting its power |
The pathologies are real
The honest objection is that bureaucracies fail in specific, well-catalogued ways, and nothing about calling the paperwork "governance infrastructure" exempts you from any of them.
Goal displacement, to borrow Robert Merton's term for it, is the big one: the rule stops being a means to an outcome and becomes the outcome, so the form gets filed correctly about a decision nobody examined. Veto points accumulate, each individually reasonable, until the median change takes a fortnight. The process becomes the product, and the department's output is a description of its own diligence.
Agentic systems can acquire all of this faster than human organisations, because the clerical work is nearly free. A fleet that emits four thousand audit entries a day that nobody reads is not accountable; it is compliance theatre running at machine speed, with a storage bill. An approval queue that a tired human clears in batches of fifty is a rubber stamp with extra latency. Cheap paperwork is exactly what lets the pathological version scale, and it is cheaper here than it has ever been.
So the design constraint is narrower than "add governance". Automate the clerical part and keep the accountable part human. Keep the gates few so the ones that exist are taken seriously. Write logs to be read — sampled, reviewed, with someone's name against the review — rather than retained. Generate the register from what is actually running instead of maintaining it by hand, because a hand-maintained register is fiction with a review date. The bureaucracy that works is the one whose records are a by-product of doing the work, not a second job performed afterwards by whoever is least able to refuse.
The shape of it
The demo will go on being one agent doing something impressive in ninety seconds, because that is what demos are for. The deployment will go on being a small civil service, because that is what running things on other people's behalf has always required.
The teams that get this right won't be the ones that eliminate the paperwork. They will be the ones that make it cheap enough to be honest: a register that maintains itself, a log worth reading, a handful of gates that mean something, and a change record that answers the only question an incident review really asks.
Fewer agents. Better paperwork. The route to running a lot of them safely goes through the filing system, not around it.