Supervision Trees

On why authority in multi-agent systems should flow down, not up
The previous three pieces have all been about a single agent: how it gets credentials, how it decides what to do, how the surrounding architecture should constrain it. The harder version of the same question is what happens when the agent isn't alone. Multi-agent systems are not a curiosity any more — they are the most-pitched architectural pattern of the year, and the demo videos are getting better. They are also, mostly, being designed in the wrong direction.
The typical multi-agent demo shows N autonomous agents collaborating: each holds its own credentials, each pulls work from a shared queue, each makes its own decisions about what to do and what it needs. The emergent behaviour is presented as the magic. This is the wrong starting point. It is wrong for the same reasons single-agent autonomy is wrong, only it is wrong in N places at once, and the failure modes compose multiplicatively.
There is a better shape. It comes from operational systems engineering, not from agentic-systems marketing, and it has been working in production for thirty years.
Three patterns, briefly
Federated. Each agent has its own identity, its own credentials, and pulls work from a shared queue. Coordination is emergent. Authority is distributed across the agents. This is the demo shape.
Mesh. Agents negotiate authority among themselves, sometimes via reputation systems, sometimes via cryptographic capability tokens passed between agents. Coordination happens through agent-to-agent communication. This is the research shape.
Hierarchical. A coordinator agent decomposes work, allocates tasks to workers, and grants each worker the narrow scope needed to do its task. Workers hold no long-lived credentials. Authority flows down from the coordinator. This is the operations shape, and it is what Erlang/OTP has been doing with supervision trees since 1996.
The federated pattern is the most popular because it is the most photogenic. The hierarchical pattern is the most operationally tractable because it concentrates authority into auditable surfaces and bounds the blast radius of any individual worker failure.
Why hierarchical, not federated
The case is more or less direct.
Credential surface multiplies in the federated model. Each agent holds its own credentials. The credential foraging problem from Article 1 now has N agents that might forage credentials from N filesystems. The aggregate authority of the system is the union of every agent's reachable tokens, which is, for any non-trivial system, more than any single human had in mind when authorising the agents.
Coordination is emergent and untraceable. Two federated agents that decide independently to delete the same resource produce a race. Two federated agents that decide independently to do conflicting work produce inconsistency. There is no authoritative ordering of decisions, because there is no authority. The system's behaviour is the sum of local decisions, and that sum is harder to predict than any individual agent's behaviour.
Cumulative effects are invisible. Each agent's individual actions might be in policy. The aggregate effect across all of them might not be. A single agent making one schema migration is fine. Five agents each making one schema migration in the same hour is the kind of thing a policy engine should notice — but only if there is a coordinator that sees all five.
Failure isolation is harder. When one agent goes wrong in a federated system, the consequences propagate through whatever shared state the agents touch. There is no clear supervisor to restart it, kill it, or quarantine the work it had already done.
The hierarchical pattern fixes all four. Authority lives at the coordinator, not the worker. Coordination is centralised by definition. Cumulative effects are visible because the coordinator sees them. Failure is isolated by killing the worker and reassigning its task. This is not a new insight; it is the operational instinct that produced supervisor processes in OTP, the kube-apiserver in Kubernetes, and cell-based architecture in every well-designed cloud system. It applies to agents for the same reasons it applied to processes.
The inversion is the point
Most multi-agent demos present autonomous workers as the desirable pattern: agents that decide for themselves what to do, what they need, and how to get it. The hierarchical pattern is the inverse: workers do not decide for themselves. The coordinator decides. The workers execute.
Said differently: the coordinator knows what each worker needs to access, because the coordinator gave the worker the task. The worker doesn't have to decide; the worker doesn't have the authority to decide; the worker doesn't have credentials reaching beyond what the coordinator scoped. If the worker reasons its way into wanting more, it asks. The coordinator either grants narrowly-scoped access or escalates to policy.
This sounds restrictive. It is restrictive. That's the point. The restriction is what makes the system tractable to reason about, audit, and supervise. The cost is that workers cannot improvise outside their scope. The benefit is that workers cannot improvise outside their scope. These are two phrasings of the same property, and the second one is the reason to have it.
What hierarchical needs to do well
For the pattern to be genuinely better and not just an architectural fashion, several specific things have to be true.
Workers hold no long-lived credentials. The coordinator mints scoped credentials per task. Credentials expire when the task completes. There is nothing for the worker to forage, because there is nothing on the worker's filesystem worth foraging. This is the credential foraging defence from Article 1, applied at the worker level.
Each worker has a narrow tool surface. Specialisation is the feature. A code review worker has read access to a repo and nothing else. A test runner worker spins up ephemeral environments and tears them down. A migration worker has very specifically scoped database write access. Narrow tools mean narrow blast radius and policies that can reason about whether a given action is in scope.
The coordinator is itself bounded. The coordinator does not get to bypass policy on the grounds that it is acting on behalf of the user. Coordinator requests go through the same policy gate as anything else. This is the bit that gets skipped in most demos, and it is the bit that turns "supervisor" into "single point of failure with all the credentials." If the coordinator can do whatever it wants, the architecture has only moved the credential foraging problem up one level.
No worker-to-worker delegation. Workers cannot grant other workers access to anything. All scoping routes through the coordinator. Allowing workers to share credentials with each other reintroduces the multiplied authority surface that the hierarchical pattern was supposed to eliminate.
Orchestration plans are auditable. The coordinator's decomposition — "I am going to send these tasks to these workers, and grant these scopes to each" — is itself an artefact that can be reviewed before execution and after. The plan is the request that the coordinator submits to policy; the plan is the audit trail. If the coordinator's planning is opaque, the rest of the architecture is gestural.
Scope is per-task, not per-worker. A code review worker invoked five times in a week is invoked with fresh scope each time. Credentials don't accumulate across invocations. The worker's identity persists; its authority does not.
These are not optional. A hierarchical system that gets any of them wrong is a hierarchical system in name only.
Failure modes
Even done well, the pattern has its own failure modes.
The coordinator becomes a bottleneck. Every decision routes through one process. Throughput is bounded by the coordinator's decision rate. The mitigation is to make the coordinator's policy decisions cheap and parallelisable, and to compose hierarchies — multiple coordinators each handling a sub-domain, with a higher-level coordinator above them. OTP allowed supervisor trees for exactly this reason.
The coordinator becomes a target. Concentrated authority is, by construction, concentrated risk. A compromised coordinator is worse than a compromised worker, because its authority spans every worker it supervises. The coordinator therefore has to be itself supervised — by policy, by audit, by humans on the escalation path. The coordinator's own access to credentials should be the most heavily-gated surface in the system.
The coordinator becomes too smart. The seductive failure: the coordinator starts making decisions on behalf of the user without out-of-band confirmation, because it is convenient and the user has implicitly granted ambient trust. This reintroduces the HITL problem from Article 2, just at a higher level. The fix is structural: the coordinator's decisions about destructive operations go through the same policy gate as a worker's would. There is no privileged class of agent that policy doesn't apply to.
Workers acquire side-channel access. If workers can read each other's filesystems, share state through a database, or pass tokens via shared message queues, the hierarchical model has been silently violated. The supervision tree only works if the boundaries are real. In containerised environments this is mostly automatic; in single-process or shared-filesystem deployments it has to be enforced explicitly.
Plans go unaudited. The coordinator's planning surface is a privileged position. If nobody is reviewing the plans, the coordinator's decisions become the de facto policy, and the de jure policy becomes a rubber stamp. This is the same drift that affects un-audited single-agent policy, only it is harder to spot because the plans are upstream of any individual destructive action.
Where this is hard
Some cases don't fit cleanly.
Long-running agents that span tasks. The orchestrated model assumes work decomposes into bounded tasks with clear scope. A long-running agent that operates continuously — a monitoring agent, a research agent, an always-on assistant — doesn't fit this model directly. The pragmatic answer is to decompose the long-running agent into a coordinator that itself spawns bounded sub-tasks, each with their own scope. The agent's persistent identity is a coordinator; the work the agent does is a sequence of supervised tasks. This is a refactor that most agent designs haven't been through yet.
Discovery and recon. Before the coordinator can decompose work, it has to know what work exists. That requires read access — to filesystems, codebases, ticket queues, recent activity. The discovery surface is itself a credential surface, and a compromised coordinator with broad read access leaks information about the entire system. The mitigation is to give the coordinator narrow, audited read access, and to push the actual recon into bounded sub-tasks where possible.
Negotiation between workers. Some problems are inherently coordinative — two workers that need to converge on a shared answer, or split a piece of work, or hand off state. The hierarchical model handles this by routing the negotiation through the coordinator: workers don't talk to each other directly; they propose alternatives to the coordinator and the coordinator chooses. This adds latency. For some problems it is the wrong tradeoff, and the federated or mesh pattern is genuinely better. The honest answer is that hierarchical is the right default but not the right universal.
Cross-tenant coordination. When the system spans multiple users, organisations, or trust domains, the question of which coordinator owns which work gets complicated. The same task touching two users' data needs the coordinator that has authority for both, or two coordinators that explicitly hand off across the boundary. This is where most current multi-agent demos fall apart, because the demos all run within a single trust domain.
Cold-start coordinators. A coordinator that hasn't yet learned the patterns of work it supervises will allocate poorly, escalate too often or too rarely, and will make plans that the policy engine has to deny frequently. The cold-start problem from Article 3 applies to coordinators with extra force, because their decisions affect every worker beneath them.
The shape that comes next
The honest summary is that the right multi-agent architecture today is a supervision tree where authority flows down, workers are narrow, coordinators are bounded, and the policy engine sits above the coordinator rather than beneath it. This is not the architecture the demos are showing. It is the architecture that, in operational practice, doesn't lose customer data.
The federated pattern will continue to dominate the marketing because it is the more impressive thing to point at. The hierarchical pattern will continue to dominate the production deployments because it is the more impressive thing to actually run. These are the same dynamics that played out in distributed systems generally over the last twenty years: the demonstrably-clever pattern lost to the demonstrably-tractable one in every domain that grew up.
For agents, the bet is that the same dynamic applies. Workers that don't decide for themselves are workers that can be reasoned about. Coordinators that route through policy are coordinators that can be supervised. Authority that flows down is authority that can be revoked.
There's another piece worth writing about how the policy gate itself behaves differently when the requester is a coordinator rather than a worker — what changes about the inputs, the classification, the escalation thresholds. The coordinator's request is a plan, not an action; the policy has to evaluate the plan as an artefact, which is a different shape of evaluation than evaluating a single API call. That is genuinely a different question.
For now, the framing for the multi-agent question is this: do not start by giving N agents N credentials. Start by giving one coordinator the authority to scope tasks, and let everything else inherit.