Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

AI Doesn't Fix the Bus Factor — It Might Make It Worse

An engineer holding aloft a vast structure of interlocking blueprints and glowing reasoning-threads while a boxy mid-century robot photocopies only the flat outer pages, a red bus small on the horizon — 1960s gouache.

The bus factor is the number of people on your team who need to be hit by a bus before the project is in serious trouble. It's a morbid metric, which is why engineers love it. A bus factor of one means you're one unfortunate pedestrian incident away from organisational collapse. Most teams know, in the abstract, that theirs is probably lower than they'd like. Most teams do approximately nothing about it.

Enter AI. The pitch is obvious: AI can explain code it didn't write, reconstruct the reasoning behind architectural decisions, answer questions about the codebase at 3am without waking anyone up, and onboard new engineers faster than any internal wiki ever managed. The optimistic read is that AI flattens the knowledge gradient across a team — the tribal knowledge that used to live exclusively in Dave's head is now, sort of, accessible to everyone. When Dave leaves, you ask the AI.

This is wrong, but it's wrong in an interesting way. AI doesn't reduce the bus factor. In many teams it's actively making it worse, and doing so invisibly, which is the worst possible way for a risk to grow.


Output Is Not Knowledge

The thing AI is actually good at is producing outputs. Code, documentation, explanations, summaries, decision matrices — all of it flows out at a pace that would have seemed implausible five years ago. The quality is often genuinely good. The problem is that outputs and knowledge are not the same thing, and treating them as interchangeable is where the fragility enters.

Knowledge, in the sense that matters for bus factor, is the reasoning behind decisions — the context that isn't visible in the artefact, the path not taken and why, the constraint that shaped the solution that now looks arbitrary. When an experienced engineer makes a decision, they're drawing on a substrate of accumulated context: the failed experiment from eighteen months ago, the production incident that ruled out the elegant solution, the conversation with a customer that reframed the problem entirely. None of that substrate lives in the code. Precious little of it ends up in documentation. Most of it lives in a person.

AI doesn't capture that substrate. It produces outputs that are consistent with the visible artefacts, which is not the same thing. When a junior engineer asks the AI why the state machine is structured the way it is, the AI will generate a plausible and coherent explanation. It may even be correct. But "plausible and coherent" and "accurate to the original reasoning" are different properties, and in a codebase where the original reasoning was never externalised, there's no way to tell them apart.

This is the failure mode: not that the AI gives no answer, but that it gives a confident answer that nobody can verify. The team proceeds on a foundation of reconstructed reasoning, accumulated across months of queries, none of it anchored to what actually happened. The bus hasn't hit anyone yet, but the knowledge has already left the building.


The AI Has Its Own Bus Factor

There's a second problem that's easier to miss because it doesn't involve people at all. AI models are dependencies. They have versions, deprecation schedules, terms of service, vendor roadmaps, and API behaviour that drifts between releases — often subtly, sometimes dramatically.

If your team has built workflows, prompting chains, internal tools, or institutional muscle memory around a specific model's behaviour, you have taken on a dependency with a bus factor of approximately one provider relationship. The model you built around gets deprecated. The API changes. The context window that fit your documentation ingestion pipeline no longer does. The behaviour that made your code review automation work shifts in a way nobody announced explicitly, because LLM behaviour changes aren't always captured in a changelog.

I have watched teams build non-trivial internal tooling on top of a specific model's quirks — its tendency to produce a particular output format, its handling of edge cases in a prompt, its behaviour at a certain temperature — and then spend weeks debugging when a model update changed those quirks. The tooling worked fine. Then it didn't. Nobody hit a bus; a deprecation notice did the same job.

This is not an argument against AI tooling. It's an argument for treating AI models with the same scepticism you'd apply to any other third-party dependency: abstract the interface, test the behaviour, plan for substitution.


The Slow Leak of Collective Competence

Bus factor discussions usually focus on what happens when someone leaves. The subtler problem is what happens while they're still there.

A team that routes a significant portion of its problem-solving through AI is a team where junior engineers are getting correct answers without developing the reasoning skills that produced those answers. This is not unique to AI — it's a version of the same critique levelled at Stack Overflow in 2013 — but the AI case is more acute because the answers are more comprehensive and more immediately satisfying. Stack Overflow told you what to do; modern LLMs tell you what to do, explain why, handle your follow-up questions, and generate the tests. There's very little pressure to struggle through to an independent understanding.

The result, compounded over a year or two, is a team whose collective competence is shallower than it looks. The outputs are fine. The velocity metrics are fine. But the team's ability to reason about novel problems — the problems the AI handles poorly because they're genuinely new — has quietly atrophied. When the senior engineer who was quietly providing the framing for every AI query leaves, what remains is a team that's very good at using AI to solve problems that are similar to problems they've seen before, and significantly less equipped for anything else.

Antifragility, in Nassim Taleb's framing, is the property of systems that don't just survive stressors but get stronger because of them. The way competence is built is through exposure to difficulty — through being wrong and understanding why, through reaching for explanations that don't come easily and building the mental models that make the next problem faster. A team that skips that process isn't fragile in the bus factor sense; it's fragile in a broader sense. It has optimised for output at the cost of the capacity that generates output.


The Worst Case, Concretely

Picture a team of four: one senior engineer and three people in their first or second roles. The senior is good, experienced, and has been using AI heavily for eighteen months. The codebase is clean. The output is solid. The architectural decisions are sensible.

The senior leaves. Not because of a bus — people leave for better offers, for personal reasons, for the usual assortment of human motivations. They give two weeks' notice and do what they can in the handover.

The remaining team now has a codebase full of good decisions with invisible reasoning. They have AI to help them. When they ask the AI why the event sourcing model was chosen over a simpler CRUD approach, the AI explains event sourcing clearly and generates a coherent justification. The justification happens to be wrong — the actual reason was a specific client requirement that arrived mid-build and has since been dropped — but it's coherent enough that nobody questions it. Six months later, a new feature breaks the model's assumptions in a way that makes no sense until someone digs up a contract clause from three years ago.

The team didn't fail because they were bad engineers. They failed because the reasoning that should have been externalised lived in the senior engineer's head, was never captured during the work, and is now genuinely unrecoverable. The AI filled the gap with something plausible. Plausible was worse than nothing, because nothing at least prompts investigation.


Mitigation Is the Wrong Frame

The standard response to bus factor risk is mitigation: documentation, cross-training, pair programming, knowledge-sharing sessions. These are all fine. They're also reactive and expensive and treated as overhead by most teams, which is why most teams' bus factors are lower than they'd like. The mitigation framing asks "how do we recover when knowledge walks out?" It should be asking "how do we design systems where knowledge externalisation is a side-effect of the work itself, not a separate activity?"

The difference matters. A separate documentation activity competes with shipping. It's the first thing dropped when the sprint gets tight. It captures what someone remembers to write down after the fact, which is not the same as what was known during the decision. It reads like archaeology: accurate in outline, thin on the reasoning that drove the details.

What actually works is process harnesses — structures where producing a reasoning artefact is part of completing the task, not a tax on top of it. Architecture Decision Records are the obvious example: the practice of writing a short, structured document — the context, the options considered, the decision, the consequences — at the point when a significant decision is made, as part of the PR or the ticket, before the context fades. ADRs have been around long enough that their effectiveness is fairly well established. Most teams don't write them, because nothing in the standard development workflow requires it.

AI-augmented workflows create a specific opportunity here that's worth naming explicitly. Every significant AI-assisted decision involves a conversation. That conversation contains the reasoning — the options discussed, the constraints surfaced, the path taken and the alternatives rejected. Treating that conversation as ephemeral is a choice. Treating it as a first-class artefact is also a choice, and it's the better one. Export the session. Summarise the reasoning. Link it to the commit. This costs thirty seconds and turns the AI interaction itself into a knowledge capture mechanism rather than a knowledge substitution mechanism.

The distinction between substitution and augmentation is where the bus factor question actually lives. AI as substitution means the AI answers instead of the engineer understanding. AI as augmentation means the AI helps the engineer understand faster, and the engineer's understanding gets externalised as part of the process. Same tool; completely different outcome for the team's fragility.


Designing for Bus-Resistance

The goal isn't mitigation. It's design. A bus-resistant team is one where knowledge externalisation is load-bearing — baked into the workflow, not bolted on afterwards.

In practice this means a few things. Shared context over private sessions: if a significant architectural conversation happened in someone's personal AI session, it didn't happen as far as the team is concerned. Shared or exported sessions, linked to the relevant ticket, change this. Reasoning trails as a PR requirement: not a documentation page that nobody reads, but a short human-readable explanation of why this change exists, what was considered, and what the constraints were. The AI can help write it — the point is that it exists somewhere that isn't a browser tab.

Explicit knowledge audits are useful and underused. Sit down quarterly and map what knowledge is single-threaded — what decisions, systems, or processes only one person fully understands. Then treat that map as a backlog. Not a moral failing to address with urgency, but a risk register to work through methodically. The act of making the bus factor legible tends to produce the social pressure needed to address it.

The more ambitious version of this is designing the team's AI usage to actively build competence rather than substitute for it. This means, occasionally, requiring engineers to understand what the AI produced before shipping it — not as an audit exercise but as a learning mechanism. It means using AI to generate explanations that get discussed in review, not just code that gets merged. It means treating "I used AI to do this and don't fully understand it" as a flag, not a confession.

None of this is incompatible with moving quickly. The overhead is real but small, and it's front-loaded rather than deferred to the point where someone leaves and the knowledge leaves with them. Paying the externalisation cost during the work is almost always cheaper than paying the reconstruction cost afterwards — and unlike the reconstruction cost, it actually works.


The Number to Watch

Bus factor is a proxy for a more important question: how much of what makes your team functional is in people's heads, and what happens to it when those people leave? For most of software history, the answer was "too much" and "it goes with them." AI was supposed to change that. In teams that haven't thought carefully about the distinction between output and knowledge, it hasn't — it's just made the outputs shinier while the knowledge problem quietly got worse.

The fix isn't complicated, even if it requires discipline to execute. Treat reasoning as an artefact. Treat AI conversations about significant decisions as documentation, not ephemera. Build process harnesses that make externalisation automatic rather than optional. And audit your bus factor explicitly, because the number that makes everyone uncomfortable is the number worth knowing.

If your bus factor is one, and that one person happens to be the one who knew how to prompt your internal tooling correctly, you have a more interesting problem than you probably realised.