Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

The Real Bottleneck in Engineering Isn’t Compute — It’s Attention

A montage representing scarce attention. A man sits at a computer with many thoughts competing for his attention.

There’s a peculiar kind of progress happening in engineering right now. We have more automated code generation, faster CI pipelines, near-instant infrastructure provisioning, and AI pair-programmers that can produce a working authentication service in the time it used to take to find the right Stack Overflow answer. Output capacity has never been higher. And yet, somehow, engineering teams don’t feel faster. They feel more overwhelmed.

The reason is that we’ve spent years optimising the wrong constraint.

In systems terms, you don’t improve throughput by throwing resources at a non-bottleneck. If your database is the limiting factor, adding web servers does nothing except give your load balancer more options for delivering traffic to the same queue. We know this. The Theory of Constraints is on the reading list of every SRE who’s traced a latency spike back to a serial step nobody thought to instrument. And yet, at the team level, we’ve been provisioning more and more compute — AI tools, automation, accelerated pipelines — without addressing the resource that was already saturated before any of this started: the human attention required to steer it.

What Attention Actually Is (and Why It Doesn’t Scale)

Attention isn’t focus. That’s a popular simplification that makes it sound like a setting you can toggle. It’s a neurological process by which the brain selects and prioritises a fraction of the available sensory and cognitive input for deeper processing — the mechanism that lets you function without being paralysed by the sheer volume of information hitting your sensory apparatus at any given moment. As Herbert Simon put it back in 1971, with characteristic precision: a wealth of information creates a poverty of attention. He was describing the early internet era. He had no idea.

The brain runs two broadly distinct attention systems. The bottom-up system is fast, reflexive, and largely involuntary — it responds to novelty, movement, threat, and emotional content. The top-down system is slower, deliberate, and metabolically expensive — it’s what you use when you’re tracing a race condition through a distributed system at 2am, or reviewing a pull request for security implications rather than just style. These two systems are in permanent competition, and modern digital environments have been explicitly engineered to bias the outcome towards the bottom-up: notifications, badge counts, infinite scroll, algorithmic feeds surfacing the most emotionally charged content to maximise dwell time. A US civil court recently found Meta and YouTube liable for deliberately designing their platforms to addict users, particularly the young, by systematically exploiting this asymmetry. That verdict is interesting context, but it’s slightly beside the point here. The same asymmetry that makes social platforms adversarial to sustained thought is operating inside your engineering workflow, with nobody standing trial for it.

The taxonomy that cognitive psychologists use divides attention into four operational modes: sustained (holding focus on a single task over time), selective (attending to one signal while filtering others), alternating (switching cleanly between tasks), and divided (attempting to process multiple streams simultaneously). Every one of these has a capacity limit, a degradation curve, and a recovery requirement. None of them are in the on-call runbook.

The SRE Framing

SRE doctrine distinguishes toil from engineering work not merely by whether the task is manual, but by whether it has enduring value. Toil is repetitive, tactical, and scales linearly with the size of the system — if you’re doing more of it this quarter than last quarter for reasons that aren’t “the system grew”, that’s a signal. The SRE book’s upper bound is 50% toil. The reasoning is that above that threshold, the people who should be improving the system are instead servicing it, and the system falls further and further behind its own growth.

Cognitive load works the same way. There is a finite budget of high-quality attentional capacity per engineer per day — call it four to six hours of deep, focused work for a person in good condition, less under interruption, less again under chronic stress — and that budget is being consumed by an expanding category of attention-taxing toil: reviewing AI-generated code that is syntactically correct but semantically questionable, triaging an alert queue that never quite clears, context-switching between Slack, Jira, the IDE, and three browser tabs that were relevant to something else an hour ago. The tasks themselves may be real and necessary. The problem is that they’re consuming the same attentional resource as the work that actually moves things forward.

A useful way to think about it: engineers have a working memory budget analogous to a system’s CPU time. Deep cognitive work — architectural design, debugging subtle failures, security modelling, understanding someone else’s clever idea in a pull request — is high-CPU. Interruptions don’t just consume CPU during the interrupt; they incur a context-switch penalty. Research on task-switching puts the recovery time at somewhere between fifteen and twenty-five minutes to return to full cognitive depth after an interruption. Most engineering environments generate interruptions at a frequency that makes that recovery time structurally impossible. The net result is a team operating permanently in a degraded mode, with nobody’s attention budget fully replenished, and nobody alerting on it because we don’t instrument human cognitive load the way we instrument CPU and memory.


AI Doesn’t Help With This. At the Moment, It Makes It Worse.

The standard narrative is that AI tools reduce engineering load. This is true in a narrow sense: they reduce the cost of producing code, documentation, tests, and configuration. The production rate of artefacts goes up. What doesn’t go up — what arguably goes down — is the human bandwidth available to validate, integrate, and make decisions about those artefacts.

Code review is the most obvious example. Reviewing AI-generated code at the pace AI generates it is a category error. The review process for a PR written by a human contributor who has been thinking about the problem for two days is qualitatively different from reviewing a 400-line function that an LLM produced in three seconds based on an underspecified prompt. The second is faster to produce and, in my experience, faster to get wrong in ways that are harder to spot — because the surface confidence of the output is high even when the semantic accuracy is not. Rubber-stamping it is dangerous; reviewing it properly costs as much attention as writing it yourself, possibly more. The economics only improve if you restructure the review process around that reality, and most teams haven’t.

Broader than code review: every time AI adds something to the pipeline — a generated ADR, a suggested architecture, a scaffolded test suite — it creates a validation requirement that lands in a human attention queue. If the rate of generation exceeds the rate of thoughtful validation, you don’t have an improved system. You have an unreviewed one that looks reviewed because someone clicked Approve.

Think about what happens when you add web tier capacity to a system whose database is already saturated. You haven’t fixed anything; you’ve just made the queue arrive faster. AI tooling deployed without regard for the human validation stage downstream is exactly this: more throughput at the generation layer, same fixed throughput at the review layer, growing backlog of artefacts that look approved because someone clicked Approve.

That backlog doesn’t stay invisible. It shows up in incident rates, in subtle integration bugs, in the review that waved through the authenticated-as-any-user endpoint because it was the ninth PR of the afternoon and the code looked confident. And incidents, once they happen, consume orders of magnitude more attention than the review would have — which is why good SRE teams treat post-incident work as debt repayment, not penance.

The Hostile Environment

You can’t have a serious conversation about developer attention without acknowledging the environment it’s operating in. The average knowledge worker — and engineers are not exempt — is operating in a communication infrastructure that has been optimised, at considerable engineering expense, to prevent sustained focus. Not as a side effect. Deliberately.

Slack, Teams, email, Jira, GitHub notifications, monitoring dashboards: any of these individually is manageable. Together, and with the social norm that responsiveness is a signal of engagement, they constitute a continuous-interrupt system that keeps the human attentional stack in shallow mode. The bottom-up system is constantly triggered — “oh, someone reacted to that comment I left”, “there’s an amber status on the build I’m not on” — and the top-down work, the stuff that actually moves the engineering forward, never gets the runway it needs.

What’s interesting from a systems perspective is that we’ve implemented this interrupt architecture with all the thoughtfulness of an early-2000s monitoring setup: alert on everything, no severity triage, no escalation policy, assume that the person receiving the page will work it out. We spent years fixing that problem in our monitoring systems. We have not started on the human equivalent.

Social media sits at the edge of this, slightly off-site but still consuming the same resource. The average engineer who picks up their phone during a deep work session doesn’t re-enter deep work when they put it down; they re-enter a lower-quality shallow processing state and gradually rebuild. At which point the phone usually goes again. The asymmetry that Social Europe described — it is far easier to hijack attention than to sustain it — is not a metaphor. It is a neurological fact, and the engineering workflow is full of hijack surfaces.

Neurodivergence Is Not an Edge Case

Any discussion of attention in engineering teams that doesn’t acknowledge neurodivergence is modelling the wrong population. ADHD and ASD are both significantly over-represented in technical roles compared to the general population — the exact numbers are contested, but “higher than average” is not. This matters for two reasons.

First, the standard attention model — steady-state focus available on demand, degraded by interruption, restored by rest — doesn’t describe ADHD accurately. ADHD attention is better characterised as highly state-dependent: either locked in (hyperfocus, which is remarkable productive capacity when pointed at the right problem) or locked out (unable to engage with a task that feels underspecified, under-stimulated, or externally imposed). The interruption economy hits ADHD engineers asymmetrically: the Slack ping that costs a neurotypical engineer fifteen minutes of recovery can cost someone with ADHD their entire afternoon’s deep work capacity.

Second, the environmental interventions that help neurotypical engineers — noise-cancelling headphones, focus blocks, no-meeting periods — are often necessity rather than preference for neurodivergent ones. Designing for the edge case is, in this instance, designing for a substantial fraction of your engineering team, and the people who are pushing back hardest on open-plan offices and always-on communication culture are frequently the ones who are most acutely aware that the environment isn’t working.

Load Management, Not Productivity Hacks

The reframe I want to argue for is this: cognitive load is a systems problem, not a personal discipline problem. “Tips for staying focused” is roughly as useful as “tips for keeping your database fast” — not entirely useless, but missing the architectural level at which the problems are actually caused and fixed.

At the personal level, the standard deep-work toolkit has been well-documented: time-blocking for high-cognitive work, notification lockdown during those blocks, a physical and digital environment structured to minimise bottom-up attentional triggers, capture mechanisms for intrusive thoughts (a notepad, not your phone), and rituals that signal to your brain that the mode has changed. These work. They’re not the whole solution, but they work, and teams that implement them individually report substantial improvements in the quality of their output even when the volume of work hasn’t changed. The philosophy of when to apply them varies — some people do better with rhythmic daily sessions, others with extended focus sprints and then recovery periods — but the underlying mechanism is the same: you’re scheduling access to your top-down attention system and protecting that schedule from preemption.

At the team level, the interventions look more like SRE practice and less like self-help. Interrupt budgets: what fraction of the team’s aggregate attention should be consumed by reactive work, and is that fraction being tracked? On-call rotation for cognitive interrupts, not just production incidents — someone takes point on Slack for the morning while others work in flow, and the role rotates. No-meeting blocks enforced at the calendar system level, not as a gentleman’s agreement. Alert fatigue triage for communication channels, treating persistent high-volume notifications with the same urgency you’d treat a noisy alerting system. If your monitoring system fires fifty times a day and nobody looks at it anymore, that’s a known failure mode. Apply the same analysis to your Slack channels.

The AI tooling question re-enters here. If AI is increasing the rate of artefact production, you need to explicitly budget the attention required to validate those artefacts, or accept that you’re accumulating unreviewed technical debt at an accelerated rate. The right answer is probably structured review workflows that acknowledge the different risk profiles of AI-generated and human-authored code, with attention allocation adjusted accordingly. The wrong answer is assuming that the throughput gain on the generation side translates directly into a throughput gain on the validation side, because it doesn’t, and the lag shows up in incident rates rather than sprint velocity metrics.

Attention as Infrastructure

Here’s the argument I keep coming back to: organisations treat compute as infrastructure. They provision it deliberately, monitor its utilisation, plan for its growth, protect it from abuse (rate limiting, throttling, DDoS mitigation), and invest in its efficiency. They do not do this for cognitive capacity. They hire engineers, assume that full attention is available by default, expose that attention to a continuous-interrupt environment, saturate it with AI-generated work arriving faster than it can be processed, and then express surprise when output quality degrades or engineers burn out.

The engineering profession has spent twenty years building better and better tools for managing system load. We have autoscaling, circuit breakers, backpressure mechanisms, graceful degradation under saturation. We have runbooks for what to do when a service is at capacity. We have post-incident reviews that systematically examine whether load contributed to a failure. We have, in short, a mature discipline for managing constrained resources in complex systems under variable load.

We have not applied any of that discipline to the cognitive capacity of the people operating those systems. The irony is fairly pointed.

This isn’t an argument for treating engineers as machines to be optimised — it’s the opposite. Machines can run at 100% utilisation indefinitely without degrading. Humans cannot. The SRE toil budget exists precisely because Google’s SRE team recognised that humans have a finite capacity for unrewarding work, and that exceeding it has consequences that don’t show up immediately but do show up eventually, in the form of attrition, error rates, and system fragility. The same principle extends to attention more broadly.

What would it mean to treat developer attention as infrastructure? It would mean measuring it — at least approximately, by tracking interrupt rates, meeting loads, and self-reported focus capacity. It would mean provisioning for it deliberately, by structuring the working day around what is known about human cognitive cycles rather than around meeting schedules and communication tool defaults. It would mean protecting it from abuse, which includes the abuse of the AI tooling being deployed in a way that maximises artefact generation without regard for the downstream validation cost. And it would mean having a recovery plan, which includes acknowledging that a team that has been operating at cognitive saturation for three months doesn’t recover in a sprint.

The compute isn’t the bottleneck. It hasn’t been for some time. The question is whether we’ll notice before the queue gets much longer.

Further reading: The SRE Book’s chapter on Eliminating Toil remains the cleanest articulation of the toil-to-engineering-work ratio argument. Cal Newport’s Deep Work is the standard text on structured attention management, though it was written before AI tooling changed the production rate problem. Tim Wu’s The Attention Merchants provides the commercial and historical context for how we ended up with an environment engineered against focus. For the neurological detail, the British Psychological Society’s ongoing work on attention design, and the growing body of research on ADHD in technical roles, are worth the detour.