Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Context Is a Budget: Writing an AGENTS.md That an Agent Will Actually Use

A mid-century librarian hands a mechanical automaton a single index card, chosen from an overflowing armful of scrolls — 1960s gouache.

There is a particular file that almost every team using coding agents now has, and that almost nobody is honest about. It sits at the root of the repository, called AGENTS.md or CLAUDE.md depending on the tooling, and it has grown the way these things grow — one well-intentioned commit at a time. Someone added the build command. Someone else, after a bad afternoon, added four paragraphs explaining why the agent must never touch the migrations directory. A third person pasted in the entire style guide because it seemed tidier than linking to it. By the time anyone looks properly, the file is two thousand lines long, and the agent is quietly ignoring most of it.

That last part is the bit people resist. The file feels authoritative — it is long, it is specific, it is checked into version control — so the assumption is that the agent reads it the way a diligent new hire reads the onboarding wiki: thoroughly, once, and then carries it forward. It does not. The file is loaded into the context window, competes for attention with everything else in there, and past a certain length the marginal instruction stops changing behaviour at all. You did not write a contract. You wrote a very long suggestion, and the model is reading the first third of it.

The window is a budget, and you are overspending

The mental model that fixes this is to stop thinking of the context window as a container you fill and start thinking of it as a budget you spend. It is finite — a fixed number of tokens, shared across the system prompt, your AGENTS.md, the code the agent has pulled in, the conversation so far, and the headroom it needs to actually reason and reply. Every token your instructions file consumes is a token unavailable to the problem at hand. That is the trade, and it is real, even on the large-context models, because the constraint was never only "does it fit". The constraint is "does the relevant material stay salient against everything competing with it".

This is where the research has caught up with what anyone pairing with these tools for a few months already suspected. The "Lost in the Middle" effect — instructions buried in the centre of a long context get attended to less reliably than the same instructions at the edges — is not a quirk you can write your way around with firmer language. It is a property of how attention distributes across a long sequence. A four-thousand-line AGENTS.md is not four thousand lines of instruction. It is forty lines of instruction the model will probably follow, wrapped in three thousand nine hundred and sixty lines of expensive, attention-diluting noise that pushes those forty lines towards the unreliable middle.

So the discipline is not "write a comprehensive file". It is "spend the budget well". And spending it well means most of what is currently in your file should not be there.

What earns its place

The test for any line in an AGENTS.md is simple, and it is not "is this true" or "is this important". Plenty of true, important things belong nowhere near this file. The test is: can the agent infer this from the repository itself, and if it gets it wrong, will the consequences be caught? If the agent can read the code and deduce the convention, you are paying tokens to tell it something it already knows. If it cannot deduce it — and getting it wrong costs you — that is exactly what the file is for.

Three categories survive that test reliably.

The first is the commands the agent cannot guess. It cannot know that your tests run under uv run pytest -q rather than bare pytest, that the linter is ruff rather than flake8, or that there is a make check target that does the lot. These are pure signal — short, deterministic, and the difference between an agent that can close its own feedback loop and one that hands you broken code because it never ran the suite.

The second is the conventions that are decisions rather than deductions. No amount of reading the code reveals that a particular library is banned for political rather than technical reasons, or that "we always return early instead of nesting" was settled in a review eighteen months ago and is not up for relitigation. The code shows what the team did; it does not show what the team decided, and the decisions are the part the agent will otherwise cheerfully overturn one reasonable-looking PR at a time.

The third is the genuine landmines. The migration that must run before the seed script. The service that looks stateless but holds a connection pool you must not exhaust. The directory that is generated and must never be hand-edited. These are the things that are cheap to state and ruinous to discover by experiment.

Everything else is a candidate for the cut.

Keep, cut, relocate

Most of the bloat in a real AGENTS.md is not wrong. It is misfiled — context that belongs somewhere the agent will encounter it in situ, not piled into a single file it reads once at the start.

Content Verdict Why
Exact build/test/lint commands Keep Cannot be inferred; gates the agent's self-correction loop.
Banned libraries, non-obvious conventions Keep Decisions, not deductions — invisible in the code.
Known gotchas and ordering constraints Keep Cheap to state, expensive to hit by accident.
Full coding style guide Cut → relocate Put it in the linter config. A rule the formatter enforces needs no prose.
Architecture deep-dives, history Cut → relocate Belongs in a docs/ file the agent loads only when relevant.
Per-module detail Cut → relocate A short README or docstring next to the code lands in context with the code.
"Please write clean, maintainable code" Cut Pure noise. It changes nothing and costs tokens.

The relocation column is the important one. Cutting something from AGENTS.md is not deleting it — it is moving it to where it earns its tokens on demand. The style guide becomes ruff configuration the formatter enforces without prose. The architecture explanation becomes a doc the agent pulls in only when it is working on architecture. The per-package quirk becomes a docstring that arrives in the window precisely when the relevant file does, and not a moment before.

Write it deterministically

What survives should read like a runbook, not an essay. The agent is not moved by your reasoning, and the paragraph explaining why the migrations directory is sacred spends ten times the tokens of the instruction itself while being easier to lose in the middle. State the rule. Make it imperative, specific, and checkable.

# AGENTS.md

## Commands
- Install: `uv sync`
- Test: `uv run pytest -q`
- Lint + format check: `uv run ruff check . && uv run ruff format --check .`
- Types: `uv run mypy .`
Run all four before proposing a change. All must pass.

## Conventions (decisions, not style — those are in ruff)
- Return early; do not nest beyond two levels.
- No `requests` — use the shared `http` client in `app/clients/`.
- SQLAlchemy 2.x style only. No legacy `Query` API.

## Landmines
- `alembic/` is generated. Never hand-edit; create a revision.
- `app/billing/` touches money. No changes without an explicit ask.

That is perhaps thirty lines, and it will outperform a two-thousand-line predecessor on every axis that matters, because every line of it is load-bearing and none of it is buried. The agent can run the commands, knows the three things it could not have guessed, and is warned off the two places where being wrong is expensive.

The discipline is pruning

The failure mode is not writing a bad AGENTS.md on day one. It is letting a good one rot into a bad one, because the file only ever grows. Every incident adds a defensive paragraph; nothing is ever removed, because removing a line feels like inviting the bug back. So the file accretes, and the signal-to-noise ratio falls, and one morning the agent does the exact thing your file spent four paragraphs forbidding — because those four paragraphs were on line nineteen hundred, and the model never really read them.

Treat the file the way you would treat the hot path it effectively is. Prune it. When a rule moves into the linter, delete the prose. When a gotcha is fixed at the source, delete the warning. Re-read the whole thing occasionally and ask of each line whether it is still earning its budget, because a constraint the agent can no longer find is worse than no constraint at all — it gives you the comfort of having written it down and none of the behaviour.

The goal was never a thorough file. It was more behaviour change from fewer tokens, and those two things are in tension the moment you forget that context is a budget and start treating it like a filing cabinet. Spend it on the handful of things the agent genuinely cannot work out for itself. Everything else, the repository can tell it — if you let it.