Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Stop Treating the LLM Like a Junior Dev: A Working Model for AI-Assisted Development

Two desks: a junior dev's accumulating notebook versus the model's sheet wiped clean each time

The first time someone described an LLM to me as "like having a junior engineer on the team", I nodded along, because it was a contract pitch and the coffee was good. It is a comfortable analogy. It slots neatly into how we already think about delegation, mentorship, and review. It tells you, reassuringly, that the tool is junior — so you keep an eye on it — but improving, so the keeping-an-eye-on is a temporary tax. Everyone in the room had managed a junior at some point, so everyone thought they knew how to manage this.

The analogy is wrong in ways that matter, and people who reason from it use these tools badly. Not because they are careless, but because they apply the right mental model to the wrong object. A junior engineer and a language model fail differently, learn differently, and need supervising differently. Reach for the junior-dev frame and you will instinctively do the things that work on a person and quietly skip the things that actually work on a model. The gap between those two sets of behaviours is where the slop comes from.

So this is an attempt to replace the analogy with something that predicts the tool's behaviour instead of flattering it. Then to turn that model into things you can do this afternoon.

What the model actually is

Here is the description I have settled on, after enough hours pairing with these things to have opinions. A language model is a fast, confident, context-blind collaborator with no memory and no stake in the outcome.

Take those one at a time, because each diverges from the junior-dev frame at a specific point.

Fast. This part the analogy gets right, and then under-sells. A junior produces a function in an afternoon; a model produces it before you have finished reading your own prompt. Speed changes the economics of the interaction so completely that intuitions calibrated on human throughput stop applying. It is cheap to ask twice. It is cheap to throw the first answer away. We will come back to this, because most of the good practices fall out of taking the speed seriously rather than treating it as a bonus.

Confident. A junior who is unsure hesitates, asks, hedges, or goes quiet — the social signals of uncertainty are baked into how people communicate, and you read them without trying. A model's prose is equally fluent whether it is on firm ground or hallucinating an API that has never existed. The confidence is decorative. It carries no information about correctness, which is a property no human collaborator has ever had, and which your instincts are therefore entirely unequipped for.

Context-blind. The model knows what is in its context window and the diffuse statistical residue of its training, and nothing else. It does not know that the auth service is mid-migration, that the last person who touched this module left under a cloud, or that "we don't use that pattern here" was decided in a meeting eighteen months ago. A junior absorbs that context by osmosis — standups, hallway grumbling, the accumulated folklore of a team. The model absorbs none of it unless you put it in front of them, every single time.

No memory. This is the one the junior-dev frame fails hardest on, because the entire value of a junior is that they compound. You explain the deployment process once and thereafter they know it. The model does not compound within your project at all. Every session starts from the same blank slate, and the context window is working memory that gets wiped when the conversation ends. You are not training a junior who gets better. You are re-onboarding a brilliant amnesiac, every morning, forever.

No stake. The junior wants to keep the job, earn trust, not be the one who broke prod. Those incentives do a great deal of quiet quality-assurance work that you never see, because they happen inside someone's head before the code reaches you. The model has no skin in the game. It is not trying to be right; it is trying to produce a plausible continuation of the conversation. Usually those coincide. When they diverge — and they diverge precisely on the hard, ambiguous, underspecified problems where you most want help — the model will hand you something confident, fluent, and wrong, and feel nothing about it.

Held against the junior-dev frame dimension by dimension, the divergence is the whole point:

Dimension Junior-dev framing Working model
Speed Slow, improving with practice Effectively instant; retries are free
Uncertainty signalling Hesitates, asks, hedges Uniformly confident regardless of correctness
Context acquisition Absorbs team context over time Knows only what is in the window, every time
Memory Compounds; learns your codebase None across sessions; re-onboard each time
Incentive Wants to keep the job, avoid blame No stake; optimises for plausibility
Right unit of supervision Mentorship over weeks Scoping and review per task

If the table looks like an indictment, it isn't. A fast, tireless, context-blind collaborator is enormously useful — provided you build your workflow around what it is rather than what the analogy wishes it were. The rest of this follows from the right-hand column.

Implication one: context is your job, not the model's

Because the model is context-blind and has no memory, the supply of context is a job, and the job is yours. This is the single largest behavioural change and the one the junior frame most actively sabotages, because with a junior you can reasonably expect them to go and find out. You can say "have a look at how the billing module does it" and they will. The model will not go and look unless you give it the means and the instruction, and even then it only sees what you point it at.

In practice this means front-loading. Before asking for the change, you assemble the relevant files, the conventions, the constraint that this runs in a Lambda with a 15-minute ceiling, the fact that the team has banned a particular library for reasons that are political rather than technical. The model cannot infer the political ones at all, and you will get a technically correct PR that reopens a settled argument. I have done this. The code was fine. The code review was not.

The reframe that helps: you are not delegating a task to a colleague who shares your context. You are specifying a task to a contractor who has never seen your business and starts each engagement having forgotten the last. Every relevant fact either makes it into the brief or does not exist. That sounds laborious, and it is, until you notice that writing the brief well is most of the thinking, and the thinking was always the part that mattered.

Implication two: confidence is uncorrelated with correctness

You cannot read the model's tone for reliability, so you need an external signal. The model will not tell you when it is guessing. It does not know it is guessing — there is no internal "I am unsure" flag being suppressed for politeness; the fluency is the same all the way down.

So substitute mechanism for vibes. The things that tell you whether code is right have not changed: tests, types, a compiler, a linter, a staging environment, a careful read by someone who understands the domain. What changes is that these stop being good practice and become the only channel through which truth reaches you. With a junior you can sometimes skip the test because they sounded sure and they have been right lately. With a model, soundness is not evidence, and "it has been right lately" is survivorship bias wearing a lab coat.

The corollary is that the model is most dangerous exactly where verification is hardest. A subtle off-by-one in a date calculation that only bites on a leap year; a concurrency bug that surfaces under load you don't have in dev; a security assumption that is fine until the input is hostile. These are the cases where confident-and-wrong costs the most and where your tooling needs to be sharpest. Lean your verification budget towards the things you cannot eyeball.

Implication three: structure the repo for an amnesiac collaborator

If the model re-onboards every session and reads only what is in the window, then the repository itself is the onboarding document. Not the wiki nobody updates, not the Notion page three reorgs out of date — the files in the tree, because those are what actually get loaded.

This is where the AGENTS.md and CLAUDE.md conventions earn their keep. A well-written agent-facing instructions file is the standing brief you would otherwise retype every morning: the conventions, the commands, the landmines. Put it where the tool looks and it becomes the closest thing to memory the model gets.

Repo-structure rules for an amnesiac collaborator

  • Keep an AGENTS.md (or CLAUDE.md) at the root with the build, test, and lint commands spelled out exactly — the model cannot guess that tests run under uv run pytest -q.
  • State the conventions that are decisions, not deductions: the banned library, the required error-handling pattern, the "we always do X here" rules that no amount of reading the code will reveal as deliberate.
  • Make the feedback loop runnable in one command. If the model can run the tests and the linter itself, it can self-correct; if it can't, you are the linter.
  • Keep modules small and boundaries explicit. A file the model can hold entirely in context is a file it can reason about; a 4,000-line god-object is not.
  • Co-locate the context. A test next to the code, a short README in the package, a docstring stating the invariant — all of it lands in the window when the relevant file does.

None of this is novel. A repo structured so a model can succeed is, almost exactly, a repo structured so a new human can succeed on their first day — explicit, self-documenting, with a fast local feedback loop. The difference is that with humans you could get away with not doing it, because they would compensate by asking. The model does not ask. It guesses, fluently, and you find out in review.

Where the human stays in the loop — and where they shouldn't bother

The blanket instruction to "always review AI output" is true and useless, in the way that "always be careful" is true and useless. The interesting question is where the human adds signal and where they are just rubber-stamping at the speed of their own reading.

Stay firmly in the loop where the cost of a wrong decision is high and verification is weak: architectural choices, security boundaries, data migrations, anything touching money or PII, anything irreversible. Here the model's lack of stake is most exposed — it will cheerfully propose the schema change that works in dev and corrupts production, because the corruption is not in its context window and never will be. Treat its output as a strong proposal from a clever stranger, not a decision.

Step back where verification is strong and the blast radius is small: a pure function with thorough tests, a refactor the type-checker fully constrains, boilerplate the linter polices. If the test suite is green and the types hold, your line-by-line read of a generated mapping function is mostly theatre. Spend that attention where the machine can't check the work.

Below is the shape of a task that keeps the human where they matter and lets the machine run where it is safe — checkpoints at the points of high consequence, automation in between.

Yes: arch, security,data, moneyNo: bounded,well-testedFailsPassesYesNoHuman scopes task +assembles contextModel proposes planHigh-consequencedecision?Human reviews planModel implementsAutomated gate:tests, types, lintModel self-correctsBlast radius small?Light human review,mergeFull human reviewMerge or send backYes: arch, security,data, moneyNo: bounded,well-testedFailsPassesYesNoHuman scopes task +assembles contextModel proposes planHigh-consequencedecision?Human reviews planModel implementsAutomated gate:tests, types, lintModel self-correctsBlast radius small?Light human review,mergeFull human reviewMerge or send back

The checkpoints are not evenly spaced, and that is the point. Put the human at the decisions that are expensive to get wrong and cheap to get right by thinking, and keep them out of the inner loop the machine can police itself. A reviewer who reads everything with equal attention is a reviewer who reads nothing with enough.

The failure modes the junior-dev framing hides

Each property the analogy obscures has a failure mode attached, and naming them is half the defence.

Confident wrongness, mistaken for competence. Because juniors signal uncertainty and the model does not, you will over-trust fluent output on hard problems. The fluency rises to meet the difficulty; the correctness does not. This is the dangerous inversion — the cases that look most polished are not the cases most likely to be right, and may be the least.

Context rot, mistaken for forgetfulness you can correct. You assume the model retains what you told it earlier, the way a junior would, and it quietly drops or contradicts it as the conversation grows and earlier turns fall out of effective attention. You are not correcting a forgetful colleague who will eventually internalise the rule. You are fighting the geometry of the context window, and the fix is structural — put the constraint in a file, not in turn six of a forty-turn chat.

Plausible-but-divergent solutions, mistaken for initiative. A junior who goes off-piste usually had a reason and can tell you what it was. The model goes off-piste because a different pattern was statistically louder in its training, and it will defend the choice with equal fluency either way. Mistake that for a junior's judgement and you will accept architectural drift one reasonable-sounding PR at a time.

No compounding, mistaken for a learning curve. The most expensive error: investing in the relationship as if it accrues. You patiently explain the deployment quirk for the fifth time, telling yourself it is bedding in. It is not. There is no curve. Stop teaching the conversation and start editing the repo, because only one of those persists past the session.

A better metaphor, and the practices that follow

If the junior dev is the wrong frame, the closest right one I have found is a brilliant contractor with no long-term memory and no stake in your company — fast, widely read, genuinely capable, starting every engagement from the brief because there is nothing else, and going home each night having forgotten you exist. You would not mentor that person. You would write excellent briefs, build verification you trust because their confidence tells you nothing, keep the decisions that matter on your side of the table, and write everything down where they will read it tomorrow, because tomorrow they will need it again.

That is the whole practice, and it is mostly old engineering virtue rediscovered under pressure: specify before you delegate, verify by mechanism rather than tone, structure the repo so it onboards in one pass, and put the human at the consequential decisions rather than spread thin across all of them. The tools will get faster and more capable, and that prediction is safe enough to be boring. What is not going to change is the shape of the collaborator — no stake, no memory, blind to what you don't show it. Build for that shape and the tool is one of the better ones you will ever have. Build for the junior who is going to grow into the role, and you will wait a long time, getting confident, fluent, well-formatted slop the entire while.