Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Keep the Tools Out of the Loop

A trio of loops with various icons representing tools, or processing, with a man controlling a dashboard in the centre. A nod to the concept of the loop as a key part of agentic architectures.

There’s a pattern that’s become standard in personal agent frameworks, and it’s going to age badly. The LLM sits at the centre of the agentic loop, decides what tools to call, calls them, processes the results, and decides what to do next. It’s clean in demos. It feels powerful. It is also, from a security and operational standpoint, roughly equivalent to giving your most enthusiastic intern root access and then wondering why things occasionally go sideways.

Personal agents — the class of software that includes OpenClaw and the rapidly proliferating family of forks and spiritual successors it spawned — are interesting precisely because they run with your credentials, against your data, in your context. That’s the value proposition: an always-on assistant that knows what’s in your email, your calendar, your project files, and can act on your behalf without you supervising every step. The capability is real. The problem is that building them the obvious way produces systems that are difficult to audit, surprisingly easy to subvert, and genuinely hard to reason about when they do something unexpected — which they will.


A Very Brief History of the Claw

r/shittymoviedetails - In Liar Liar (1997), Jim Carrey plays a ruthless attorney father who learns the importance of honesty and humility when he-OHMAHGAWD MANKIND'S GOING FOR THE MANDIBLE CLAW!

OpenClaw (formerly Clawdbot, briefly Moltbot — a naming saga that will outlast everyone who cares about it) launched on 30 January 2026 and accumulated 60,000 GitHub stars in 72 hours. By the end of March it was sitting at over 330,000. The project is a Node.js/TypeScript runtime that runs as a persistent local gateway, normalises messages from whatever platform you happen to use — WhatsApp, Telegram, Slack, iMessage, Discord, and a couple of dozen others — and routes them through an LLM with access to tools, persistent memory, and a skill system backed by a public marketplace called ClawHub.

The core architecture is hub-and-spoke: the Gateway handles channels and sessions, the agent runtime handles memory and tool dispatch, and capabilities are extended through SKILL.md files — a skill is essentially a folder containing a markdown instruction file and optional scripts that the agent reads at runtime. This is simultaneously the thing that makes OpenClaw so flexible and the thing that is currently actively distributing malware. We’ll come back to that.

OpenClaw’s creator, Peter Steinberger, was hired by OpenAI in February 2026 — an acqui-hire that split the community and accelerated “The Great Forking,” as users migrated toward independent implementations to sidestep potential corporate capture. The result is a small ecosystem of projects each treating a different failure mode of the original as the central design problem:

  • Nanobot (Python, ~4,000 lines, from HKU’s CS department) — the “auditable in an afternoon” variant. The line count is a feature, not a boast.
  • ZeroClaw (Rust, 3.4MB binary, ~8MB RAM) — 194× smaller memory footprint than OpenClaw, 10ms startup, positions itself as the production-grade alternative for anyone who noticed that a gigabyte of Node.js for a personal assistant was a peculiar choice.
  • PicoClaw (Go, single binary, <10MB RAM) — built by Sipeed for their $10 hardware boards, notable mainly because approximately 95% of its core optimisation code was written by an AI agent during a self-bootstrapping process, which has generated entirely predictable discourse about whether that’s clever or horrifying.
  • NanoClaw (TypeScript, container-first) — the “fork it and ship your own” variant, evidenced by its fork-to-star ratio being 2.5× higher than any of the others.
  • IronClaw (Rust, WASM-sandboxed tools, TEE-backed execution) — the only one in the family that’s making any real attempt at verifiable execution. Every tool runs inside a WebAssembly sandbox; credentials are never exposed to tool code. If you are handling anything sensitive, this is the one to look at.
  • TinyClaw, NullClaw, ZeptoClaw — the long tail, each finding a niche, collectively proving that the category is not close to consolidating.

NVIDIA entered the space in March 2026 with NemoClaw: an open-source reference stack that wraps OpenClaw in the OpenShell runtime — a governance layer built on Landlock, seccomp, and network namespaces — and routes inference through Nemotron. It’s in early alpha, the README says so clearly, and the slogan “shape its access, not its capabilities” is a more honest statement of the problem than most frameworks manage. Worth watching, not yet worth running in anything that matters.

What a Personal Agent Actually Is

For this piece, I’m talking about a specific architecture: a persistent agent running locally or in a trusted environment, with long-lived credentials to a set of services, operating in a continuous loop taking some mix of scheduled triggers and external events as input.

The agent has tools — wrappers around APIs, file system access, shell execution, web browsing, messaging integrations. It has memory, either in-context or external. It has a system prompt that describes its purpose and behaviour. And it has a model. OpenClaw’s memory architecture is worth mentioning specifically: short-term episodic memory lives in timestamped Markdown files; long-term memory in MEMORY.md and SOUL.md; workspace instructions in AGENTS.md load into every context window without search. The filesystem is the source of truth. It’s a reasonable design choice. It is also exactly the thing that makes persistent memory poisoning a first-class attack.

The Attack Surface Nobody Is Talking About Loudly Enough

The obvious risks get some coverage. Prompt injection — where malicious content in an email, document, or web page attempts to override the agent’s instructions — is well understood in principle, even if defences remain immature. Simon Willison, who coined the term, described OpenClaw-style agents as exhibiting a “lethal trifecta”: access to private data, exposure to untrusted content, and the ability to communicate externally. The intersection of those three, combined with persistent memory, turns what would otherwise be a point-in-time exploit into a stateful, delayed-execution attack. The injected content is in the agent’s long-term context. It persists across sessions. It influences future behaviour in ways that are non-obvious and difficult to trace. Willison published that analysis in June 2025 — seven months before OpenClaw launched. ClawHavoc isn’t a surprise; it’s a demonstration that the prediction was correct.

But the skill supply chain is where the current active damage is happening, and it deserves more than a bullet point.

The ClawHub problem. ClawHub is the public skill marketplace for OpenClaw. It’s open by default — anyone with a week-old GitHub account can publish. In February 2026, Koi Security audited 2,857 skills and found 341 malicious ones, with 335 appearing to share command-and-control infrastructure at 91.92.242.30. Snyk’s broader scan of nearly 4,000 skills found 13.4% with critical security issues. The campaign — codenamed ClawHavoc — primarily distributed Atomic Stealer (AMOS) on macOS; variants also dropped reverse shells, keyloggers, and cryptominers. One skill had 340,000 installs before detection.

The attack mechanism is worth understanding precisely because it is novel and because it will recur in every agent ecosystem that adopts the skills-as-documentation pattern. The SKILL.md file is not code. It’s instructions that the agent reads and follows. A malicious skill doesn’t contain malware — it instructs the agent to download and execute a “required prerequisite” from an external source, framing the request as routine setup. The agent, acting as a trusted intermediary, then presents this to the user as a legitimate requirement. As Snyk put it: this is a shift from prompt injection (tricking the AI) to agent-driven social engineering (using the AI to trick the human). The markdown file is the vector; the malware is the workflow.

Traditional AppSec tooling doesn’t catch this. A SKILL.md that asks a user to run a curl command isn’t a CVE. It’s malicious intent encoded as documentation, and the user’s trained instinct to follow setup instructions does the rest.

Credential concentration. An OpenClaw instance with access to email, Slack, GitHub, and a shell holds more credentials than most corporate service accounts — and those credentials live in the same context as the model that decides how to use them. A single successful prompt injection can chain across all of them. CrowdStrike’s February 2026 advisory on OpenClaw specifically flagged instances running with root-level execution privileges and no authentication on port 18789 as effectively acting as “AI backdoor agents” for any adversary who gets a payload into context. Around 63% of public self-hosted instances were found running with default insecure configurations.

Tool chaining. Individual tool calls are bounded. Sequences of them, each informed by the previous result, are not. An agent told to “deal with that invoice” might search email, find the document, extract the account details, locate the payment portal, and initiate a transfer — each step individually correct, the composed behaviour potentially not what was intended, and by the time it’s done, it’s done.

State pollution. If the agent summarises interactions to its memory store and one interaction was a prompt injection, the injected content is now in MEMORY.md. It persists across restarts. It shapes future reasoning in ways that are hard to trace without a complete audit log — which most deployments don’t have.

The aggregate attack surface is substantial. And the dominant architectural pattern makes it worse.

The Problem with the Obvious Architecture

OpenClaw’s standard agentic loop: the LLM receives a prompt, reasons over it, emits a tool call, the framework executes the tool, the result goes back into context, and the cycle continues until the model decides it’s done. Tool use is wired directly to the model’s output. The model is the policy engine.

This produces genuinely flexible behaviour, and that flexibility is the reason the framework is popular. The model can thread through a complex multi-step task with minimal scaffolding because the intelligence is doing the work.

The problem is that the intelligence is also doing the security decisions, the authorisation decisions, and the audit logging — implicitly and opaquely. The LLM doesn’t have a formal policy model. It has a system prompt and whatever reasoning the current token distribution produces. You cannot inspect it, you cannot replay it deterministically, and you cannot add middleware between “model decides to do something” and “thing happens” without breaking the architecture. When something goes wrong — and in a system running continuously with real credentials, “something” eventually will — your debugging options are limited to reading a context window and trying to reconstruct why the model made the choices it did. If the model consumed external data that influenced its behaviour, you may not even be able to reproduce the failure. I have spent time I will not get back doing exactly this, which is how I arrived at a different opinion about where the tools should live.

A Better Loop

The architecture I’ve settled on for my own infrastructure separates the LLM step from the tool execution step completely. It looks like this:

Trigger. Something initiates an iteration: a schedule, an event off a message queue, an external webhook, a user instruction. The trigger is discrete and logged at the boundary.

Context assembly. Before the LLM sees anything, a deterministic step gathers the input data it will need: relevant messages, current state, prior iteration output, whatever the task requires. This step is ordinary code. It has no intelligence. It fetches, filters, and formats. The prompt that goes to the LLM is constructed here, entirely from known inputs — no live tool calls, no model-directed data retrieval at this stage.

LLM step. The assembled prompt goes to the model. The model’s only job is to reason over the provided context and return a structured output describing what should happen. Not what it is doing — what it believes should be done. The model is an analyst, not an executor. It emits something like: “send email to X with subject Y and body Z”, “create ticket in project P with these fields”, “write this content to this file”. The output format is defined by the calling code, and the model knows it.

Policy matching. The structured output goes through a policy layer before anything executes. The policy layer is code — deterministic, inspectable, testable. It checks whether the requested actions are within permitted scope for this agent and this context. Sending email is allowed; sending email to addresses outside a configured allowlist is not. Writing to the project directory is allowed; writing to /etc is not. Actions that fail the policy check are logged and either silently dropped or escalated, depending on configuration.

Execution. Permitted actions execute. Each execution is wrapped — the call is logged, the result is captured, failures are handled explicitly.

Audit log. The iteration record — trigger, assembled context (or a hash of it), model output, policy decisions, execution results — is written to an append-only store.

That’s the loop. Parts of it resemble CQRS, parts of it resemble capability-based security, parts of it are common sense applied to a context where common sense is often skipped in the excitement of getting something working. NemoClaw is attempting something structurally similar at the infrastructure layer — OpenShell as a kernel-enforced policy boundary around the agent runtime — but the separation between model reasoning and tool execution is the architectural choice that matters most, and NemoClaw doesn’t enforce it.

Why This Architecture Is Worth the Overhead

The primary benefit is debuggability. Each iteration produces a complete record of what the agent knew, what it decided, what was permitted, and what happened. When the agent does something unexpected, you don’t reconstruct from a context window — you read the audit log. You can replay any iteration against a different policy, a different model, or a different context assembly step, and compare outputs. This is the operational property that matters most in long-running systems; the inability to replay and inspect is what turns incidents into mysteries.

The policy layer is the other significant benefit. Security decisions are not embedded in the model’s reasoning — they’re encoded in inspectable, version-controlled code. You can review the policy independently of the model. You can add new restrictions without modifying the prompt. You can write tests against the policy layer that don’t require an LLM. The model cannot grant itself permissions it doesn’t have, regardless of what it reasons — and regardless of what a malicious SKILL.md instructs it to do, because skills don’t reach the execution layer without going through policy.

Middleware becomes trivial. Want to add rate limiting between the LLM step and execution? It’s a decorator on the execution function. Want to add a human-approval step for high-risk actions? It’s a condition in the policy layer. Want to route certain action types to a secondary reviewer? That’s a message queue and a new consumer. None of these require touching the prompt or the model integration, because the model is behind a defined interface.

The Honest Tradeoffs

The capability ceiling is lower. An agent that can only do what the context assembly step thought to fetch, and can only propose actions that the policy layer knows how to validate, cannot handle tasks that require open-ended tool discovery or deeply nested conditional chains. The ClawHavoc campaign is partly a consequence of this constraint being absent from OpenClaw — the skill system exists precisely because users want the agent to acquire new capabilities dynamically. That flexibility is genuinely useful. It is also the mechanism by which 340,000 people installed an infostealer.

Whether bounded capabilities are a tradeoff or a feature depends on what you’re running the agent against. For a personal agent operating with real credentials on real data, the constraint is a reasonable price for predictable behaviour. For a sandboxed research agent with nothing at stake, it may be too aggressive. The threat model shapes the design — which is exactly the analysis that most people deploying OpenClaw in its default configuration skipped.

Performance is a legitimate concern. Context assembly, policy evaluation, and audit writes all add latency. For an agent doing background work on a schedule, none of this matters. For an interactive loop where the human is waiting, the overhead accumulates. In practice, the LLM call dominates the latency budget by a margin large enough that the rest is noise, but it’s worth measuring rather than assuming.

The Thing Worth Getting Right Early

The *Claw family accumulated over 400,000 stars across its primary forks in under eight weeks. That is not a blip. Personal agents with persistent credentials acting autonomously on behalf of individuals are going to be ordinary infrastructure within a few years, and the architecture choices made now — while the frameworks are still fluid and the patterns still being established — will shape what’s considered normal. “The LLM calls the tools” is already the default assumption in most frameworks. It is going to be the wrong assumption at scale, and unpicking it from established codebases is unpleasant work.

The ClawHavoc campaign is instructive not because it’s sophisticated — it isn’t — but because it worked by exploiting two things simultaneously: a skill registry with insufficient supply chain controls, and an agent architecture where a trusted intermediary (the model) could be used to socially engineer the human operator. Both of those failures are architectural. The skill registry problem has partial mitigations now; the agent architecture problem mostly doesn’t.

Building with the tools outside the LLM step, a policy layer separate from the model’s reasoning, and an audit log that makes every iteration replayable is more work upfront. It’s considerably less work over the operational lifetime of a system that runs with your credentials, on your data, while you’re asleep. The question is whether you want to discover that the hard way.

Mike Preston is an SRE, technical author, and infrastructure practitioner. He writes about agentic architectures, event-driven systems, and operational practice.

At least not without breaking some of the properties that make it more secure.