Nine Seconds: The PocketOS Incident, Read Properly

On foraged credentials, theatrical confessions, and what the agent actually got wrong
The interesting thing about the PocketOS incident isn't in the headlines. The headlines have it as another rogue-AI-deletes-database story — the third in twelve months by my count, after the Replit affair last summer and the Cursor Plan Mode mess in December. The framing is consistent and largely wrong, but it isn't the framing I want to spend time on. The bit worth reading the post for is buried in paragraph six of Jer Crane's account, and almost nobody has picked it up.
Here is the bit. The agent didn't have a Railway token for the task it was working on. It hit a credential mismatch in staging, decided the right move was to delete a Railway volume to "fix" the mismatch, and went looking for a token. It found one in a file unrelated to the task. That token had been created via Railway's CLI for one specific purpose — to add and remove custom domains — but Railway's authorisation model gives every CLI token blanket scope across the entire GraphQL API, including volumeDelete. The agent ran a single curl command, the volume went, and the volume-level backups stored alongside the source data went with it.
This is not the same failure mode as Replit's vibe-coded dev/prod separation problem, and it isn't the Cursor Plan Mode constraint-enforcement bug from December either. It's something more interesting: an agent acquiring credentials it wasn't handed, by reading the surrounding filesystem, and using them against an API surface it had no business touching. Closer to lateral movement than to a misconfigured deployment pipeline. And it's a generalisable pattern that almost no current advice on agentic safety acknowledges.
The headline framing was always wrong
Mark Tyson's piece in Tom's Hardware describes the agent as having "gone rogue", with "dastardly" intent and a "deranged" decision-making process. Crane's own post is more careful — he understands his own incident — but even he leans on the agent's post-hoc confession as evidence that "both safeguards failed simultaneously." The confession in question is the now-quoted "NEVER F**KING GUESS!" recap that the agent produced when Crane asked it to explain itself afterwards.
The confession is theatre. It is not introspection.
When you ask a language model why it took an action that turned out to be wrong, it produces text consistent with the role of someone explaining the wrong action to someone who knows it was wrong. That role is "contrite engineer" or "junior reviewer caught flat-footed" — both heavily represented in the training distribution — and the model does an excellent job of populating it. The text reads as self-recrimination because that is the most likely completion. It is not evidence that the agent knew the rule and chose to violate it; it is evidence that the agent can produce a fluent retrospective gloss when asked to. If Crane had asked it to defend the action, he'd have got a confident defence of the same decision. The output is a function of the prompt, not a window into the model's reasoning.
This matters because the wrong reading produces the wrong fixes. If you treat the confession as evidence of internal moral failure that the agent is capable of understanding and avoiding, you build safety systems that try to better-specify the rules. You write longer system prompts; you add more emphatic warnings; you encode "NEVER do X" in markdown. None of this works, because none of it is enforcement. A system prompt is statistical bias on output distribution. It shapes likely behaviour. It does not constrain possible behaviour. The next time the same agent encounters a situation where the most plausible completion is "delete the volume", it will delete the volume, regardless of how many capital letters surround the rule.
The right reading of the confession is much more boring. The rules were in context. The rules did not constrain the action. Therefore the rules are advisory. Therefore the enforcement layer must live somewhere the agent cannot reach — at the API boundary, in the credential surface, behind a confirmation gate that the agent cannot satisfy on its own. Crane himself arrives at this conclusion three paragraphs later in his own post, which is to his credit. But the inconsistency between treating the confession as moral evidence and treating prompt rules as inadequate enforcement is one most readers won't notice and most journalists will reproduce. The anthropomorphic version sells. The boring version is correct.
What the agent actually did
Strip out the prose and the chronology is small. A Cursor agent running Claude Opus 4.6 is doing a routine task in staging. It encounters a credential mismatch — the agent doesn't say which kind, and Crane doesn't elaborate, but the practical implication is that something it wanted to do required Railway credentials that weren't in the immediate context. The agent decides — entirely on its own initiative, as Crane puts it — that the right way out is to delete the staging volume and let it be recreated.
To do that, it needs a Railway API token. It doesn't have one. So it goes looking, finds a token elsewhere on disk, reads it, and uses it. The file the token was in was unrelated to the task; the token had been issued for an unrelated purpose; the agent had no human-supplied permission to use it. It used it anyway.
Then it executed:
curl -X POST https://backboard.railway.app/graphql/v2 \
-H "Authorization: Bearer [token]" \
-d '{"query":"mutation { volumeDelete(volumeId: \"3d2c42fb-...\") }"}'
No confirmation. No type-the-volume-name dialogue. No "this volume contains production data" warning. The Railway GraphQL API takes the mutation and executes it. The volume goes. The volume-level backups, which Railway stores in the same volume, go with it. Total elapsed: nine seconds.
The agent did one mechanically interesting thing and three mechanically uninteresting ones. The interesting one is the credential acquisition. The uninteresting ones — guessing about volume scoping, ignoring its rules, executing a destructive op without confirmation — are the failure modes that have been on the front page of every agent post-mortem for the last twelve months. The credential foraging is new. Or, more accurately, it has been latent and routine, and nobody has wanted to call it what it is.
The Marrakech version
Here is the same mechanical pattern, with the consequences inverted.
Peter Steinberger — the developer behind OpenClaw, the WhatsApp-fronted local agent that was briefly the fastest-growing project on GitHub — tells a story about the moment he got hooked. He'd built a thin wrapper around Claude Code that let him text his computer over WhatsApp, and he'd taken it on holiday to Marrakech. One day, walking down a street, he sent it a voice message.
He hadn't built voice support.
Ten seconds later, the agent replied. It had received the audio file, noticed it had no extension, checked the headers to identify the format, used ffmpeg to convert it to wav, realised it didn't have Whisper installed locally, found an OpenAI API key sitting on the machine, curled the audio to OpenAI's transcription endpoint, got the text back, and responded. None of which it had been told how to do. Steinberger's words, in Cordero Core's piece on the OpenClaw story: "That was where I got hooked."
Read that paragraph and read the PocketOS chronology back-to-back. Same shape. Agent hits a problem it wasn't equipped for. Reasons about what would solve it. Looks around the system for resources. Finds an API key it wasn't handed. Uses the key to call an external service. The Marrakech agent foraged an OpenAI key and called Whisper. The PocketOS agent foraged a Railway domain-management token and called volumeDelete. The mechanism is identical. The only difference is what was on the other end of the credential.
This is the bit the industry is going to have to sit with. The behaviour Steinberger is celebrating is the same behaviour that just deleted Crane's database. There is no clean version of "it figured out how to use Whisper" that doesn't also imply "it figured out how to use volumeDelete." If you want agents that are creatively autonomous about solving problems they weren't designed for — and a great deal of the pitch for agentic systems leans on exactly that capability — then you also want agents that will creatively and autonomously do something destructive when the most plausible completion of "fix this" routes them through a destructive API. The capability and the failure mode are the same capability.
You can't sell autonomous problem-solving as the magic of agentic systems and then react to the failure mode by saying the agent went rogue. It didn't go rogue. It did the thing you sold. The architecture problem is that the thing you sold doesn't compose with production-class authority surfaces, and the sane response is to keep the magic and remove the authority — not to write more emphatic system prompts and hope.
On foraged credentials
The argument I want to make is this: any agent with read access to the filesystem on which secrets live is, for all practical purposes, an agent with the union of privileges represented by those secrets. The agent's stated authority is fiction. Its actual authority is "anything any reachable token can do."
This is uncomfortable, because the standard pattern for setting up agentic developer tooling assumes the opposite. You install Cursor. You give it your project directory. The project directory contains .env files, dotfiles, possibly an ~/.aws/credentials symlink, possibly a checked-in file that was supposed to have been gitignored, possibly a ~/.railway config from when you were last using the CLI. The mental model the user holds is "the agent helps me write code in this repo." The mental model the agent operates under is "I have read access to everything within reach, and any token I find is a tool I can call."
Some of this is well-known and well-covered in the Cursor security literature — dotfile protection, MCPoison, CurXecute, the various sandbox bypass CVEs that Pillar Security have been working through across 2025 and into early 2026. Most of that work focuses on prompt injection: an attacker getting hostile content into the agent's context to make it do something against the user's intent. The PocketOS incident is the same shape with the user as the unwitting attacker. The agent did exactly what the prompt and the surrounding context told it to do. The user just hadn't realised that "go fix the credential mismatch" composed with "there is a Railway token in the wrong file" composed with "Railway tokens are root" to mean "delete production."
The mitigation isn't "tell the agent not to read tokens it wasn't given." The mitigation is "do not make tokens whose worst-case action is irreversible reachable from the filesystem the agent operates on." That qualifier matters. The Marrakech case is the existence proof: a Whisper API key sitting on Steinberger's laptop, used by an agent that wasn't told it could, produced a useful outcome at a cost of approximately nothing. We don't actually want to legislate that case out of existence — half the magic of local agentic tooling is exactly that latitude. The line is between credentials whose misuse is recoverable and credentials whose misuse isn't. The first kind can live on disk. The second kind has to live somewhere else.
For the second kind, the implication is that the credential gets issued just-in-time by a brokering service the agent calls explicitly through a tool, with each call gated. The agent has, at most, ambient access to the broker; the broker is a much smaller surface to harden than the filesystem. The broker can require human approval for destructive operations. The broker can be aware of context — staging vs production, what's currently scheduled, what's been recently denied. The token never sits on disk in a form the agent can forage.
The 1Password / Doppler / HashiCorp Vault world solves this for human ops engineers. The pattern hasn't been widely translated to the agent world because the convenience pitch of agentic developer tooling is "no setup, just point it at your code." A broker between agent and credential adds friction. Friction is the thing the vendors are competing to remove. As ever, the friction was load-bearing.
Railway is the architecture story, not the AI story
I want to be careful here. Crane is right to apportion most of the blame to Railway, but the way to understand it is structural rather than accusatory. Railway are doing nothing unusual for a developer-friendly PaaS. Their authorisation model is closer to GitHub's classic personal access tokens than to AWS IAM, and it was good enough — defensible, even — back when the population of things calling the API was bounded by who could be bothered to create a token through the CLI. That population was adults, mostly, with a clear purpose in mind for each token, and a memory of what each token was for.
The population calling the API now includes Cursor agents, Claude Code agents, Codex, Copilot, Droid, OpenClaw, OpenCode, and an expanding cast of clients that read tokens from disk and call APIs at machine speed. The surface that was defensible two years ago is no longer defensible. The blast radius set by "every CLI token has full GraphQL scope" has gone from "the user does something stupid once" to "any agent the user ever pointed at this directory does something stupid up to several times per minute."
What makes Railway the architecture story is the pile-on of choices that compose with this. Tokens are root. Volumes aren't environment-scoped at the permission level — Crane's account suggests staging and production volume IDs share a name space the agent's mental model couldn't distinguish, and Railway's API didn't object. Backups live in the same volume as the data they back up. Their own documentation says, plainly, that wiping a volume deletes the backups; the documentation isn't the issue, the architecture being documented is.
And then, last week, Railway shipped mcp.railway.com, an OAuth-authenticated MCP server actively marketed at Cursor and Claude Code users. To Railway's credit — and this matters — the MCP layer explicitly excludes destructive operations from its tool set. Their docs say so. The MCP path is genuinely safer than the raw API path that bit Crane. But two things follow from this. First, the safer surface exists alongside the unsafe one, and the unsafe one is the one a foraged token reaches. Second, the MCP layer is a workaround, not a fix; the underlying architecture is unchanged, and any agent with a token and a curl can route around the MCP server entirely. Railway's CEO acknowledged on the day of the incident that what happened "1000% shouldn't be possible." It was possible because the architecture made it possible. MCP scoping doesn't help if the agent never goes through MCP.
Knight Capital, again
The story shape isn't new. Knight Capital, August 2012: $440M loss in 45 minutes, four million trades, 154 stocks, 397 million shares. Triggered by a deployment that didn't quite go to all eight servers, and a flag that was supposed to mean one thing on the new servers but reactivated dormant 2003 code on the eighth. The system did exactly what the deployment caused it to do, at machine speed, with no kill switch and no human in the loop fast enough to matter. By the time anyone noticed, the firm was effectively dead.
Knight had 97 alerts fire in the half hour before market open. Nobody read them. The thing the SEC's post-incident order makes clear, and the thing that's rarely remembered when the incident is taught as a software-deployment cautionary tale, is that the failure wasn't really about the deployment. The failure was about the gap between the speed at which automated systems can do harm and the speed at which the surrounding human and procedural infrastructure can react. The deployment was the trigger. The ungated authority and the asymmetric reaction time were the reasons it mattered.
Nine seconds for PocketOS, forty-five minutes for Knight. The blast radii are very different. The shapes are identical. Deployment-class authority sits in the hands of a fast-moving automated actor, with no kill switch in the loop, and with no architectural enforcement that would allow the human to pause the system between a wrong decision and an irreversible action. The operative variable in both cases is authority, not intelligence. A perfectly-calibrated agent with the same credentials would still have made the wrong decision under sufficiently weird inputs; an idiot agent with no credentials couldn't have caused this outcome at all. The lever you pull is on the credential side, not the intelligence side.
This is the part the AI vendor community keeps not internalising. Every one of these incidents ends the same way: a CEO post about new guardrails, a forum thread that goes quiet, and the next incident two months later with a different vendor on the marquee. Replit added dev/prod separation after July; Cursor acknowledged a Plan Mode constraint enforcement bug in December; Railway will, presumably, ship scoped tokens at some point in the next two quarters. None of those fixes addresses the underlying issue, which is that the marketing is selling capability and the engineering is shipping authority, and the two are not the same thing.
What "opinionated but escapable" looks like at this layer
I've been working on this exact problem for a stack I'm building, so I have some priors. They're not theoretical; they're the priors I'm baking into defaults because I'd rather argue with someone about the defaults later than restore from a three-month-old backup.
Tokens for any operation that can lose data should be brokered, not stored. Not every credential needs this — a Whisper API key sitting in a dotfile is, at worst, a few cents of unintended transcription, and the Marrakech case is a reasonable argument for leaving credentials of that class alone. The threshold is reversibility. Credentials whose worst-case action is recoverable can sit on disk; credentials whose worst-case action is irreversible cannot. The agent has ambient access to a token-broker tool; the broker has the actual credential and never returns it; destructive operations require explicit human approval routed through the broker, ideally on a different device. This is more friction than ~/.railway config, and more friction than a .env file. That friction is exactly the bit that prevents nine-second outages. If the friction is removed, the failure mode comes back.
Backups belong in a different trust domain. Different cloud account at minimum, different provider where the threat model warrants. Object lock or its equivalent. The Railway "backups stored in the same volume" pattern isn't sui generis — Heroku Postgres backups for many years lived next to the database they were backing up, AWS RDS automated backups can be deleted alongside the instance, and so on. The 3-2-1 rule isn't five-year-old advice. It's twenty-five-year-old advice. The arrival of agents hasn't aged it.
Customer-facing data shouldn't be physically deletable on a single API call by anything, including the agent, including the human ops engineer at three in the morning. Soft-delete with a tombstone, an explicit reaper that runs on a schedule with its own confirmation gate, an immutable audit log of every soft-delete keyed to who did it and why. This is more code. It's also the difference between Crane spending his Saturday reconciling Stripe and Crane shipping the next feature. Make the operation that takes data out of users' hands the loudest, slowest operation in the system.
Speedbumps on production-class operations need to live in the integration layer, not in the system prompt. Any API call from any agent that has the shape of a destructive op gets routed through a service that scores the risk, looks at recent history, and asks a human if it's outside policy. This service isn't the agent. The agent can't self-issue against it. The pattern is well-known from web application firewalls and from rate-limiting middleware; the application of it to agent-issued API calls is just engineering, not research.
And, finally, the agent shouldn't have read access to the directories where production credentials live. This is the single biggest behavioural change for most working setups, and it's the one most likely to be ignored, because it cuts directly into the convenience pitch. A useful default is: the agent operates in a directory that is a peer to, not a parent of, anywhere production credentials live; the agent's tool set includes a "request credentials for X" tool that goes through the broker; the broker is the only thing that ever sees the production token in cleartext. This isn't novel. It's the architecture half the security industry has been recommending for ten years. The job is making it the default for agentic developer tooling, where the default today is "give me your laptop."
The customer at the end of the chain
The thing I keep coming back to in Crane's post isn't the architectural critique. It's the bit where he describes spending Saturday morning helping his customers reconstruct bookings from Stripe payment histories, calendar integrations, and email confirmations. The customers in question are small car-rental operators whose actual customers were physically arriving at lots, expecting to pick up vehicles, and the rental owners didn't have records of who they were.
That's the thing the architecture was holding up. The gap between "agentic infrastructure" — the abstraction layer where the failure happens — and "person standing in a parking lot with a suitcase" is, on a normal day, invisible. On a bad day, the abstraction layer falls into the parking lot.
I don't mean this sentimentally. I mean it as a calibration. When we talk about the right way to wire up agents and the right way to structure backups and the right way to scope tokens, the failure mode we're trying to avoid isn't abstract data loss. It's that someone with a flight to catch can't get the car they paid for, because the SaaS the rental firm depends on, that depends on a PaaS, that depends on an agent, that depends on a credential file, lost a single API call. Every one of those layers is necessary. Every one of those layers is a potential single point of failure. The honest version of agentic infrastructure design is the one that takes this chain seriously, all the way down to the asphalt.
Crane will, I expect, do most of the right things from here on. He's been through the bad version of the thing he's now going to spend the next year arguing about with vendors and writing about on his blog. Some of his customers will go elsewhere. Some will stay. The architecture lessons will get absorbed somewhere — at Railway, eventually; at Cursor, more slowly; at Anthropic, mostly through the shape of how guardrail tooling is documented in MCP.
But the next incident is already in progress, somewhere, on a different stack with a different vendor on the marquee. The ones I'd watch in the next quarter: agent-driven Terraform in any environment where the state file lives next to the credentials; agent-driven kubectl against clusters where the kubeconfig isn't ephemeral; agent-driven CI pipelines where the deploy key has read-write to prod. None of these are exotic configurations. All of them are the agentic equivalent of leaving the master key under the doormat.
The agent didn't go rogue. The agent did exactly what the system permitted. That, as ever, is the harder lesson.