Letting AI Debug Production: Helpful or Dangerous?

In March 2026, an engineer at FutureSearch opened his laptop to find it unresponsive — thousands of Python processes consuming everything the machine had. He opened Claude Code to diagnose it. For roughly twenty-five minutes, Claude gave him confident, detailed, and substantially wrong answers: the base64-encoded subprocess pattern in the newly installed litellm package looked, to Claude, exactly like Claude Code's own tool execution. That's how Claude Code passes scripts to the interpreter. Same encoding, same structure, same invocation pattern. It took persistent scepticism from the engineer — explicitly asking Claude to consider the possibility that the code was malicious rather than legitimate — before the analysis shifted. When it did, the identification was fast and correct: a supply chain attack with a credential-stealing payload hidden in a Python .pth file, exfiltrating to an attacker-controlled domain registered the previous day. From that confirmation to public disclosure took under ten minutes.
That incident is worth sitting with, because it contains most of what you need to know about AI-assisted incident response. It went well. The attack was found, disclosed, and the malicious package quarantined within hours. The AI was genuinely useful — the engineer did not need to remember the specific shell commands for pulling a fresh Docker container, or exactly how PyPI quarantine requests work, or whose email address to send to. Claude handled the procedural scaffolding while the engineer did the thinking. But the most consequential part of the investigation was a human refusing to accept the AI's first answer.
What people are actually using AI for during incidents
The honest picture of AI in incident response, as of now, is less "autonomous AI SRE" and more "better tooling for the parts that were already tedious." Log summarisation is the clearest win. An LLM can read five hundred lines of Kubernetes event logs and tell you which three things are probably relevant faster than a tired on-call engineer can scroll. Not always correctly, but fast enough that the engineer can verify rather than search. That changes the work: instead of hunting for signal, you're confirming a hypothesis. That's a better use of attention under pressure.
Runbook suggestion is similar. Given an alert type and a service context, most systems can now surface the relevant runbook, highlight the sections that match the current symptom profile, and draft the Slack message that needs to go to stakeholders. None of this requires correctness in the strong sense — the engineer still executes the runbook, still writes the message — but it removes the cognitive overhead of finding and formatting, which is surprisingly significant at 3am. Post-incident summarisation is where AI probably delivers the most reliable value, because the pressure is off, the facts are established, and the output is a document rather than an action. Generating a timeline from Slack threads and PagerDuty logs is exactly the kind of pattern-matching task that LLMs do well.
The harder question is what happens when you push beyond summarisation into action.
The overconfidence problem
The LiteLLM investigation illustrates something structural about how LLMs fail under conditions that closely resemble training data. When the malware's encoding pattern matched a pattern the model had seen thousands of times as legitimate behaviour, it produced a confident explanation in that direction. This is not a hallucination in the usual sense — it wasn't fabricating facts — it was pattern-matching correctly on the wrong hypothesis. The model had no way to know it was operating on novel attacker behaviour that had been deliberately designed to resemble legitimate tooling. That's precisely the point of a supply chain attack.
The broader principle matters here. MIT researchers in 2025 found that LLMs use more confident language when they're wrong than when they're right — roughly 34% more likely to reach for "definitely" and "certainly" when generating incorrect output. The wronger the model, the more certain it sounds. This is a fairly catastrophic property for incident response tooling, where the temptation to act on a confident summary is highest at exactly the moment when correctness matters most. A 3am on-call engineer who has just been woken up, is looking at a dashboard full of red, and has a clear, confident, step-by-step remediation plan from the AI is not in a great epistemic position to push back on it.
The failure mode here is not that the AI is bad at its job. It's that the AI is good enough to be persuasive, and the conditions under which you use it during incidents are precisely the conditions that erode your capacity to evaluate its outputs critically.
Where context collapse bites
A related problem is that the AI doesn't know what you know. A senior SRE looking at an alert carries mental context that isn't in the logs: this service has a known memory leak that manifests under specific traffic patterns; the deployment from yesterday made a subtle change to the circuit breaker timeout; the on-call team changed their escalation policy two weeks ago. None of that is in the runbook or the log stream. The AI sees the same surface the logs present, which is often a plausible picture of a different problem.
This bites hardest in the "similar incident" pattern — where AI suggests a fix because it worked last time. It's seductive because it's often right. But "this looks like that thing from November" is doing a lot of work as a reasoning chain, and the AI is not well placed to notice when the structural similarity is superficial. The November incident was a connection pool exhaustion caused by a bad release. The current incident looks the same in the logs but is actually downstream of a network partition in a different availability zone. The fix is not the same. If you've handed the AI write access and it's already applied the November playbook, you're now debugging two problems.
The read-only constraint is not bureaucratic timidity. It's the direct response to this failure mode.
Human-in-the-loop is load-bearing
There's a tendency in the current tooling discourse to frame human-in-the-loop as a transitional state — something you maintain while you build confidence, until eventually the system has enough of a track record that you let it act autonomously. That framing is probably wrong for incident response specifically, because the failure modes don't follow a power law. It's not that 95% of incidents are routine and 5% are novel. Novel incidents are more common than that, and they're the ones where AI assistance is most likely to be confidently incorrect.
The better model is that human-in-the-loop is the steady state, and the question is where the loop is. Log summarisation with human review before action is a tight loop; the human is a confirmation step before something changes. Autonomous remediation with human review of the post-incident report is a much wider loop; by the time the human is in the picture, the changes have been made and the blast radius is set. The width of the loop should be proportional to the reversibility of the action. Acknowledging an alert is reversible. Restarting a pod is generally reversible. Rolling back a database migration is not reversible, and "AI suggested it" is not a rollback strategy.
The incident.io AI SRE and similar products have converged on something close to this in practice: investigation and summarisation are automated, suggested actions are presented for human approval, and the set of actions the AI can take autonomously is tightly scoped. That's the right architecture. It also means the product is doing less of what the marketing suggests and more of what the engineering requires.
Guardrails in practice
Read-only mode for the initial investigation phase is the minimum. This means the AI can query logs, metrics, and traces; examine service configurations; pull runbooks; and search incident history. It cannot restart services, modify configurations, issue kubectl commands, or interact with deployment tooling. This is less restrictive than it sounds: most of the value during the first thirty minutes of an incident is diagnostic, not remedial, and read-only access is sufficient for everything diagnostic.
The best public example of this going wrong isn't strictly incident response, but the structural lesson transfers directly. In February 2026, Alexey Grigorev was using Claude Code to migrate infrastructure for the DataTalks.Club course platform. The setup was already slightly unsafe — he'd added new infrastructure to an existing Terraform workspace to save a few dollars a month, despite Claude explicitly warning him to keep them separate. He overrode it. When Terraform later ran without access to the state file from his old machine, it assumed the infrastructure didn't exist and began creating duplicates. He caught that, stopped the apply, and asked Claude to clean up using AWS CLI. The cleanup was going fine until Claude decided — quite reasonably, as a piece of reasoning — that terraform destroy would be cleaner than deleting resources piecemeal via CLI. Before running it, Claude had silently unpacked a Terraform archive he'd pointed it to, which replaced the current state file with the old one containing the production infrastructure. The destroy completed. The VPC, RDS instance, ECS cluster, load balancers, and automated backups were gone. Recovery took 24 hours and required AWS Business Support, which now costs him a permanent 10% surcharge.
There are a few things to notice about this. The root cause was the missing state file — he'd moved to a new machine and hadn't migrated Terraform. Claude didn't invent that problem. When he handed Claude the archive to compare resources against, Claude unpacked it as instructed, which restored the old production state, and then ran destroy against it. Each step followed from the last. The AI's reasoning was locally coherent throughout; terraform destroy is genuinely cleaner than deleting resources piecemeal via CLI when you have a state file to work from. What Claude had no way to model was the accumulated context: that the state file was wrong, that the archive contained production data, and that "clean up the duplicates" didn't mean "destroy everything the state references." The failure wasn't the AI going rogue; it was a chain of individually reasonable decisions compounding into an irreversible outcome because write access was granted without a shared understanding of the blast radius. Claude had warned him about the initial risky decision. He'd dismissed it.
This is also the case that refutes the "AI would be fine if you just supervise it" argument. He was supervising it. He was present, actively engaged, and approving each step. The problem was that he didn't have enough context to know which steps needed more scrutiny.
Scoped actions are the next layer. When you do grant the AI the ability to act, the scope should be defined by service and operation type, not by "the AI will figure out what's appropriate." The AI can restart pods in the dev namespace. It cannot touch prod. It can roll back to the previous deployment tag for service X. It cannot modify IAM policies. The scope should be written down as code, enforced by the tooling, and reviewed periodically rather than growing organically as someone adds "just one more" capability.
Audit logging of AI actions during incidents is underused. If the AI queries your logs, summarises a hypothesis, and the engineer acts on it, you want a record of what the AI said and what the engineer did with it. This is useful for post-incident analysis but also for calibrating the system over time — if the AI's initial hypothesis was wrong in three of the last five similar incidents, that's a signal about the reliability of that class of suggestion.
Post-incident analysis of AI involvement
The FutureSearch transcript is public, which makes it unusual. Most AI-assisted incident investigations leave no record of what the AI said during the investigation — only the final outcome. That's a problem for learning from incidents, because the near-miss cases are as informative as the successes. Knowing that Claude was confident and wrong for twenty-five minutes before identifying a novel supply chain attack is useful operational data. It tells you something about the kind of novelty that creates blind spots, and it suggests that "has the AI considered a malicious explanation" is a reasonable checklist item when the symptom pattern is unusual.
Post-incident reviews should include the AI's contributions as a first-class element — not just "we used Claude Code to investigate" but "the AI's initial hypothesis was X, we accepted/rejected it because Y, and the actual cause was Z." If the AI was right and the human accepted it without verification, that's worth noting as a process risk even if the outcome was fine. If the AI was wrong and the human caught it, that's the loop working as intended, and you want to understand what enabled the catch. The engineer at FutureSearch caught it because he was sceptical that a routine investigation would turn up an undisclosed supply chain attack — he knew the base rate. That kind of Bayesian reasoning is not something you can offload.
The meta-layer
There's a detail in the LiteLLM case that didn't get as much attention as it deserved. The attack itself used an AI agent. TeamPCP's toolchain included something called hackerbot-claw, an automated attacker that uses an AI agent for targeting. The attack on LiteLLM was phase nine of a coordinated campaign that also compromised the Trivy security scanner, exfiltrated the LiteLLM maintainer's PyPI publish token from a GitHub Actions runner, and used consistent infrastructure — same RSA key pair, same exfiltration bundle format — across all operations. This wasn't a one-off opportunistic attack. It was systematic, and the AI component was on the offensive side.
The defensive tooling is getting more capable at roughly the same pace as the offensive tooling. That's roughly what you'd expect, and it means the question "should I use AI in incident response?" is probably already settled — not using it is also a choice with costs, and those costs are increasing. The question is how you use it, and the answer starts with: read-only first, human in the loop for actions, treat AI-generated hypotheses as hypotheses rather than conclusions, and log everything the AI does during an incident so you can audit it afterwards.
A frozen laptop with eleven thousand Python processes was, in this case, enough of a signal. In a containerised CI environment with proper egress controls, the malware would have exfiltrated silently and nobody would have found the .pth file until much later. The bug in the malware that triggered disclosure was not something you can rely on. The process that found the attack despite the AI's early misdirection — persistent scepticism, willingness to reject the confident answer, domain knowledge about what's plausible — that's what you can build on.