The LiteLLM Attack, in Hindsight

At 10:52 UTC on 24th March 2026, someone uploaded litellm==1.82.8 to PyPI. By 10:58, a developer at FutureSearch had it on his machine without knowing anything had changed. Six minutes later, the malware had collected his SSH keys, cloud credentials, Kubernetes config, and anything else it could reach — encrypted the haul with a 4096-bit RSA key, and posted it to models.litellm.cloud, a domain registered the previous day. The fork bomb that followed was, in a bitter irony, a bug in the malware. It brought his laptop to a halt with 11,000 Python processes, which is what tipped him off that something was wrong. Without that bug, the whole thing might have gone quietly undetected.
Seventy-two minutes after first symptoms, the compromised package had been confirmed live on PyPI via a Docker sandbox, reported to security@pypi.org, disclosed publicly, and shared across the relevant Reddit communities — all with Claude Code doing the forensic heavy lifting. That last part is interesting in its own right and I’ll come back to it, but the more durable lesson here is about the chain of failures that made the attack possible in the first place.
How it worked
The attacker compromised the PyPI account belonging to litellm’s maintainer. No corresponding GitHub tags exist for versions 1.82.7 or 1.82.8; the official release history on GitHub stops at v1.82.6. Both versions were published directly to PyPI, bypassing the project’s normal CI/CD pipeline entirely.
Version 1.82.7 embedded the payload in litellm/proxy/proxy_server.py, which at least required someone to import that module. Version 1.82.8 was more aggressive: it added a litellm_init.pth file, 34KB, placed at the top level of site-packages. A .pth file is a standard Python mechanism — it’s processed by the interpreter at startup to extend sys.path. Place executable code in one and it runs automatically, in every Python process, without any import, without any trigger. The attacker used this to turn the malware from something that fires when the library is called into something that fires when Python starts.
The payload itself was three stages. First, collection: SSH private keys and configs, .env files, AWS, GCP, Azure, and Kubernetes credentials, shell history, database passwords, crypto wallet files, anything matching common secret patterns, and a dump of all environment variables. Second, exfiltration: AES-256-CBC with a random session key, that key encrypted with the hardcoded RSA-4096 public key, the bundle POSTed to models.litellm.cloud — not litellm.ai, which is the legitimate domain. Third, persistence and lateral movement: a ~/.config/sysmon/sysmon.py backdoor with a systemd user service, and if a Kubernetes service account token was accessible, creation of privileged alpine:latest pods in kube-system to replicate the persistence onto cluster nodes.
The fork bomb came from the .pth launcher calling subprocess.Popen([sys.executable, ...]). Each child Python process inherits site-packages, finds the same .pth file, and spawns another child. The attacker had no guard against re-entry. On the FutureSearch engineer’s machine, the malware tried to install persistence at 11:07; he force-rebooted at 11:09, leaving ~/.config/sysmon/sysmon.py at 0 bytes. Two minutes either way and the story is different.
After disclosure, the attacker — still in control of the GitHub account — closed the original security issue and flooded it with hundreds of spam bot comments to suppress visibility. That detail is worth sitting with: the same compromised maintainer account that uploaded the malware was then used to bury the report. The PyPI package was eventually quarantined and the legitimate maintainers recovered control, but the window between upload and quarantine was several hours.
The patterns that enabled it
Supply chain attacks aren’t new, and this one didn’t require any novel technique. What it required was a sequence of individually unremarkable gaps that, combined, gave the attacker a clean run.
Account takeover with no downstream verification. Once the attacker had the maintainer’s PyPI credentials, there was nothing between them and publishing. PyPI has been rolling out mandatory MFA for critical packages, but the enforcement wasn’t in place for this account, or was bypassed. More importantly, even with MFA on the PyPI side, there was no mechanism for the build artefact itself to prove it came from the expected source. The wheel carries no attestation tying it to a specific GitHub repository, a specific commit, or a specific CI/CD run.
Loose transitive dependencies without lock files. The FutureSearch MCP plugin declared a dependency on litellm with an open upper bound. When uvx futuresearch-mcp-legacy ran at 10:58, it resolved to the newest available version — 1.82.8 — and installed it. That’s the design working correctly, which is part of the problem. Lock files with hash verification would have caught the new version before it ran; without them, “latest” means “whatever the registry says, unverified”.
The .pth mechanism as an execution vector. Python’s .pth file processing is a legitimately useful feature — it’s how many namespace packages and editable installs work — but it’s not widely understood outside packaging toolchain work. When the interpreter starts, it scans site-packages for files ending in .pth. Lines beginning with import are executed directly; other lines are treated as path extensions. This is entirely by design: it’s the mechanism that allows packages like pytest plugins and editable installs (pip install -e .) to wire themselves into an environment without modifying interpreter internals. A similar hook exists via sitecustomize.py and usercustomize.py — files Python will execute unconditionally at startup if they exist anywhere on sys.path. The attacker used .pth rather than these because it’s embeddable in a wheel and installs automatically with the package, whereas sitecustomize.py requires writing to the interpreter’s own site-packages directory, which typically needs elevated permissions.
From a security standpoint, .pth files are a standing promise to execute any code that ends up in site-packages at every interpreter startup, in every context — scripts, servers, test runners, CI jobs, everything. Most developers have no tooling that flags unexpected .pth additions; pip install will not warn you that a new startup-execution file just appeared in your environment, and there is no native mechanism to audit what .pth files currently exist across your virtual environments.
No out-of-band release verification. The absence of a corresponding GitHub tag for v1.82.8 was a clear signal — one that Claude Code picked up on immediately when the FutureSearch engineer was investigating. But nothing in the standard packaging toolchain checks whether a published version has a matching upstream tag. That cross-reference is currently a human-performed check, which means it only happens when someone already has reason to look.
MCP server auto-installation as an implicit trust extension. This deserves specific mention because it’s a pattern that’s going to come up repeatedly. The agentic tooling ecosystem — Claude Code, Cursor, and the various MCP plugins — has normalised installing Python packages on demand as part of normal operation. The infection vector here wasn’t a developer running pip install litellm; it was Cursor reconnecting an MCP server after a system update. The developer didn’t make a package installation decision; the tooling did, automatically, as a background task. The trust boundary has shifted, and the security model hasn’t entirely kept up.
What would have stopped it
Some of these are organisational controls, some are toolchain features, and some are environmental hygiene. None of them individually would have been sufficient; that’s somewhat the point.
PyPI Trusted Publishers (OIDC attestation). If litellm had configured PyPI Trusted Publishers — publishing via GitHub Actions OIDC rather than stored credentials — then stolen credentials alone wouldn’t have been enough. The attacker would also need to compromise the GitHub Actions workflow or the repository itself. This doesn’t eliminate the risk, but it raises the bar meaningfully. PEP 740 extends this with provenance attestations that consumers can verify. It is, at this point, the most direct control for this exact attack pattern.
Lock files with hash verification. uv.lock, poetry.lock, or pip-compile output pinned at the specific wheel hash eliminates the “install newest version” vector entirely. The malicious 1.82.8 would have been ignored because the lock file specified 1.70.4 with its expected hash. The FutureSearch engineer’s own production environment was clean for exactly this reason — it had litellm pinned at 1.70.4 with an upper bound of <1.77.3. The infection came through the MCP plugin, which wasn’t subject to the same lock.
Wheel content inspection. Scanning for .pth files, unexpected top-level executables, and outbound network calls in wheel contents is achievable with tooling like pip-audit, inspector, or custom pre-install hooks. Few teams do this routinely. It’s the kind of control that sounds obvious in retrospect and expensive in practice, which is why it mostly doesn’t happen — but in environments where developers install packages from public registries onto machines with access to production credentials, the cost calculus is probably wrong.
Egress controls on developer machines. The exfiltration relied on a POST to models.litellm.cloud succeeding from the developer’s laptop. A corporate egress proxy with allowlisting, or even DNS-based blocking of newly registered domains, would have prevented the data leaving. The persistence domain was registered the day before the attack; most organisations’ threat intel feeds would not have caught it in time, but certificate transparency monitoring for typosquats on dependency domains is a real (if rarely deployed) control.
MCP server dependency isolation. The tooling ecosystem has a pattern of running MCP servers via uvx, which creates ephemeral virtual environments but shares the same uv cache and, critically, the same host filesystem access. Running MCP servers in containers with explicit capability grants — no host home directory, no cloud credential files, outbound network limited to known endpoints — would have contained the blast radius significantly. The K8s lateral movement failed in the FutureSearch case because the malware ran on macOS with no service account token in the expected Linux paths; in a containerised CI environment with a mounted service account token, the outcome could have been substantially worse.
On the response
The disclosure was unusually fast, and the Claude Code transcript is worth reading not primarily for the AI angle but for what it shows about the forensic process. The initial hypothesis was a runaway Claude Code loop — the base64-encoded subprocess pattern is identical to how Claude Code itself passes scripts to python -c. The investigator had to push through several confident wrong answers before the malware was correctly identified. That’s normal for incident response; the tooling happened to be AI rather than grep and strace, but the structure of the investigation was familiar.
The part that was genuinely different was the speed from “confirmed” to “publicly disclosed”. From Docker confirmation to blog post, PR, and Reddit posts was under ten minutes. For a package with 47,000 weekly downloads, that window matters. Faster disclosure is good. The tradeoff is that acting quickly on an AI-assisted analysis rather than a full manual review introduces its own error modes — in this case the analysis was correct, but the precedent of treating LLM-generated forensic output as ground truth without additional verification is one to be cautious about.
The actual lesson
Supply chain attacks succeed not because defenders are incompetent but because the default configuration of the Python packaging ecosystem optimises for ease of consumption over integrity verification. Signed releases, provenance attestations, and hash-pinned lock files exist, but they’re opt-in and most packages don’t use them, most tools don’t enforce them, and most developers don’t check for them until something goes wrong.
The LiteLLM attack was a straightforward credential theft with a delivery mechanism that’s been in threat models for years. The exfiltration domain was registered the day before. The malicious versions had no GitHub tags. The .pth mechanism is documented. None of this required a sophisticated attacker.
The chain that needed to hold included: the maintainer’s account credentials, PyPI’s publishing controls, the dependency resolution policy of one MCP plugin, the absence of a lock file in that plugin’s environment, and the absence of egress controls on the developer’s machine. Any one of those holding would have broken the attack. None of them did, which is the more useful thing to take away.
If your team’s threat model doesn’t currently include “what if a package we use as a transitive dependency is compromised on PyPI”, it should. The agentic tooling wave is adding new automatic package installation surface faster than the security toolchain is adapting to it.
This is not isolated
While I was writing this, axios was compromised.
In the early hours of 31st March 2026, seven days after the litellm incident, an attacker used stolen credentials for jasonsaayman — the primary npm maintainer of axios — to publish poisoned versions 1.14.1 and 0.30.4. Axios has somewhere north of 100 million weekly downloads and is a transitive dependency in a substantial fraction of the Node.js ecosystem. The malicious versions were live for roughly two and a half hours before npm removed them. Based on Wiz’s analysis, that was long enough to get execution in approximately 3% of environments that had axios installed.
The mechanism was different from litellm’s .pth trick — npm doesn’t have that particular footgun — but the equivalent is the postinstall script hook. The attackers pre-staged a trojanised copy of the legitimate crypto-js library as plain-crypto-js@4.2.1, published 18 hours beforehand, to avoid triggering “brand-new package” alarms in static scanners. The poisoned axios releases simply added it as a runtime dependency. When npm resolves and installs a package, it runs postinstall scripts automatically; the dropper contacted a C2 server and delivered platform-specific second-stage RAT payloads — separate binaries for macOS, Windows, and Linux — then erased itself and replaced its own package.json with a clean decoy. Fifteen seconds from npm install to full compromise, with no trace in node_modules afterwards unless you knew to look for the presence of the plain-crypto-js directory itself.
The axios attack may or may not be directly connected to the group behind litellm. The litellm compromise is attributed to TeamPCP — a financially motivated threat actor who first compromised Trivy, a security scanner used in litellm’s own CI pipeline, extracted the PyPI publishing credentials from there, and used those to push the malicious versions. That’s worth reading twice: they didn’t attack litellm directly, they attacked a tool that litellm trusted to run in its build pipeline, then used what they found to attack litellm downstream. The litellm PyPI account credentials were the payload of the Trivy compromise, not its goal. Security researchers at Wiz suggest TeamPCP may be using credentials harvested from the litellm attack as initial access for further campaigns; whether axios is part of that chain or a separate actor running the same playbook remains unclear at the time of writing.
What is clear is the pattern. Compromised maintainer account. Versions published directly to the registry, bypassing CI/CD. No corresponding source tag on GitHub. Registry metadata that looks identical to a legitimate release. A startup-execution hook — .pth on Python, postinstall on npm — as the execution vector. Credential theft, persistence, and Kubernetes lateral movement as the goals.
The registries are different. The languages are different. The specific hook mechanisms are different. The underlying problem — that package registries extend implicit execution trust to anything a compromised credential can publish, and most consuming environments have no means to verify that trust is warranted — is exactly the same across all of them. PyPI has Trusted Publishers. npm has SLSA provenance attestations, which were present in axios 1.14.0 and conspicuously absent from 1.14.1. Both are available, both are opt-in, and in both cases the opted-in controls were the canary that confirmed the compromise rather than the gate that prevented it.
The threat model you actually need covers PyPI and npm and whatever comes next. The agentic tooling wave is adding new automatic package installation surface faster than the security toolchain is adapting to it, and the attacker’s side of this equation scales much better than the defender’s.
LiteLLM CVE: PYSEC-2026-2. Affected versions 1.82.7 and 1.82.8; exposure window approximately 10:52–20:15 UTC on 24th March 2026. Axios CVEs: GHSA-fw8c-xr5c-95f9 / MAL-2026-2306. Affected versions 1.14.1 and 0.30.4; exposure window approximately 00:21–03:15 UTC on 31st March 2026. In both cases: if the affected versions ran in your environment during the window, assume all credentials accessible from that environment are compromised and rotate accordingly.