Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Your CI Runner Said privileged: false. It Was Still Root.

On 2 October a critical vulnerability, CVE-2026-73802, was published for the runner that executes Gitea Actions jobs. It scores 9.9 on CVSS, which in practice means network-reachable, low complexity, low privileges, and a changed scope: the attacker starts inside a job and ends up outside it. The weakness is filed as CWE-269, improper privilege management. That is accurate and tells you almost nothing, so here is the mechanism.

A workflow can set jobs.<job>.container.options, a free-text field of extra flags for the job's container. The runner takes those options and builds them into the Docker HostConfig it hands to the daemon. Runner operators can also disable privileged mode, and the runner honoured that: whatever the workflow asked for, the Privileged flag in the finished config was forced to false.

Forced to false. Not "the config was validated". The one field the operator had a switch for was overwritten, and everything else the workflow author supplied was copied through. Host PID and IPC namespaces, additional capabilities, security-profile overrides: all of it survived into the final configuration. The advisory's conclusion is that a workflow author could use this to enter host namespaces and run commands on the runner host as root.

The affected range is every version before 1.0.9-0.20260731160927-34bfa1915022. The advisory's recommended action is to upgrade, and I'd do that before reading further.

Docker did nothing wrong

It is tempting to file this under container escapes, and the instinct is to look for the clever bit. There isn't one. Nobody found a kernel bug or a runc race or a clever symlink. The daemon received a configuration that said "share the host's PID namespace and add these capabilities", and it did precisely that, because that is what a trusted caller asking for those things gets.

The trusted caller was the runner. The runner had translated a string from a repository file into a request with root-equivalent authority, and the only filter on the way was one boolean.

This is the shape worth remembering:

workflow YAML          (written by anyone who can push a branch)
      ↓
container.options     (free text)
      ↓
runner builds HostConfig   ← the trust boundary was supposed to be here
      ↓
Docker daemon         (trusts its caller, as designed)
      ↓
host namespaces, capabilities, security profile
      ↓
runner host

CI systems spend their whole existence converting low-trust input into high-authority operations. A repository file is editable by a contributor, sometimes by a pull request from a fork, and the thing consuming it holds credentials, a container runtime, and usually a network position inside your estate. Every field the runner forwards is a place where the author's intent becomes the host's behaviour. The runner's job is to treat that file as hostile input, and for these fields it did not.

privileged is not a boolean, it's a vector

The mistake underneath this one is conceptual, and I have made a version of it myself. privileged: false reads like a statement about the container's isolation. It is a statement about one flag.

A container's effective privilege is the sum of several independent things: the Privileged flag (which grants nearly all capabilities, all devices, and relaxes the confinement profiles in one go); the capability set, which can be widened without the flag; the namespaces it shares with the host, PID, IPC, network, and UTS among them; devices; bind mounts, including the daemon's own socket; the seccomp profile; the AppArmor or SELinux profile; and whether root inside the container is root outside it, which is the user-namespace question. Each of those can be tightened or loosened separately.

Setting one of them to the safe value neutralises that one. The others carry on. A container with Privileged false, --pid=host, and CAP_SYS_ADMIN added is, for most purposes, not meaningfully contained, and a container with the Docker socket mounted needs none of the rest. Which is why "we disabled privileged mode" is an answer to a narrower question than the one the person asking usually has in mind.

The same pattern turns up in code review as the difference between a denylist and an allowlist. The runner, in effect, had a denylist of one entry. An allowlist would have taken the small set of options a build legitimately needs and rejected the rest, and then this CVE would be a feature request about supporting --shm-size.

The question to ask of your runners

Here is the audit I'd run on any self-hosted runner, whatever it is and wherever it's hosted:

If an attacker can modify a workflow, what stops that workflow becoming the runner host?

Then try to prove it, from inside a job, rather than reading the runner's config and feeling reassured. It's a short list.

Namespaces. Compare readlink /proc/self/ns/pid (and ipc, net, uts) inside the job with the same on the host. If /proc/1/cmdline inside your job container shows the host's init rather than your entrypoint, you are in the host's PID namespace and you can see, and in the right circumstances signal or nsenter into, everything on the machine.

Capabilities. grep Cap /proc/self/status, then decode with capsh --decode. You are looking for anything beyond the default container set. CAP_SYS_ADMIN is the one that ends the conversation.

Confinement. grep Seccomp /proc/self/status should say a filter is active, and cat /proc/self/attr/current should show the AppArmor or SELinux profile you expect. An unconfined there was somebody's decision, and it should be one you can name.

Mounts and sockets. Is /var/run/docker.sock, or a Podman socket, reachable from the job? Anything that can talk to the daemon can ask it for a privileged container, so the flag-level hardening above it is moot. Check mount for anything from the host that isn't obviously yours, and look at what /proc and /sys are exposing and whether they are writable.

Devices. ls /dev against what a plain container shows. Extra device nodes are a direct line to the hardware.

User namespace. cat /proc/self/uid_map. If uid 0 inside maps to uid 0 outside, then root in the container is root on the host, merely wearing fewer capabilities. Rootless runtimes and user-namespace remapping change that, and it is among the cheapest mitigations available.

Network position. What can the job reach? Cloud metadata endpoints, the internal registry, the admin interface of the CI server, the management VLAN. The runner host's network is usually far more privileged than the repository's authors.

Credentials. What is in the job's environment, in files on the runner, in the runner's own registration token? Which of those could be used to do something to a repository other than the one that triggered the job?

Persistence. Does anything survive from one job to the next: workspace directories, caches, a long-lived daemon with a warm layer cache, a modified ~/.bashrc? A job that lands on a long-lived host can leave something behind for the next job, and the next job may belong to a different team.

I haven't run this list against every runner implementation, and it would be a poor use of anyone's time to take my word for what GitHub's runner or GitLab's executors do with option strings. container.options exists in more than one CI system, and what each does with it is a matter of reading the code and then testing the behaviour with a throwaway workflow on a throwaway runner. Do that. The point of the exercise is the habit, not the verdict.

Disposable per trust boundary

You can harden a long-lived runner, and you should, but it is the wrong place to put your confidence. A shared machine running jobs from mutually distrusting authors has to get every one of the vectors above right, every time, in every release, including the ones nobody has found yet. This CVE is what happens when one of them slips.

The architecture that survives a slip is a runner that is created for a job (or for a trust boundary), runs it, and is destroyed. An ephemeral VM, ideally, since a VM boundary is a considerably harder thing to cross than a namespace one; an ephemeral rootless container as the cheaper alternative. If the runner is a throwaway, then "the workflow became the host" costs you one throwaway host, with nothing on it worth taking and no neighbours to reach. The trust boundary is then wherever you decide to put the disposal line: per job for anything that builds untrusted code, per repository or per team for things you trust a bit more, never one pool for everybody because it was convenient.

Pair that with scoped, short-lived credentials (OIDC federation where the target supports it) so that the runner has little to lose in the first place. The best credential to find on a compromised runner is one that has already expired.

What to do this week

Upgrade the runner to the patched version. Then, until you are sure everything has been upgraded, consider rejecting container.options wholesale at the workflow-validation layer if your setup lets you; almost nobody needs arbitrary Docker flags in a build job. Check who can push to the repositories that your runners serve, including forks and pull-request triggers, because that is the population who could have exploited this. And if a runner of an affected version sat exposed to contributors you don't fully trust, treat its credentials as potentially exposed and rotate them. I'd rather say that than assume nobody noticed.

Then run the audit above on whatever you have left, because the next one will not be this field.

privileged: false was a true statement. It was also an incomplete one, and the gap between the two was exactly the size of the vulnerability.