Podman Quadlets and systemd
Declarative systemd integration for Podman: Quadlet unit types, multi-container pods, networks, timers, sdnotify, and healthchecks.
Podman Quadlets and systemd
Declarative systemd integration for Podman: Quadlet unit types, multi-container pods, networks, timers, sdnotify, and healthchecks.
Overview
Quadlet is a systemd generator that turns short, declarative unit files — .container, .pod, .network, .volume, .kube, .build, .image — into ordinary systemd service units at boot and on every daemon-reload. You write intent; Quadlet writes the ExecStart=podman run … line, the dependency graph, the stop and cleanup commands, and the sdnotify plumbing.
It replaces podman generate systemd, which is deprecated (bug fixes only, no new features). Quadlet is the supported way to run containers under systemd, and it is what makes a single-node Podman host behave like a small, self-healing orchestrator: ordered start-up, restart policies, health-gated readiness, journald logging, scheduled jobs, and unattended image updates.
flowchart TB
A["web.container<br/>app.pod<br/>app.network"] --> B["Quadlet generator<br/>podman-system-generator"]
B -->|"boot and daemon-reload"| C["Generated units in /run/systemd/generator"]
C --> D["web.service<br/>app-pod.service<br/>app-network.service"]
D --> E["systemd"]
E -->|"ExecStart"| F["podman run / podman pod create"]
F --> G["conmon"]
G --> H["Container process"]
H -.->|"sd_notify READY=1"| E
E -.->|"journald"| I["journalctl -u web.service"]
Version map
Quadlet has moved fast. Everything below is written against Podman 5.4+ unless flagged; Podman 6.1 is current at time of writing.
| Version | Released | What landed |
|---|---|---|
| 4.4 | 2023-02 | Quadlet arrives: .container, .volume, .network, .kube |
| 4.7 | 2023-09 | podman generate systemd deprecated; VolumeName=, NetworkName= |
| 4.8 | 2023-11 | .image units |
| 5.0 | 2024-03 | .pod units; Notify=healthy |
| 5.2 | 2024-08 | .build units |
| 5.3 | 2024-11 | ServiceName=, StartWithPod=, [Quadlet] DefaultDependencies= |
| 5.6 | 2025-08 | podman quadlet CLI (install, list, print, rm); .pod ExitPolicy=, Label= |
| 6.0 | 2026-06 | CNI, slirp4netns, cgroups v1 and iptables support removed; a user-set Restart= in a .pod file is finally honoured; Quadlet man pages split per unit type; /usr/share/containers/systemd/users search paths |
| 6.1 | 2026-08 | ImageVolume= on .container |
Podman 6.0 dropped several things that older guides still assume. If you are moving a host from 5.x, check these first:
# cgroups v2 is now mandatory (it was already required by Quadlet)
podman info --format '{{.Host.CgroupsVersion}}' # must print v2
# netavark is the only network backend; CNI is gone
podman info --format '{{.Host.NetworkBackend}}' # netavark
# pasta is the only rootless network stack; slirp4netns is gone
podman info --format '{{.Host.Slirp4NetNS.Executable}}' # empty on 6.x
# nftables only — iptables-legacy hosts need migrating before upgrade
nft list ruleset | head
Quadlet Fundamentals
Unit types
Each file type owns one section ([Container], [Pod], …) which Quadlet consumes; every other section ([Unit], [Service], [Install]) is passed through to systemd untouched.
| File | Section | Generated service | Podman resource name | Systemd Type |
|---|---|---|---|---|
foo.container |
[Container] |
foo.service |
container systemd-foo |
notify |
foo.pod |
[Pod] |
foo-pod.service |
pod systemd-foo |
forking |
foo.network |
[Network] |
foo-network.service |
network systemd-foo |
oneshot |
foo.volume |
[Volume] |
foo-volume.service |
volume systemd-foo |
oneshot |
foo.image |
[Image] |
foo-image.service |
(pulled image) | oneshot |
foo.build |
[Build] |
foo-build.service |
(built image) | oneshot |
foo.kube |
[Kube] |
foo.service |
pods from the YAML | notify |
The systemd- prefix on resource names is deliberate: it keeps generated resources from colliding with ones you created by hand. Override it with ContainerName=, PodName=, NetworkName=, or VolumeName=. Override the unit name with ServiceName=.
Gotcha: a
.podfile producesfoo-pod.service, notfoo.service.systemctl --user start foowill fail; you wantsystemctl --user start foo-pod.
Where files go
Rootful, in precedence order (/run beats /etc beats /usr):
/run/containers/systemd/ # transient, for testing
/etc/containers/systemd/ # sysadmin-managed
/usr/share/containers/systemd/ # distribution-packaged
Rootless:
$XDG_RUNTIME_DIR/containers/systemd/ # transient
~/.config/containers/systemd/ # the usual place
/etc/containers/systemd/users/${UID} # admin-managed, one user
/etc/containers/systemd/users/ # admin-managed, all users
/usr/share/containers/systemd/users/${UID} # distro-packaged, 6.0+
/usr/share/containers/systemd/users/ # distro-packaged, 6.0+
Subdirectories are searched recursively, and symlinks work both for the search-path roots and for files inside them.
The generator loop
# Rootless: drop the file in, reload, start
mkdir -p ~/.config/containers/systemd
$EDITOR ~/.config/containers/systemd/web.container
systemctl --user daemon-reload
systemctl --user start web
# Rootful
sudo $EDITOR /etc/containers/systemd/web.container
sudo systemctl daemon-reload
sudo systemctl start web
# See what Quadlet actually generated, without installing anything
/usr/lib/systemd/system-generators/podman-system-generator --user --dryrun
Enabling on boot
Generated units are transient — systemd will not let you systemctl enable them, because the unit file does not exist on disk between reloads. Quadlet applies the [Install] section itself at generation time, which is the equivalent:
[Install]
WantedBy=default.target
Only Alias=, WantedBy=, RequiredBy=, UpheldBy= and (for templates) DefaultInstance= are honoured. For rootful units, multi-user.target is the conventional target; default.target is right for user units.
# This will NOT work on a Quadlet-generated unit
systemctl --user enable web.service
# Failed to enable unit: Unit /run/user/1000/systemd/generator/web.service
# is transient or generated
Drop-ins
Quadlet supports systemd-style drop-ins on the source file, which is how you keep a packaged unit intact while changing one value:
mkdir -p ~/.config/containers/systemd/web.container.d
cat > ~/.config/containers/systemd/web.container.d/10-image.conf <<'EOF'
[Container]
Image=docker.io/library/nginx:1.29-alpine
EOF
systemctl --user daemon-reload
Merge order is alphabetical, and prefix directories apply too: for web-front-eu.container, Quadlet reads container.d/, web-.container.d/, web-front-.container.d/, then web-front-eu.container.d/, each overriding the last.
Container Units
A minimal unit
# ~/.config/containers/systemd/web.container
[Unit]
Description=Front-end web server
[Container]
Image=docker.io/library/nginx:1.29-alpine
PublishPort=8080:80
Volume=%h/site:/usr/share/nginx/html:ro,z
[Service]
Restart=always
[Install]
WantedBy=default.target
What that generates
Reading the generated unit is the fastest way to debug Quadlet, and worth doing once so the defaults stop being mysterious:
[Unit]
Wants=podman-user-wait-network-online.service
After=podman-user-wait-network-online.service
Description=Front-end web server
SourcePath=/home/mike/.config/containers/systemd/web.container
RequiresMountsFor=%t/containers
[X-Container]
Image=docker.io/library/nginx:1.29-alpine
PublishPort=8080:80
Volume=%h/site:/usr/share/nginx/html:ro,z
[Service]
Restart=always
Environment=PODMAN_SYSTEMD_UNIT=%n
KillMode=mixed
ExecStop=/usr/bin/podman rm -v -f -i systemd-%N
ExecStopPost=-/usr/bin/podman rm -v -f -i systemd-%N
Delegate=yes
Type=notify
NotifyAccess=all
SyslogIdentifier=%N
ExecStart=/usr/bin/podman run --name systemd-%N --replace --rm --cgroups=split --sdnotify=conmon -d -v %h/site:/usr/share/nginx/html:ro,z --publish 8080:80 docker.io/library/nginx:1.29-alpine
[Install]
WantedBy=default.target
Points worth internalising:
--replace --rm— the container is recreated from scratch on every start. Anything not in a volume is gone. This is a feature: the unit file is the source of truth.Type=notifywithNotifyAccess=all— systemd waits for a readiness message before considering the service started, and before starting anything orderedAfter=it.--cgroups=split— the container's cgroup lives under the service's cgroup, sosystemd-cgtopandsystemctl statusshow real resource usage.- Specifiers are passed through, not resolved.
%h,%t,%Nand friends reach the generated unit verbatim; systemd expands them at start. That is why relative paths beginning with%need a./prefix — Quadlet cannot tell them from a specifier. - The
[X-Container]block is a copy of your source section, kept for reference. Only[Unit],[Service]and[Install]do anything. - Rootful units get
network-online.targetinstead ofpodman-user-wait-network-online.service; user units cannot wait on system targets, hence the shim.
Common [Container] keys
The full list runs to about 120 keys (man podman-systemd.unit); these are the ones that carry most of the weight.
| Key | podman run equivalent |
Notes |
|---|---|---|
Image= |
image argument | Required. Use a fully-qualified name |
ContainerName= |
--name |
Defaults to systemd-$unitname |
Exec= |
command after the image | Overrides CMD |
Entrypoint= |
--entrypoint |
|
PublishPort= |
--publish |
Repeatable |
Volume= |
--volume |
Accepts foo.volume:/dest to reference a .volume unit |
Mount= |
--mount |
Long-form syntax |
Network= |
--network |
Accepts foo.network and foo.container |
Environment= |
--env |
Repeatable |
EnvironmentFile= |
--env-file |
Relative paths need ./ if they start with % |
Secret= |
--secret |
Secret=name,type=env,target=VAR |
Pod= |
--pod |
Must name a .pod unit, not a bare pod name |
AutoUpdate= |
--label io.containers.autoupdate=… |
registry or local |
Notify= |
--sdnotify |
false (default), true, healthy |
HealthCmd= and friends |
--health-* |
See Health Checks |
User= / Group= |
--user |
|
UserNS= |
--userns |
keep-id, auto, nomap |
ReadOnly= |
--read-only |
|
NoNewPrivileges= |
--security-opt no-new-privileges |
|
DropCapability= / AddCapability= |
--cap-drop / --cap-add |
Space-separated, repeatable |
PodmanArgs= |
anything else | Escape hatch, appended before the image name |
PodmanArgs= exists for options Quadlet has no key for. It is unvalidated and can silently conflict with generated flags, so treat it as a last resort rather than a habit:
[Container]
Image=docker.io/library/redis:7-alpine
PodmanArgs=--memory-reservation=256m --oom-score-adj=-500
Ordering between units
Quadlet rewrites Podman-flavoured dependencies in [Unit] into real service dependencies. Reference the source file name, not the generated service name:
[Unit]
Requires=database.container
After=database.container
Wants=, Requires=, Requisite=, BindsTo=, PartOf=, Upholds=, Conflicts=, Before= and After= are all translated. Because container units are Type=notify, After= genuinely means "after the dependency reported ready" — which is only as truthful as its readiness signal (see sdnotify).
Multi-Container Pods
A pod is a set of containers sharing namespaces — by default ipc, net and uts, matching Kubernetes. Containers in a pod reach each other over localhost, and ports are published on the pod, not on individual containers.
graph TB
subgraph pod["Pod: systemd-app (app-pod.service)"]
I["Infra container<br/>holds the namespaces"]
W["web<br/>nginx :80"]
A["api<br/>uvicorn :9000"]
C["cache<br/>valkey :6379"]
end
H["Host :8080"] --> I
W -.->|"localhost:9000"| A
A -.->|"localhost:6379"| C
N[("app.network")] --> I
V[("app-data.volume")] --> W
A complete worked example
Four files describing one application. Ports are declared once, on the pod.
# ~/.config/containers/systemd/app.network
[Network]
NetworkName=app
Driver=bridge
Subnet=10.89.7.0/24
Gateway=10.89.7.1
Label=app=demo
# ~/.config/containers/systemd/app-data.volume
[Volume]
VolumeName=app-data
Copy=true
Label=app=demo
# ~/.config/containers/systemd/app.pod
[Unit]
Description=Demo application pod
[Pod]
PodName=app
PublishPort=8080:80
Network=app.network
Volume=app-data.volume:/srv/shared:z
[Service]
Restart=always
[Install]
WantedBy=default.target
# ~/.config/containers/systemd/api.container
[Unit]
Description=API service
[Container]
Image=localhost/api:latest
Pod=app.pod
Environment=PORT=9000
Notify=true
HealthCmd=curl -fsS http://localhost:9000/healthz
HealthInterval=30s
HealthStartPeriod=20s
HealthOnFailure=kill
# ~/.config/containers/systemd/web.container
[Unit]
Description=Web front end
After=api.container
Requires=api.container
[Container]
Image=docker.io/library/nginx:1.29-alpine
Pod=app.pod
Volume=app-data.volume:/usr/share/nginx/html:ro,z
HealthCmd=curl -fsS -o /dev/null http://localhost/
HealthInterval=30s
HealthOnFailure=kill
Notify=healthy
systemctl --user daemon-reload
systemctl --user start app-pod # note the -pod suffix
systemctl --user status app-pod api web
podman pod ps
How the dependency graph is wired
Quadlet generates this automatically from Pod=:
flowchart LR
N["app-network.service"] --> P["app-pod.service"]
V["app-data-volume.service"] --> P
P -->|"Wants + Before"| A["api.service"]
P -->|"Wants + Before"| W["web.service"]
A -.->|"BindsTo + After"| P
W -.->|"BindsTo + After"| P
A -->|"api first, because web.container<br/>declares Requires + After"| W
- The pod unit gets
Requires=/After=on its network and volume units, andWants=/Before=on each member container. - Each member gets
BindsTo=/After=the pod, so stopping the pod stops its containers, and a pod failure takes the containers with it. - Ordering between members is yours to declare, as
web.containerdoes above.
Pod-specific keys
| Key | Effect |
|---|---|
PodName= |
Pod name; defaults to systemd-$unitname |
PublishPort= |
Publish on the infra container — the only place ports work in a pod |
Network= / NetworkAlias= |
Pod-level networking |
Volume= |
Mounted into every container that joins |
ExitPolicy= |
stop (Quadlet default) or continue |
UserNS=, UIDMap=, GIDMap= |
User-namespace mapping for the whole pod |
PodmanArgs= |
--cpus, --memory, --share, and other podman pod create options |
Two behaviours catch people out:
ExitPolicy=stop versus Restart=on-failure. Quadlet pods default to ExitPolicy=stop, so when the last container exits the pod stops cleanly with exit code 0. The generated Restart=on-failure therefore never fires. If you want the pod to come back, set Restart=always explicitly in [Service]. (On Podman before 6.0, a Restart= in a .pod file was overwritten with on-failure regardless — if you are on 5.x and a pod refuses to restart, this is why.)
StartWithPod=. Defaults to true: members start, stop and restart with the pod. Set it to false for a container that should be started on demand — a migration job, a debug shell — while still being torn down with the pod:
[Container]
Image=localhost/api:latest
Pod=app.pod
StartWithPod=false
Exec=/app/migrate.sh
[Service]
Type=oneshot
Imperative pods
Still useful for exploration and for generating YAML:
# Create a pod; ports and limits belong to the pod, not its containers
podman pod create --name app \
--publish 8080:80 \
--network app \
--cpus 2 --memory 1g \
--share ipc,net,uts \
--exit-policy stop
podman run -d --pod app --name api localhost/api:latest
podman run -d --pod app --name web docker.io/library/nginx:1.29-alpine
podman pod ps
podman pod top app
podman pod logs -f app # interleaved, all containers
podman pod stats app
podman pod inspect app --format '{{.InfraContainerID}}'
# Add PID namespace sharing (so containers can see each other's processes)
podman pod create --name app --share +pid
.kube units for larger sets
When an application is already described in Kubernetes YAML, skip the per-container units:
# Round-trip: build it imperatively, then capture it
podman kube generate app > ~/.config/containers/systemd/app.yaml
# Healthchecks are emitted as a livenessProbe from 6.1 onwards
podman kube generate --type deployment --replicas 3 app > app.yaml
# ~/.config/containers/systemd/app.kube
[Unit]
Description=Application from Kubernetes YAML
[Kube]
Yaml=app.yaml
ConfigMap=app-config.yaml
PublishPort=8080:80
AutoUpdate=registry
ExitCodePropagation=any
KubeDownForce=true
SetWorkingDirectory=yaml
[Install]
WantedBy=default.target
That generates podman kube play --replace --service-container=true … with a matching podman kube down on stop. ExitCodePropagation= decides how container failures surface to systemd — any, all, or none (the default, which hides them).
Podman supports Deployments, DaemonSets, Jobs, PVCs, ConfigMaps and Secrets, but not Services, Ingress, or anything requiring a scheduler. .kube is for describing a workload, not for pretending to be a cluster.
Networks, Volumes, Images, and Builds
.network units
The point of a .network unit is not just creating the network, but making containers depend on it, with the options you chose rather than Podman's defaults.
# frontend.network
[Network]
NetworkName=frontend
Driver=bridge
Subnet=10.89.0.0/24
Gateway=10.89.0.1
IPRange=10.89.0.128/25
IPv6=true
Internal=false
Options=isolate=true
Label=tier=frontend
# In a container that should join it
[Container]
Network=frontend.network
| Key | Notes |
|---|---|
NetworkName= |
Without it, the network is called systemd-frontend |
Internal=true |
No outbound routing — the right default for database tiers |
Options=isolate=true |
netavark blocks traffic to other networks |
DisableDNS=true |
Turns off container-name resolution (on by default for named networks) |
NetworkDeleteOnStop=true |
Delete the network when the unit stops |
IPAMDriver= |
host-local (default), dhcp, or none |
Gotcha: editing a
.networkfile does not reconfigure an existing network. The generated unit runspodman network create --ignore, which is a no-op if the network exists. To apply a subnet change:systemctl --user stopthe consumers,podman network rm frontend, then start again — or setNetworkDeleteOnStop=trueup front.
Two special forms of Network= are worth knowing:
Network=frontend.network # join a Quadlet-managed network, with ordering
Network=sidecar.container # share another container's network namespace
Network=host # host networking, no ordering added
Network=none # no networking at all
The sidecar.container form is how you build a sidecar without a pod — useful when you want two containers on one network stack but different lifecycles.
Systemd-level network ordering
Quadlet adds network-online.target (rootful) or podman-user-wait-network-online.service (rootless) to every generated unit, so image pulls do not race the network coming up. Disable it when a unit must start early and pulls nothing:
[Quadlet]
DefaultDependencies=false
Note the section: DefaultDependencies= in [Quadlet] controls Quadlet's network ordering, whereas the identically-named key in [Unit] is systemd's own and means something else entirely.
For hosts with a specific interface that must be up first, order against systemd-networkd's per-interface unit:
[Unit]
After=sys-subsystem-net-devices-eth1.device
Wants=sys-subsystem-net-devices-eth1.device
.volume units
# app-data.volume
[Volume]
VolumeName=app-data
Copy=true
UID=1000
GID=1000
Label=app=demo
# Bind-mount style with a device
[Volume]
VolumeName=nfs-data
Driver=local
Type=nfs
Device=nfs.example.com:/exports/data
Options=rw,noatime,rsize=8192,wsize=8192,nfsvers=4
Reference from a container with the unit name so ordering is generated:
[Container]
Volume=app-data.volume:/var/lib/app:z
Volume units are oneshot with RemainAfterExit=yes; the underlying command is podman volume create --ignore, so they are idempotent and never destroy data.
.image and .build units
.image pulls ahead of time, which keeps the first start of a container off the network:
# base.image
[Image]
Image=docker.io/library/alpine:3.22
Policy=newer
Retry=3
RetryDelay=5s
.build builds from a Containerfile, so a container can depend on an image that exists nowhere but this host:
# site.build
[Build]
ImageTag=localhost/site:latest
File=/srv/site/Containerfile
SetWorkingDirectory=unit
Label=org.opencontainers.image.source=https://git.example.com/site
Pull=newer
# The consuming container
[Container]
Image=site.build
Referencing site.build (rather than localhost/site:latest) is what generates Requires=site-build.service. Layer caching makes subsequent starts fast, but [Service] TimeoutStartSec= still needs raising for a first build.
sdnotify and Readiness
The problem
Restart=always and After= are only useful if systemd knows when a service is actually ready. Without a readiness protocol, "started" means "the process was spawned", and a dependent service starts against a database that has not finished recovery.
Podman supports systemd's sd_notify in four modes:
--sdnotify= |
MAINPID |
READY=1 sent when |
Socket reaches container |
|---|---|---|---|
container (CLI default) |
conmon | the app inside calls sd_notify |
yes |
conmon (Quadlet default) |
conmon | the container has started | no |
healthy |
conmon | the container passes its healthcheck | no |
ignore |
— | never; NOTIFY_SOCKET is stripped |
no |
Observed on the wire with --sdnotify=conmon, a datagram arrives on NOTIFY_SOCKET about 0.2s after podman run:
MAINPID=3952
READY=1
Quadlet mapping
Quadlet's Notify= key selects the mode, and always sets Type=notify plus NotifyAccess=all:
| Quadlet | Generated flag | Use when |
|---|---|---|
(unset) or Notify=false |
--sdnotify=conmon |
The workload has no readiness signal |
Notify=true |
--sdnotify=container |
The app calls sd_notify itself |
Notify=healthy |
--sdnotify=healthy |
You have a healthcheck and no in-app signal |
[Service] Type=oneshot |
none — no sdnotify, no -d |
Batch jobs that exit |
Note the inversion: the Podman CLI defaults to container, Quadlet defaults to conmon. The CLI default makes systemd wait for a message that most images never send; Quadlet picks the safe one. podman generate systemd did the same rewriting for the same reason.
Notifying from inside the container
The container needs Notify=true, and the application needs to write to $NOTIFY_SOCKET. Podman passes NOTIFY_SOCKET — and LISTEN_FDS, LISTEN_PID, LISTEN_FDNAMES for socket activation — through to the container process.
Shell, for images that ship systemd-notify:
#!/bin/sh
set -eu
/app/warm-caches.sh
systemd-notify --ready --status="serving"
exec /app/server
Python, with no dependency beyond the standard library:
"""Minimal sd_notify client for containers run with --sdnotify=container."""
from __future__ import annotations
import logging
import os
import socket
logger = logging.getLogger(__name__)
def sd_notify(state: str, *, unset_env: bool = False) -> bool:
"""Send a service-manager notification.
Args:
state: Newline-separated protocol assignments, e.g. "READY=1"
or "STATUS=warming caches". See sd_notify(3).
unset_env: Remove NOTIFY_SOCKET from the environment after sending,
so forked children cannot notify on the service's behalf.
Returns:
True if the datagram was sent, False if there was no socket to send
it to (i.e. the process is not running under a notify-type unit).
Raises:
OSError: The socket exists but could not be written to. Callers
generally want to log and continue rather than abort start-up.
"""
addr = os.environ.get("NOTIFY_SOCKET")
if not addr:
logger.debug("NOTIFY_SOCKET unset; not running under systemd notify")
return False
try:
# A leading '@' denotes the abstract namespace, encoded as a NUL byte.
path = "\0" + addr[1:] if addr[0] == "@" else addr
with socket.socket(
socket.AF_UNIX, socket.SOCK_DGRAM | socket.SOCK_CLOEXEC
) as sock:
sock.connect(path)
sock.sendall(state.encode("utf-8"))
logger.debug("sd_notify sent: %r", state)
return True
finally:
if unset_env:
os.environ.pop("NOTIFY_SOCKET", None)
def main() -> None:
logging.basicConfig(level=logging.INFO)
warm_caches() # anything slow that must finish before traffic
try:
sd_notify("READY=1\nSTATUS=serving")
except OSError:
# Never let a notification failure take down the service; systemd
# will time out and restart us if this was genuinely fatal.
logger.warning("could not notify systemd of readiness", exc_info=True)
serve_forever()
Useful assignments beyond READY=1:
| Assignment | Meaning |
|---|---|
READY=1 |
Start-up complete |
RELOADING=1 |
Reload in progress; pair with MONOTONIC_USEC= |
STOPPING=1 |
Shutdown started |
STATUS=… |
Free text shown by systemctl status |
WATCHDOG=1 |
Keep-alive ping when WatchdogSec= is set |
EXTEND_TIMEOUT_USEC=… |
Ask for more start-up time |
If the app never sends READY=1, the unit sits in activating (start) until TimeoutStartSec (90s by default) expires, and systemd then kills it. Raise it deliberately rather than guessing:
[Service]
TimeoutStartSec=300
Image pulls are already covered: from 4.7, Podman sends EXTEND_TIMEOUT_USEC=30s every 25 seconds while pulling, up to ten times, buying five minutes beyond TimeoutStartSec without any configuration. Builds, slow application start-up, and pulls over about five minutes are not covered — those need the explicit timeout, or an .image/.build unit that does the work in a separate oneshot service the container depends on.
Health-gated readiness
Notify=healthy is the pragmatic option for third-party images: no code changes, and systemd's ordering becomes genuinely meaningful.
[Container]
Image=docker.io/library/postgres:17-alpine
Notify=healthy
HealthCmd=pg_isready -U postgres -d postgres
HealthInterval=10s
HealthStartPeriod=60s
HealthRetries=5
[Service]
TimeoutStartSec=180
Now a dependent unit with After=database.container starts only once Postgres accepts connections. It requires a healthcheck — Notify=healthy without one hangs until the start timeout.
Socket activation
Because Podman forwards LISTEN_FDS into the container, a socket-activated service works the same as any other:
# ~/.config/systemd/user/echo.socket (a plain systemd unit, not a Quadlet)
[Unit]
Description=Socket for the echo container
[Socket]
ListenStream=8080
FileDescriptorName=echo
[Install]
WantedBy=sockets.target
# ~/.config/containers/systemd/echo.container
[Unit]
Requires=echo.socket
After=echo.socket
[Container]
Image=localhost/echo:latest
Network=host
The container process must accept the inherited descriptor (fd 3) rather than binding its own. Systems that cannot do that should publish ports normally.
Two things to know:
- There is no Quadlet key for this, and none is needed. Podman decides it is socket-activated purely by finding
LISTEN_PID/LISTEN_FDSin its own environment, which systemd sets when it hands over the socket; it then passes the descriptors through to the container. Configuration does not enter into it, so there is nothing to switch on. - It does not survive the remote client.
podman-remote(and therefore Podman on Mac and Windows) cannot forward descriptors to the service, so theLISTEN_*variables never reach the container. Socket-activated containers are a local-Podman arrangement.
Health Checks
Defining a check
podman run -d --name api \
--health-cmd 'curl -fsS http://localhost:9000/healthz' \
--health-interval 30s \
--health-timeout 5s \
--health-retries 3 \
--health-start-period 20s \
--health-on-failure kill \
localhost/api:latest
| Flag | Default | Meaning |
|---|---|---|
--health-cmd |
from image | Command run inside the container; a bare string becomes CMD-SHELL |
--health-interval |
30s |
Gap between checks; disable stops the timer entirely |
--health-timeout |
30s |
Per-check deadline |
--health-retries |
3 |
Consecutive failures before unhealthy |
--health-start-period |
0s |
Grace window; failures inside it do not count |
--health-on-failure |
none |
none, kill, restart, stop |
--health-log-destination |
local |
local, a directory, or events_logger |
--health-max-log-count |
5 |
Entries kept |
--health-max-log-size |
500 |
Characters of output kept per entry |
How they actually run
Podman does not poll from a daemon — there isn't one. On container start it creates a transient systemd timer per check:
systemd-run --user --unit CONTAINER_ID-RANDOM \
--on-unit-inactive=30s --timer-property=AccuracySec=1s \
--property=StartLimitIntervalSec=0 \
/usr/bin/podman healthcheck run --ignore-result CONTAINER_ID
Three consequences:
- No systemd, no periodic checks. In a container-in-container build, a minimal chroot, or anything without a session bus, checks only run when you invoke them by hand.
--sdnotify=healthywill hang forever in that environment. --health-interval=disableskips timer creation; the check stays defined for manual runs.DISABLE_HC_SYSTEMD=truein the environment does the same globally.
# Find the timers for a running container
systemctl --user list-timers --all | grep "$(podman inspect -f '{{.Id}}' api | cut -c1-12)"
Reading state
# Run the check now; exit 0 healthy, 1 unhealthy, 125 error
podman healthcheck run api
# Current status and consecutive failure count
podman inspect api --format '{{.State.Health.Status}} {{.State.Health.FailingStreak}}'
# starting | healthy | unhealthy
# The rolling log of attempts
podman inspect api --format '{{json .State.Health.Log}}' | jq '.[-1]'
# {"Start":"...","End":"...","ExitCode":1,"Output":"connection refused"}
# Status column carries it too
podman ps --format '{{.Names}}\t{{.Status}}'
# api Up 3 minutes (healthy)
podman ps --filter health=unhealthy
podman events --filter event=health_status
The start period suppresses failures, not the state machine: failures inside the window leave the status at starting and FailingStreak at 0 no matter how many there are, but the first success promotes the container to healthy immediately, however much of the window is left.
Startup checks
A separate, more aggressive check that gates the regular one. Use it for slow-booting services instead of an enormous --health-start-period, because it can restart a container that is genuinely wedged rather than waiting out the window:
podman run -d --name db \
--health-cmd 'pg_isready -U postgres' \
--health-interval 30s \
--health-startup-cmd 'pg_isready -U postgres' \
--health-startup-interval 2s \
--health-startup-retries 60 \
--health-startup-success 1 \
--health-startup-timeout 5s \
docker.io/library/postgres:17-alpine
The startup check runs every 2s; after one success the regular 30s check takes over and the startup check stops. After 60 failures the container is restarted. A startup check requires a regular check to exist.
--health-on-failure under systemd
This is where healthchecks and systemd overlap, and where the wrong choice gives you two restart loops fighting each other.
| Action | Behaviour | Under systemd |
|---|---|---|
none |
Record only | Nothing restarts; useful for observation only |
kill |
Kill the container | Preferred. systemd's Restart= policy handles recovery with proper backoff |
restart |
Podman restarts it | Conflicts with Restart=; do not combine |
stop |
Stop the container | Also fine — systemd sees a clean exit; pair with Restart=always |
[Container]
Image=localhost/api:latest
HealthCmd=curl -fsS http://localhost:9000/healthz
HealthInterval=30s
HealthRetries=3
HealthStartPeriod=20s
HealthOnFailure=kill
[Service]
Restart=always
RestartSec=10
With kill plus Restart=always, an unhealthy container dies, systemd restarts the unit after 10s, and StartLimitBurst/StartLimitIntervalSec cap the flapping. That is one restart mechanism, with backoff and a journal trail — as opposed to Podman quietly restarting the container behind systemd's back.
Quadlet health keys
Every --health-* flag has a key, with the same defaults:
[Container]
HealthCmd=curl -fsS http://localhost:9000/healthz
HealthInterval=30s
HealthTimeout=5s
HealthRetries=3
HealthStartPeriod=20s
HealthOnFailure=kill
HealthLogDestination=/var/log/healthchecks
HealthMaxLogCount=20
HealthMaxLogSize=2000
HealthStartupCmd=curl -fsS http://localhost:9000/healthz
HealthStartupInterval=2s
HealthStartupRetries=60
HealthStartupSuccess=1
HealthStartupTimeout=5s
Timers and Scheduled Containers
The shape of a scheduled job
A container that runs to completion is Type=oneshot; a .timer unit triggers it. The timer is a plain systemd unit, not a Quadlet, so it goes in the systemd directory — dropping a .timer in ~/.config/containers/systemd/ silently does nothing, because the generator ignores extensions it does not own.
# ~/.config/containers/systemd/backup.container
[Unit]
Description=Nightly database backup
[Container]
Image=docker.io/library/postgres:17-alpine
Exec=/usr/local/bin/backup.sh
Volume=backup-data.volume:/backups:z
Environment=PGHOST=db
Network=app.network
Secret=pgpassword,type=env,target=PGPASSWORD
[Service]
Type=oneshot
# ~/.config/systemd/user/backup.timer <-- NOT the quadlet directory
[Unit]
Description=Run the nightly backup
[Timer]
OnCalendar=*-*-* 02:30:00
RandomizedDelaySec=900
Persistent=true
[Install]
WantedBy=timers.target
systemctl --user daemon-reload
systemctl --user enable --now backup.timer # the timer IS enable-able
systemctl --user list-timers backup.timer
systemctl --user start backup.service # run it now
journalctl --user -u backup.service -n 50
Note the asymmetry: backup.service is generated and cannot be enabled, but backup.timer is a real file and enables normally. The timer pulls in the service, so that is all you need.
Type=oneshot and RemainAfterExit
The Quadlet docs recommend RemainAfterExit=yes with Type=oneshot so the unit does not sit in inactive (dead). Do not do that for timer-triggered units — a unit that remains "started" cannot be re-triggered, and the timer silently stops firing. Use RemainAfterExit=yes for one-shot units that represent a state (a created volume, a pulled image), not for jobs that repeat.
Also worth knowing: Type=oneshot disables the start timeout entirely (it is set to infinity), so a hung job hangs forever unless you set TimeoutStartSec= yourself.
[Service]
Type=oneshot
TimeoutStartSec=1800
OnCalendar in practice
OnCalendar=hourly # *-*-* *:00:00
OnCalendar=daily # *-*-* 00:00:00
OnCalendar=weekly # Mon *-*-* 00:00:00
OnCalendar=*-*-* 02:30:00 # 02:30 every day
OnCalendar=Mon..Fri *-*-* 09:00:00 # weekdays at 09:00
OnCalendar=*-*-01 04:00:00 # first of the month
OnCalendar=*:0/15 # every 15 minutes
# Always check the expression before trusting it
systemd-analyze calendar 'Mon..Fri *-*-* 09:00:00'
systemd-analyze calendar --iterations=5 '*:0/15'
Persistent=true runs a missed job once at next boot — essential for laptops and anything that is not always on. RandomizedDelaySec= spreads load when many hosts share a schedule. OnBootSec=/OnUnitActiveSec= give relative rather than wall-clock schedules.
Templated jobs
One unit, many instances — useful for per-tenant or per-dataset jobs:
# ~/.config/containers/systemd/sync@.container
[Unit]
Description=Sync dataset %i
[Container]
Image=localhost/sync:latest
Exec=/usr/local/bin/sync.sh %i
Volume=sync-data@.volume:/data
[Service]
Type=oneshot
# ~/.config/containers/systemd/sync-data@.volume
[Volume]
VolumeName=sync-%i
systemctl --user start sync@customer-a
systemctl --user start sync@customer-b
# Pin an instance at boot with a symlink, or with DefaultInstance= in [Install]
ln -s sync@.container ~/.config/containers/systemd/sync@customer-a.container
Templates referencing templates use the template name (sync-data@.volume), not the instance name; Quadlet substitutes %i when it expands the reference.
Units Podman ships
# Auto-update: daily, ±15 minutes
systemctl --user enable --now podman-auto-update.timer
systemctl --user list-timers podman-auto-update.timer
# Restart containers with restart-policy=always after a reboot.
# Only needed for containers NOT managed by Quadlet — Quadlet units use [Install].
sudo systemctl enable --now podman-restart.service
# Clean up leftover state from an unclean boot (transient storage hosts)
sudo systemctl enable --now podman-clean-transient.service
Override the auto-update schedule without editing the packaged unit:
systemctl --user edit podman-auto-update.timer
[Timer]
OnCalendar=
OnCalendar=Sun *-*-* 03:00:00
RandomizedDelaySec=1800
The empty OnCalendar= is required: timer directives accumulate, so without the reset you get both schedules.
Auto-Updates
[Container]
Image=quay.io/example/api:stable
AutoUpdate=registry
Notify=healthy
HealthCmd=curl -fsS http://localhost:9000/healthz
HealthStartPeriod=20s
[Service]
Restart=always
| Policy | Compares against | Requires |
|---|---|---|
registry |
The remote registry digest | A fully-qualified image name — not an image ID, not a short name |
local |
The local image store digest | Something else updating local storage (a .build unit, a CI push) |
podman auto-update --dry-run # UPDATED shows "pending"
podman auto-update
podman auto-update --format '{{.Unit}}\t{{.Image}}\t{{.Updated}}'
Rollback is on by default: if the unit fails to restart on the new image, Podman reverts to the previous one and restarts again. It only works if systemd can tell the restart failed — which needs a readiness signal. With the default --sdnotify=conmon, the restart "succeeds" the moment conmon starts, and a container that crashes two seconds later is never rolled back. Pair AutoUpdate=registry with Notify=healthy or Notify=true.
podman auto-update --rollback=false # opt out
For .kube units, annotate in the YAML instead:
metadata:
annotations:
io.containers.autoupdate: "registry"
io.containers.autoupdate/api: "local" # per-container override
io.containers.sdnotify/api: "container"
Rootless in Production
# Without lingering, user units stop at logout and never start at boot
loginctl enable-linger mike
loginctl show-user mike --property=Linger
# For a service account with /sbin/nologin
sudo loginctl enable-linger svc-app
sudo systemctl --machine=svc-app@ --user daemon-reload
sudo systemctl --machine=svc-app@ --user start app-pod
# Subordinate ID ranges must exist before any rootless container runs
grep "^${USER}:" /etc/subuid /etc/subgid
sudo usermod --add-subuids 100000-165535 --add-subgids 100000-165535 "$USER"
podman system migrate # after changing ranges
# Privileged ports, host-wide
echo 'net.ipv4.ip_unprivileged_port_start=80' | sudo tee /etc/sysctl.d/99-podman.conf
sudo sysctl --system
# Journal for a user unit
journalctl --user -u web.service -f
sudo journalctl _UID=1000 -u web.service # from a root shell
podman system migrate is also the fix after changing subuid/subgid ranges — existing containers keep the old mapping otherwise, and volume ownership goes wrong in confusing ways.
Secrets
printf '%s' "$DB_PASSWORD" | podman secret create db-password -
podman secret ls
podman secret inspect db-password # metadata; add --showsecret for the value
[Container]
Secret=db-password,type=env,target=POSTGRES_PASSWORD
Secret=tls-cert,type=mount,target=/etc/ssl/private/tls.crt,mode=0400,uid=1000
Podman's default file driver stores secrets unencrypted, as base64 in secrets/filedriver/secretsdata.json under the storage root — ~/.local/share/containers/storage/ rootless, /var/lib/containers/storage/ rootful — protected only by file permissions, and readable with podman secret inspect --showsecret. That is acceptable on a single-tenant host and not acceptable on a shared one: use the pass or shell driver, or fetch from Vault at start-up, when the threat model includes other users on the box.
Debugging
# 1. What did Quadlet generate, and did it error?
/usr/lib/systemd/system-generators/podman-system-generator --user --dryrun
# 2. Narrow it to the units you are working on
mkdir -p /tmp/q && cp ~/.config/containers/systemd/web.container /tmp/q/
QUADLET_UNIT_DIRS=/tmp/q \
/usr/lib/systemd/system-generators/podman-system-generator --user --dryrun
# 3. Errors only, plus systemd's own validation
systemd-analyze --user --generators=true verify web.service
# 4. The generated file on disk, after a reload
systemctl --user cat web.service
ls /run/user/$(id -u)/systemd/generator/
# 5. Runtime
systemctl --user status web.service
journalctl --user -u web.service -n 100 --no-pager
journalctl --user -u web.service -f
podman logs -f systemd-web
# 6. Dependency graph
systemctl --user list-dependencies app-pod.service
systemd-analyze --user critical-chain web.service
A single unsupported key aborts generation for that file, and the symptom is a missing unit rather than an error:
$ systemctl --user start web
Failed to start web.service: Unit web.service not found.
$ /usr/lib/systemd/system-generators/podman-system-generator --user --dryrun
converting "web.container": unsupported key 'PublishPorts' in group 'Container'
That is nearly always a typo or a key from a newer Podman than the host runs — check the version map above before assuming the documentation is wrong.
Migrating from podman generate systemd
# What you have today
systemctl --user list-unit-files 'container-*' 'pod-*'
systemctl --user cat container-web.service | grep ExecStart
Translate the podman run line into keys, then decommission the old unit:
systemctl --user disable --now container-web.service
rm ~/.config/systemd/user/container-web.service
# write ~/.config/containers/systemd/web.container
systemctl --user daemon-reload
systemctl --user start web
podman generate systemd |
Quadlet |
|---|---|
container-foo.service |
foo.container → foo.service |
pod-foo.service + one unit per container |
foo.pod + one .container each |
--new (recreate on start) |
Always the case; --replace --rm is generated |
--restart-policy |
[Service] Restart= |
--time |
[Container] StopTimeout= |
Unit is a real file, systemctl enable works |
Transient; [Install] WantedBy= instead |
Hand-maintained After=/Requires= |
Generated from Pod=, Network=, Volume= |
Run both side by side during a migration if you like, but not for the same container — --replace means whichever starts last wins.
The podman quadlet CLI (5.6+)
Managing Quadlet files as installable units rather than hand-copied files:
podman quadlet list
podman quadlet list --filter status=active/running
podman quadlet list --format '{{.Name}}\t{{.UnitName}}\t{{.Status}}\t{{.Pod}}'
podman quadlet print web.container
podman quadlet install ./web.container
podman quadlet install --application myapp ./myapp/ # a directory of files
podman quadlet install --replace ./web.container # update in place
podman quadlet rm web.container
Files install under ~/.config/containers/systemd/; an "application" (a directory, or a .quadlets file with --- separators) installs into a subdirectory and is removed as a unit. Not yet supported by the remote client.
Quick Reference
| Task | Command |
|---|---|
| Reload after editing a Quadlet | systemctl --user daemon-reload |
| Start a container unit | systemctl --user start web |
| Start a pod unit | systemctl --user start app-pod |
| Preview generated units | /usr/lib/systemd/system-generators/podman-system-generator --user --dryrun |
| Validate one unit | systemd-analyze --user --generators=true verify web.service |
| Show generated unit | systemctl --user cat web.service |
| Follow logs | journalctl --user -u web.service -f |
| Run a healthcheck now | podman healthcheck run web |
| Health status | podman inspect web --format '{{.State.Health.Status}}' |
| List health timers | systemctl --user list-timers --all |
| Check auto-updates | podman auto-update --dry-run |
| Apply auto-updates | podman auto-update |
| Survive logout and boot | loginctl enable-linger $USER |
| Test a calendar expression | systemd-analyze calendar '*-*-* 02:30:00' |
| List installed Quadlets | podman quadlet list |
Unit file skeletons
# container
[Unit]
Description=
[Container]
Image=
[Service]
Restart=always
[Install]
WantedBy=default.target
# pod
[Pod]
PodName=
PublishPort=
[Service]
Restart=always
[Install]
WantedBy=default.target
# scheduled job
[Container]
Image=
Exec=
[Service]
Type=oneshot
TimeoutStartSec=1800
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
Unit web.service not found after daemon-reload |
Generation failed on an unsupported or misspelled key | Run the generator with --dryrun and read the error |
systemctl enable fails on a generated unit |
Generated units are transient | Use [Install] WantedBy=default.target in the Quadlet file |
systemctl start foo fails for a .pod file |
Pods generate foo-pod.service |
systemctl start foo-pod |
Unit stuck in activating (start), then killed |
Notify=true but the app never sends READY=1 |
Drop to the default conmon, or use Notify=healthy, or implement sd_notify |
Notify=healthy never becomes ready |
No healthcheck defined, or no systemd to run the timer | Add HealthCmd=; confirm systemctl is-system-running responds |
| Healthcheck never runs on its own | Timers need systemd; --health-interval=disable; DISABLE_HC_SYSTEMD=true |
Check systemctl --user list-timers --all; run podman healthcheck run to confirm the check itself works |
| Container restarts in a tight loop | HealthOnFailure=restart fighting Restart=always |
Use HealthOnFailure=kill and let systemd own recovery |
| Timer fires once and never again | RemainAfterExit=yes on a Type=oneshot job |
Remove it; keep it only for state-representing units |
.timer in the Quadlet directory does nothing |
The generator ignores non-Quadlet extensions | Put timers in ~/.config/systemd/user/ or /etc/systemd/system/ |
| Network changes have no effect | podman network create --ignore no-ops on an existing network |
Remove the network and restart, or set NetworkDeleteOnStop=true |
| Volume data vanished after an edit | Not vanished — --replace --rm recreated the container |
Anything that must persist belongs in a .volume or a bind mount |
| Unit fails with a pull timeout | Podman's automatic extension caps out at 5 minutes beyond TimeoutStartSec |
Raise TimeoutStartSec=, or pre-pull with an .image unit |
| User units die at logout | Lingering not enabled | loginctl enable-linger $USER |
| Rootless volume files owned by nobody | subuid/subgid ranges changed after creation | podman system migrate, then fix ownership with podman unshare chown |
| Auto-update never rolls back a broken image | Restart "succeeds" with --sdnotify=conmon |
Set Notify=healthy (or true) so failure is detectable |
EnvironmentFile=%n/env not found |
Quadlet leaves %-leading relative paths unresolved |
Write EnvironmentFile=./%n/env |
| Two schedules after overriding a timer | Timer directives accumulate | Reset with an empty OnCalendar= before the new value |
Related Topics
The following topics complement this cheatsheet and would be valuable additions:
- Podman - The base engine: rootless containers, images, volumes, networks, and the Docker-compatible CLI these units drive
- systemd - Unit syntax, targets, timers, drop-ins, and
journalctlin depth; everything Quadlet passes through - Container Security - Rootless isolation, user namespaces, SELinux labels, capability dropping, and secret handling
- Kubernetes - The pod semantics Podman mirrors, and the YAML that
.kubeunits consume - SELinux/AppArmor - Volume relabelling (
:z/:Z), container contexts, and the denials that break bind mounts - Linux CLI -
loginctl,systemd-analyze,journalctl, and the surrounding toolchain