systemd Service Hardening
Sandboxing a systemd service: capabilities, namespaces, seccomp, credentials, and scoring the result with systemd-analyze security.
systemd Service Hardening
Turning a service that runs as root with the whole system in reach into one that can only touch what it needs.
Overview
systemd can apply most of what you would otherwise reach for a container to get: a private mount namespace, a read-only filesystem, dropped capabilities, a seccomp filter, an IP allow-list, and a transient unprivileged user. It is all unit-file configuration, it applies to the process as it is executed, and there is no image to build.
The catch is that it is a long list of directives with interacting effects and one shared failure mode — a service that starts fine and then cannot reach something it needs. This sheet is the order to apply them in, and the tool that tells you whether it worked.
flowchart TD
A[Unmodified service<br/>root, full system access] --> B[Score it<br/>systemd-analyze security]
B --> C[Identity: DynamicUser + StateDirectory]
C --> D[Privileges: NoNewPrivileges + CapabilityBoundingSet]
D --> E[Filesystem: ProtectSystem + ReadWritePaths]
E --> F[Namespaces: Private/Protect directives]
F --> G[Kernel: seccomp + SystemCallArchitectures]
G --> H[Network: RestrictAddressFamilies + IPAddress filters]
H --> I{Still working?}
I -->|No| J[Read the journal<br/>relax the last directive]
J --> I
I -->|Yes| K[Re-score and commit]
Work in that order and each step's failures are attributable. Paste the whole list in at once and you get a service that will not start, with no clue which of twenty directives did it.
Scoring a Unit
systemd-analyze security is the feedback loop. It weighs roughly seventy checks into a single
exposure number, and it tells you which directive each point of badness came from.
# Score a loaded unit
systemd-analyze security nginx.service
# Score a file, no running manager needed - works in CI and in a container
systemd-analyze security --offline=true ./myapp.service
# Only show what is failing, above a threshold
systemd-analyze security --threshold=50 myapp.service
# Machine-readable, for a CI gate
systemd-analyze security --offline=true --json=short ./myapp.service
# Score everything on the box, worst first
systemd-analyze security
Output is one row per check — a tick or cross, the directive, what it means, and the badness that check contributed — followed by the overall exposure:
✗ PrivateDevices= Service potentially has access to hardware devices 0.2
✓ NoNewPrivileges= Service processes cannot acquire new privileges
✗ RestrictAddressFamilies= Service may allocate exotic sockets 0.3
→ Overall exposure level for myapp.service: 9.4 UNSAFE :-{
| Exposure | Band |
|---|---|
| 10.0 | DANGEROUS |
| 9.0 – 9.9 | UNSAFE |
| 7.5 – 8.9 | EXPOSED |
| 5.0 – 7.4 | MEDIUM |
| 1.0 – 4.9 | OK |
| 0.1 – 0.9 | SAFE |
| 0.0 | PERFECT |
Where none of this applies. The namespacing directives —
ProtectSystem=,ProtectHome=,PrivateTmp=,PrivateDevices=and the rest — need filesystem namespacing, which is unavailable to the per-user service manager and inside a container manager that withholds it. They are accepted and silently do nothing there. Most of them do work in a user service when combined withPrivateUsers=yes. LikewiseRestrictRealtime=andSystemCallFilter=need seccomp, which some container runtimes disable.systemd-analyze securityscores the unit file, not what the kernel actually enforced — a passing score in a container is not proof of a sandbox.
The score is a checklist, not a verdict. A unit at
MEDIUMthat correctly cannot write outside its state directory is in better shape than one atSAFEthat got there withPrivateNetwork=yeson a service that never needed the network anyway. Chase the specific crosses that matter for your threat model, not the number.
Gating in CI
#!/usr/bin/env bash
# Fail the build if any shipped unit scores worse than MEDIUM.
set -euo pipefail
threshold=50 # Internal 0-100 scale; 50 is the MEDIUM boundary
for unit in units/*.service; do
if ! systemd-analyze security --offline=true --threshold="$threshold" "$unit"; then
echo "FAIL: $unit exceeds exposure threshold" >&2
exit 1
fi
done
--threshold= makes the command exit non-zero when the unit scores worse, which is the whole CI
gate. The value is on the internal 0–100 scale, not the 0–10 one printed in the summary line.
Step 1 — Identity
Everything else is easier once the service is not root.
[Service]
DynamicUser=yes
StateDirectory=myapp
CacheDirectory=myapp
LogsDirectory=myapp
RuntimeDirectory=myapp
DynamicUser=yes allocates a UID for the lifetime of the service. There is no account to create
in a postinstall, no UID to collide on across a fleet, and nothing left behind. It implies a set of
useful defaults on its own: PrivateTmp=yes, RemoveIPC=yes, ProtectSystem=strict,
ProtectHome=read-only, and a restricted NoNewPrivileges=yes.
The *Directory= directives are what make it workable. Each one creates the directory with the
right ownership, and — critically — adds it to the writable set so ProtectSystem=strict does not
lock the service out of its own data.
| Directive | Path | Lifetime |
|---|---|---|
StateDirectory= |
/var/lib/myapp |
Persists across restarts and reboots |
CacheDirectory= |
/var/cache/myapp |
Persists; safe to delete |
LogsDirectory= |
/var/log/myapp |
Persists |
RuntimeDirectory= |
/run/myapp |
Removed when the service stops |
ConfigurationDirectory= |
/etc/myapp |
Persists across package upgrades |
All five create the directory mode 0755, chown it to the service's user, and imply
BindPaths= for it — which is what excludes them from ProtectSystem=strict. Under
DynamicUser=, the real directories live under /var/lib/private/ (inaccessible to unprivileged
users, so a recycled UID cannot reach an old service's data) with symlinks so both the host and
the service still see them at /var/lib/myapp.
DynamicUser=yesdoes not fit every service. A UID that changes between boots breaks anything that stores file ownership outside the*Directory=set, that needs a stable UID in a shared NFS mount, or that other serviceschownfiles to. For those, create a real system account (useradd --system --no-create-home --shell /usr/sbin/nologin myapp) and useUser=.
Step 2 — Privileges and Capabilities
[Service]
NoNewPrivileges=yes
CapabilityBoundingSet=
AmbientCapabilities=
NoNewPrivileges=yes means no execve() in this service or its children can gain privileges — a
setuid binary in the sandbox becomes just a binary. It is free, it almost never breaks anything,
and it is a precondition for seccomp filtering to be applied without privilege.
An empty CapabilityBoundingSet= drops every capability, permanently, for the unit and everything
it spawns. Grant back only what the service genuinely needs:
[Service]
# A service that binds port 443 but is otherwise unprivileged
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
AmbientCapabilities=CAP_NET_BIND_SERVICE
| Capability | Grants |
|---|---|
CAP_NET_BIND_SERVICE |
Bind ports below 1024 |
CAP_NET_ADMIN |
Configure interfaces, routes, firewall — very broad, avoid |
CAP_NET_RAW |
Raw and packet sockets (ping, tcpdump) |
CAP_CHOWN, CAP_FOWNER, CAP_DAC_OVERRIDE |
Bypass file ownership and permission checks |
CAP_SYS_ADMIN |
Effectively root. If you need this, the sandbox is not the answer |
CapabilityBoundingSet=is a ceiling, not a grant: it says what the service may ever hold.AmbientCapabilities=is what it actually starts with. Setting only the bounding set on a service running as a non-rootUser=gives it nothing — you need both. And because bounding-set directives are lists, a drop-in appends: writeCapabilityBoundingSet=on its own line first to reset before narrowing.
Prefer socket activation to CAP_NET_BIND_SERVICE where you can. If PID 1 opens port 443 and
passes the descriptor, the service needs no networking capability at all.
Step 3 — Filesystem
[Service]
ProtectSystem=strict
ProtectHome=yes
ReadWritePaths=/var/lib/myapp
ReadOnlyPaths=/etc/myapp
InaccessiblePaths=/srv/other-tenant
PrivateTmp=yes
Value of ProtectSystem= |
Effect |
|---|---|
no |
Default. No protection |
yes |
/usr and /boot read-only |
full |
Adds /etc read-only |
strict |
The entire hierarchy read-only, except /dev, /proc, /sys |
strict is the one to aim for. Anything the service must write to goes in ReadWritePaths=, and
if you used the *Directory= directives in step 1, its own state is already there.
[Service]
# Hide everything, then bind back exactly what is needed
TemporaryFileSystem=/var:ro
BindReadOnlyPaths=/var/lib/myapp/config
BindPaths=/var/lib/myapp/data
TemporaryFileSystem= mounts an empty tmpfs over a path, hiding whatever was there; BindPaths=
and BindReadOnlyPaths= then punch specific things back through. This is the sharpest tool in the
set and the easiest to get wrong — reach for ProtectSystem=strict plus ReadWritePaths= first.
[Service]
# Nothing under these paths may be executed - blocks dropped-payload execution
NoExecPaths=/
ExecPaths=/usr/bin /usr/lib /usr/lib64
UMask=0077
ProtectSystem=strictwith noReadWritePaths=is the single most common way to break a service when hardening it. The symptom is a write failure in the application's own error handling, not a systemd message — check the service's log, not just the unit's status.
Step 4 — Namespaces and the Kernel Interface
[Service]
PrivateDevices=yes
PrivateUsers=yes
PrivateIPC=yes
ProtectProc=invisible
ProcSubset=pid
ProtectClock=yes
ProtectHostname=yes
ProtectKernelLogs=yes
ProtectKernelModules=yes
ProtectKernelTunables=yes
ProtectControlGroups=yes
RestrictNamespaces=yes
RestrictRealtime=yes
RestrictSUIDSGID=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
RemoveIPC=yes
KeyringMode=private
| Directive | Blocks |
|---|---|
PrivateDevices=yes |
Physical devices; the service sees a minimal /dev |
PrivateUsers=yes |
UIDs outside the service's own; root in the namespace is nobody outside it |
PrivateIPC=yes |
248+. Sharing System V IPC and POSIX message queues with other services |
ProtectProc=invisible |
247+. Seeing other users' processes in /proc |
ProcSubset=pid |
247+. Everything in /proc that is not process introspection |
ProtectClock=yes |
Setting the system clock |
ProtectHostname=yes |
Changing the hostname |
ProtectKernelLogs=yes |
Reading the kernel ring buffer (dmesg) |
ProtectKernelModules=yes |
Loading and unloading modules |
ProtectKernelTunables=yes |
Writing to /proc/sys, /sys, /proc/sysrq-trigger |
ProtectControlGroups=yes |
Writing to /sys/fs/cgroup — cgroup escape |
RestrictNamespaces=yes |
Creating namespaces, i.e. building a sandbox to escape into |
RestrictRealtime=yes |
Realtime scheduling, which can be used to lock up a core |
RestrictSUIDSGID=yes |
Creating setuid/setgid files |
LockPersonality=yes |
Switching execution domain to reach a buggier syscall ABI |
MemoryDenyWriteExecute=yes |
Mappings that are writable and executable — most shellcode |
RemoveIPC=yes |
Leaving IPC objects behind after the service stops |
KeyringMode=private |
Sharing the kernel keyring with other services |
MemoryDenyWriteExecute=yesbreaks any runtime with a JIT — JVM, V8/Node, .NET, LuaJIT, PyPy, and anything embedding them. It is one of the highest-value directives for a compiled service and a non-starter for an interpreted one. Test, do not assume.
PrivateUsers=yesbreaks services that need to see real UIDs outside their own — anything reading/etc/passwdfor other users, or writing files another service owns.
Step 5 — System Calls
[Service]
SystemCallArchitectures=native
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources @obsolete
SystemCallErrorNumber=EPERM
SystemCallArchitectures=native alone is worth having: it blocks the 32-bit compat ABI on a
64-bit host, which is a recurring source of kernel bugs and a standard filter-bypass route.
Filters are built from named sets rather than individual syscalls. @system-service is the
curated "what a normal daemon needs" allow-list and the right starting point. A ~ prefix
subtracts.
# What is actually in a set
systemd-analyze syscall-filter @system-service
systemd-analyze syscall-filter @privileged
# Every set the local systemd knows about
systemd-analyze syscall-filter | grep '^@'
| Set | Contains |
|---|---|
@system-service |
The general-purpose allow-list for services |
@privileged |
Syscalls needing elevated privilege |
@resources |
Changing resource limits, scheduling, affinity |
@mount |
Mounting and unmounting |
@module |
Kernel module loading |
@reboot |
Rebooting and kexec |
@raw-io |
ioperm, iopl, raw port access |
@debug |
ptrace, and the rest of the debugging surface |
@obsolete |
Long-deprecated calls no current program uses |
@swap |
Enabling and disabling swap |
SystemCallErrorNumber=EPERMmakes a blocked call return an error instead of killing the process withSIGSYS. Useful while working out what a service actually needs — a library that probes for a syscall and falls back gracefully will surviveEPERMand die onSIGSYS. Switch back to the default kill once you have the filter right; a silently-failing syscall is worse in production than a loud one.
Step 6 — Network
[Service]
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
IPAddressDeny=any
IPAddressAllow=localhost 10.0.0.0/8
RestrictNetworkInterfaces=eth0
SocketBindDeny=any
SocketBindAllow=tcp:8080
| Directive | Effect |
|---|---|
RestrictAddressFamilies= |
Which socket families may be created. Dropping AF_NETLINK and AF_PACKET costs most services nothing |
IPAddressDeny= / IPAddressAllow= |
eBPF-based egress and ingress filtering, per unit |
RestrictNetworkInterfaces= |
250+. Which interfaces the service may use |
SocketBindDeny= / SocketBindAllow= |
249+. Which ports the service may bind |
PrivateNetwork=yes |
A private network namespace with only loopback. Total isolation |
IPAddressAllow= accepts localhost, link-local, multicast, and CIDR ranges. The two
directives compose: deny everything, then allow the specific ranges the service talks to.
PrivateNetwork=yesis the strongest option and the most disruptive — it takes away DNS as well, so anything resolving a hostname breaks. It works well for a service reached only through socket activation, since PID 1 holds the socket in the host namespace and passes the descriptor in.
IPAddressDeny=needs cgroup v2 and kernel eBPF support. On a host without it the directive is accepted and silently does nothing — confirm withsystemd-analyze security, which reports the check as failing rather than passing.
Credentials Instead of Environment Secrets
Secrets in Environment= are readable by anyone who can run systemctl show, and they sit in
/proc/<pid>/environ for the life of the process. The credentials mechanism keeps them out of
both.
[Service]
# Plaintext file, mode 0600, read by PID 1 and passed in
LoadCredential=dbpass:/etc/myapp/dbpass
# Encrypted at rest, decrypted by PID 1 at start (250+)
LoadCredentialEncrypted=apikey:/etc/credstore.encrypted/myapp.apikey
# Literal value, still kept out of the environment
SetCredential=region:eu-west-2
# Pull everything matching a glob from the credential store (254+)
ImportCredential=myapp.*
The service reads them as files from $CREDENTIALS_DIRECTORY, a 0400 tmpfs visible only to this
unit and gone when it stops:
import os
from pathlib import Path
def read_credential(name: str) -> str:
"""Read a systemd credential by name.
Raises RuntimeError when not running under systemd with credentials
configured, and FileNotFoundError when the named credential was not
passed — both are configuration errors worth failing loudly on rather
than falling back to an empty secret.
"""
directory = os.environ.get("CREDENTIALS_DIRECTORY")
if not directory:
raise RuntimeError("CREDENTIALS_DIRECTORY unset: not started by systemd?")
path = Path(directory) / name
# strip() because an editor-written secret usually carries a trailing newline
return path.read_text(encoding="utf-8").strip()
db_password = read_credential("dbpass")
# Encrypt a secret at rest, bound to this host's TPM2 or /var/lib/systemd/credential.secret
sudo systemd-creds encrypt --name=apikey plaintext.txt \
/etc/credstore.encrypted/myapp.apikey
sudo systemd-creds decrypt /etc/credstore.encrypted/myapp.apikey - # Verify
systemd-creds has-tpm2 # Is TPM2 sealing available?
# What a running service can see
sudo systemd-run --pipe --wait \
-p LoadCredential=test:/etc/myapp/dbpass \
--property=Environment=X=1 \
/bin/sh -c 'ls -l "$CREDENTIALS_DIRECTORY"'
systemd-creds encryptbinds the ciphertext to the host, and that is not optional in the way you might hope.--with-key=takeshost,tpm2,host+tpm2,tpm2-absent,autoorauto-initrd;autopicks TPM2 where available and falls back to the host key.hostmeans/var/lib/systemd/credential.secreton this machine — it is the host-bound option, not an escape from host binding, so ahost-encrypted credential cannot be baked into a shared image or copied across a fleet. Nothing here makes one portable by accident.Fleet-wide distribution needs a deliberate policy: encrypt per host at provisioning time, or seal against a TPM2 public key (
--tpm2-public-key=, with--tpm2-public-key-pcrs=and a matching--tpm2-signature=) so any host whose PCRs satisfy the signed policy can decrypt. Either way, plan the re-encryption path before adopting it — host replacement and PCR-changing firmware updates both invalidate existing ciphertext.
Testing Without Committing
systemd-run applies any [Service] directive to a one-off command, so you can bisect a broken
sandbox without editing and reloading a unit each time.
# Does the service survive this directive?
sudo systemd-run --pipe --wait \
-p ProtectSystem=strict -p ReadWritePaths=/var/lib/myapp \
/usr/bin/myapp --self-test
# Add directives one at a time until it breaks
sudo systemd-run --pipe --wait -p MemoryDenyWriteExecute=yes /usr/bin/myapp --version
For a unit already in place, layer the hardening in its own drop-in so it can be removed in one step:
sudo systemctl edit --drop-in=hardening.conf myapp.service
# ... add the [Service] block ...
sudo systemctl restart myapp.service
systemd-analyze security myapp.service
# Back it out entirely if it goes wrong
sudo rm /etc/systemd/system/myapp.service.d/hardening.conf
sudo systemctl daemon-reload && sudo systemctl restart myapp.service
A Worked Example
A plain unit, and the same unit hardened. Scores are from systemd-analyze security --offline=true
on systemd 255.
# Before: 9.4 UNSAFE
[Unit]
Description=My API
[Service]
ExecStart=/usr/local/bin/myapi --port 8080
[Install]
WantedBy=multi-user.target
# After: 0.9 SAFE
[Unit]
Description=My API
After=network.target
[Service]
Type=exec
ExecStart=/usr/local/bin/myapi --port 8080
# Identity
DynamicUser=yes
StateDirectory=myapi
# Privileges
NoNewPrivileges=yes
CapabilityBoundingSet=
AmbientCapabilities=
# Filesystem
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
UMask=0077
# Namespaces and kernel
PrivateDevices=yes
PrivateUsers=yes
ProtectProc=invisible
ProcSubset=pid
ProtectClock=yes
ProtectHostname=yes
ProtectKernelLogs=yes
ProtectKernelModules=yes
ProtectKernelTunables=yes
ProtectControlGroups=yes
RestrictNamespaces=yes
RestrictRealtime=yes
RestrictSUIDSGID=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
RemoveIPC=yes
# System calls
SystemCallArchitectures=native
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources
# Network
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
IPAddressDeny=any
IPAddressAllow=localhost
[Install]
WantedBy=multi-user.target
Intermediate points on that path, for calibration: identity plus privileges plus the basic
filesystem directives lands on 5.6 MEDIUM; adding the namespace block reaches 3.3 OK; the
seccomp and IP filters take it the rest of the way. Expect the exact numbers to drift between
systemd releases as checks are added — the direction of travel is the point, not the digits.
The
afterunit binds port 8080, not 443, which is why it needs no capabilities at all. AddCapabilityBoundingSet=CAP_NET_BIND_SERVICEand the matchingAmbientCapabilities=for a privileged port — or put a.socketunit in front and keep the empty set.
Resource Limits as a Hardening Control
Sandboxing stops a compromised service reaching things. Resource limits stop a misbehaving one taking the host with it.
[Service]
MemoryMax=1G
MemoryHigh=768M
CPUQuota=50%
TasksMax=128
LimitNOFILE=8192
LimitCORE=0
OOMPolicy=stop
TasksMax=is the fork-bomb ceiling. It counts threads as well as processes, so size it against the service's real thread pool.MemoryHigh=belowMemoryMax=gives the kernel a chance to reclaim and throttle before the OOM killer fires, which turns an abrupt kill into observable back-pressure.LimitCORE=0stops core dumps, which otherwise write process memory — including any secret it was holding — to disk.OOMPolicy=stopstops the whole unit when a process in it is OOM-killed, rather than leaving a half-dead service running.
Quick Reference
| Command | Purpose |
|---|---|
systemd-analyze security UNIT |
Score a loaded unit |
systemd-analyze security --offline=true FILE |
Score a unit file with no manager |
systemd-analyze security --threshold=50 UNIT |
Non-zero exit above the threshold, for CI |
systemd-analyze security --json=short UNIT |
Machine-readable findings |
systemd-analyze verify FILE |
Validate syntax and directive names |
systemd-analyze syscall-filter @set |
List the syscalls in a filter set |
systemd-run -p DIRECTIVE=value CMD |
Try a directive without touching a unit file |
systemd-creds encrypt IN OUT |
Encrypt a credential at rest |
systemctl show UNIT -p PropertyName |
The value systemd actually resolved |
The Ten-Minute Sandbox
Directives that are almost always safe, in the order to add them:
[Service]
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
ProtectClock=yes
ProtectHostname=yes
ProtectKernelLogs=yes
ProtectKernelModules=yes
ProtectKernelTunables=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
RestrictRealtime=yes
LockPersonality=yes
SystemCallArchitectures=native
ReadWritePaths=/var/lib/myapp
Everything above this line rarely breaks a well-behaved service. MemoryDenyWriteExecute=,
PrivateUsers=, PrivateNetwork=, SystemCallFilter= and the IP filters need testing.
Common Issues and Solutions
Service Fails to Start After Hardening
systemctl status myapp.service # Look at the exit code first
journalctl -u myapp.service -b --no-pager | tail -40
| Exit code | Means |
|---|---|
203/EXEC |
The binary could not be executed — NoExecPaths=, RootDirectory=, or a path hidden by the sandbox |
226/NAMESPACE |
A namespace or mount directive failed to set up; usually a ReadWritePaths= pointing at a path that does not exist |
228/SECCOMP |
The seccomp filter could not be installed |
31/SYS (SIGSYS) |
A blocked syscall killed the process — relax SystemCallFilter= |
Fix: bisect with systemd-run -p … rather than editing the unit. Add directives back in the
order of the six steps above; the first one that breaks it is the one that needs relaxing, not the
whole block.
Service Starts but Cannot Write
systemctl show myapp.service -p ReadWritePaths -p StateDirectory -p ProtectSystem
Fix: ProtectSystem=strict without a matching ReadWritePaths=. Prefer StateDirectory= over
a hand-written ReadWritePaths=/var/lib/myapp — it creates the directory, sets ownership, and adds
it to the writable set in one directive. Note that ReadWritePaths= requires the path to already
exist, and fails the unit with 226/NAMESPACE if it does not.
Service Cannot Resolve DNS
systemctl show myapp.service -p PrivateNetwork -p IPAddressAllow -p RestrictAddressFamilies
Fix: PrivateNetwork=yes removes all networking including the resolver. If you need DNS,
drop it and use IPAddressAllow= instead. If IPAddressDeny=any is set, systemd-resolved on
127.0.0.53 needs IPAddressAllow=localhost, and a direct external resolver needs its address
allowed too. RestrictAddressFamilies= without AF_UNIX breaks the resolver socket.
JIT-Compiled Runtime Crashes
Fix: MemoryDenyWriteExecute=yes is incompatible with every JIT. Remove it for JVM, Node,
.NET, and PyPy services. systemd-analyze security will keep scoring the check as failed — that
is correct and unavoidable for those runtimes.
Hardening Drop-in Appears to Do Nothing
systemctl cat myapp.service # Is the drop-in in the merged view?
systemd-analyze security myapp.service # Are the checks actually passing?
Fix: the usual causes are a missing daemon-reload, a drop-in file not ending in .conf, or
a trailing comment on a directive line — unit files have no trailing comments, so
ProtectSystem=strict # lock it down parses as the invalid value strict # lock it down and is
silently dropped. Put comments on their own line and re-run systemd-analyze verify.
DynamicUser Breaks File Ownership
Fix: the UID changes between restarts, so anything owning files outside StateDirectory= and
friends will break. Either move the data into a *Directory= (systemd keeps /var/lib/private/
ownership consistent for you) or drop DynamicUser= for a real system account.
Related Topics
- systemd — units, timers, socket activation, drop-ins, and the rest of the service manager
- SELinux / AppArmor — mandatory access control, the layer beneath these directives
- Container Security — the same isolation primitives, seen from the container side
- Linux User and Access Management — the accounts and sudo policy around
User= - Podman Quadlets and systemd — rootless containers as units, which take these same sandboxing directives
- Security Patterns — where service hardening sits in a defence-in-depth design