systemd
Init system and service manager for Linux: units, timers, sockets, drop-in overrides, resource control, and journald.
systemd
Init system and service manager for Linux, managing system processes, services, and boot targets.
Overview
systemd is the default init system for most modern Linux distributions. It manages the complete lifecycle of system services, handles dependencies between units, and provides extensive logging through journald. Units are the fundamental building blocks, with services, timers, mounts, and targets being the most commonly used types.
graph TB
subgraph "systemd Architecture"
A[systemd PID 1] --> B[Service Units]
A --> C[Timer Units]
A --> D[Mount Units]
A --> E[Target Units]
A --> F[Socket Units]
B --> G[journald]
C --> G
D --> G
E --> H[multi-user.target]
E --> I[graphical.target]
H --> J[network.target]
end
Unit File Structure
Unit files define how systemd manages services, timers, mounts, and other resources. They use INI-style syntax with sections denoted by square brackets.
Key Concepts
- Unit files location:
/etc/systemd/system/(admin),/usr/lib/systemd/system/(packages) - Override files:
/etc/systemd/system/myapp.service.d/*.conf(directory carries the full unit name) - User units:
~/.config/systemd/user/
Comments must be on their own line. Unit files are not shell or INI-with-trailing-comments: a
#after a value is part of the value.PrivateTmp=yes # sandboxsilently parses as the boolean"yes # sandbox", fails, and is dropped — you get a service with no sandboxing and only aFailed to parse boolean value, ignoringline buried in the journal. OnExecStart=it is worse: the comment text becomes extraargventries. Put every comment on a line of its own, and runsystemd-analyze verifybefore trusting a unit you hand-edited.
Service Unit Structure
# /etc/systemd/system/myapp.service
[Unit]
Description=My Application Service
Documentation=https://example.com/docs
After=network.target postgresql.service
Requires=postgresql.service
[Service]
Type=simple
User=appuser
Group=appgroup
WorkingDirectory=/opt/myapp
Environment=NODE_ENV=production
EnvironmentFile=/etc/myapp/env
ExecStartPre=/usr/bin/myapp-check
ExecStart=/usr/bin/myapp --config /etc/myapp/config.yaml
ExecReload=/bin/kill -HUP $MAINPID
# No ExecStop= needed: systemd sends SIGTERM to the main process, then SIGKILL
# after TimeoutStopSec=. Only set ExecStop= if the daemon needs a custom drain
# command; `ExecStop=/bin/kill -TERM $MAINPID` is a cargo-culted no-op.
Restart=on-failure
RestartSec=5
TimeoutStartSec=30
TimeoutStopSec=30
[Install]
WantedBy=multi-user.target
Timer Unit Structure
# /etc/systemd/system/backup.timer
[Unit]
Description=Daily Backup Timer
[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true
RandomizedDelaySec=300
[Install]
WantedBy=timers.target
Mount Unit Structure
# /etc/systemd/system/data.mount
[Unit]
Description=Data Volume Mount
[Mount]
What=/dev/sdb1
Where=/data
Type=ext4
Options=defaults,noatime
[Install]
WantedBy=multi-user.target
The filename is not free-form. A mount unit's name must be the escaped form of its
Where=path, or systemd refuses to load it withWhere= setting doesn't match unit name. Compute it rather than guessing:systemd-escape -p --suffix=mount /srv/datagivessrv-data.mount. The same rule governs.automount,.swap, and.deviceunits.
Service Types
| Type | Considered "started" when | Use for |
|---|---|---|
simple |
systemd forks the process (immediately) | Default; foreground processes that need no ordering guarantee |
exec |
The binary has been execve()d successfully |
Better default than simple — catches a missing/unexecutable ExecStart at start time |
forking |
The parent exits after forking the child | Traditional daemons that background themselves |
oneshot |
The process exits | Scripts and setup steps; pair with RemainAfterExit=yes |
notify |
The process sends READY=1 via sd_notify |
Services that must be ready before dependants start |
notify-reload |
As notify, plus RELOADING=1/READY=1 around reloads |
Services where systemctl reload must block until the reload completes |
dbus |
The configured BusName= appears on the bus |
D-Bus activated services |
idle |
As simple, but delayed until other jobs are dispatched |
Cosmetic only — keeps console output tidy at boot |
Choosing a type: prefer
execoversimpleunless you have a reason not to — it is the same model but reports a start failure rather than a "success" followed immediately by a crash. Onlynotify/notify-reload/dbus/forkinggive real readiness ordering: withsimpleandexec,After=says nothing about whether the dependency is actually serving, only that it was spawned.execrequires systemd 240+,notify-reloadrequires 253+.
Common Directives
Essential directives for configuring service behaviour and execution context.
Key Concepts
- Execution directives: Control how the service runs
- Resource directives: Limit CPU, memory, and I/O
- Security directives: Sandbox and restrict service capabilities
Execution Directives
[Service]
Type=exec
ExecStartPre=/usr/bin/app check
ExecStart=/usr/bin/app start
ExecStartPost=/usr/bin/notify-start
ExecReload=/bin/kill -HUP $MAINPID
ExecStop=/usr/bin/app stop
ExecStopPost=/usr/bin/cleanup
PIDFile=/run/app.pid
RemainAfterExit=yes
| Directive | Effect |
|---|---|
ExecStartPre= |
Runs before ExecStart=; a non-zero exit aborts the start. Repeatable |
ExecStart= |
The main command. Exactly one, except for Type=oneshot |
ExecStartPost= |
Runs after the service is considered started |
ExecReload= |
What systemctl reload runs |
ExecStop= |
Custom stop command; omit it and systemd sends SIGTERM |
ExecStopPost= |
Cleanup, run whether the stop succeeded or the service crashed |
ExecCondition= |
243+; a non-zero exit skips the unit without marking it failed |
PIDFile= |
Where a Type=forking daemon writes its PID. Absolute path under /run |
RemainAfterExit= |
Keep a oneshot unit "active" after its command exits |
Prefix an
Exec*=path with-to ignore a non-zero exit (ExecStartPre=-/usr/bin/optional), with+to run it with full privileges regardless ofUser=and the sandboxing, and with@to setargv[0]separately. OnlyExecStart=runs a shell if you ask for one —ExecStart=/bin/sh -c '...'— because systemd does no shell expansion of its own: globs, pipes,&&and$(...)in anExecStart=are passed through as literal arguments.
User and Permission Directives
[Service]
User=appuser
Group=appgroup
SupplementaryGroups=ssl-cert
WorkingDirectory=/opt/app
StateDirectory=myapp
CacheDirectory=myapp
LogsDirectory=myapp
RuntimeDirectory=myapp
| Directive | Effect |
|---|---|
User= / Group= |
Drop privileges to this account. The account must already exist |
DynamicUser=yes |
systemd allocates a transient UID for the lifetime of the service — no account to create, manage, or forget to remove |
SupplementaryGroups= |
Extra groups, e.g. ssl-cert for a key the service must read |
WorkingDirectory= |
cd here first; ~ means the User='s home |
StateDirectory= |
Creates and chowns /var/lib/myapp, and makes it writable under ProtectSystem=strict |
CacheDirectory= |
Likewise for /var/cache/myapp |
LogsDirectory= |
Likewise for /var/log/myapp |
RuntimeDirectory= |
Likewise for /run/myapp, removed when the service stops |
RootDirectory= |
chroot into this path before executing |
The
*Directory=directives are the ones people miss. They create the directory with the right owner and mode, add it to the sandbox's writable set automatically, and let you turn onProtectSystem=strictwithout hunting for the paths the service needs. Use them instead ofmkdirin anExecStartPre=.
Environment Directives
[Service]
Environment=HOME=/opt/app
Environment="KEY=value with spaces"
EnvironmentFile=/etc/myapp/env
EnvironmentFile=-/etc/myapp/optional
PassEnvironment=SSH_AUTH_SOCK
UnsetEnvironment=SECRET_KEY
- A
-prefix onEnvironmentFile=makes a missing file non-fatal. EnvironmentFile=is not a shell script: noexport, no command substitution, and quoting follows systemd's own rules.FOO=$BARsets the literal string$BAR.Environment=is a list, so a drop-in appends to it. Later definitions of the same key win.- Secrets in
Environment=are world-readable viasystemctl showand/proc/<pid>/environ. UseLoadCredential=instead — see the hardening sheet.
Resource Limits
[Service]
CPUQuota=50%
CPUWeight=100
MemoryMax=1G
MemoryHigh=800M
MemorySwapMax=0
IOWeight=100
IOReadBandwidthMax=/dev/sda 10M
IOWriteBandwidthMax=/dev/sda 5M
TasksMax=100
LimitNOFILE=65536
LimitNPROC=4096
| Directive | Effect |
|---|---|
CPUQuota= |
Hard ceiling as a percentage of one core — 200% is two cores |
CPUWeight= |
Relative share under contention only (1–10000, default 100) |
MemoryMax= |
Hard cap; exceeding it invokes the OOM killer on the unit |
MemoryHigh= |
Soft cap; the kernel throttles and reclaims above it instead of killing |
MemorySwapMax=0 |
No swap for this unit |
IOWeight= |
Relative I/O share (1–10000) |
IO*BandwidthMax= |
Absolute per-device ceiling; needs a real block device path |
TasksMax= |
Max processes and threads in the unit's cgroup |
LimitNOFILE=, LimitNPROC= |
Per-process setrlimit() ceilings, inherited by children |
Set
MemoryHigh=belowMemoryMax=and you get back-pressure before the kill.MemoryMax=on its own turns a slow leak into an abruptSIGKILLwith nothing useful in the journal.
Slices and the Resource Hierarchy
Every unit lands in a slice — a cgroup node that its resource limits are applied under. Setting limits on a slice budgets a whole group of services at once.
# /etc/systemd/system/myapp.slice
[Unit]
Description=Budget for all myapp services
[Slice]
# 2 cores across every unit in the slice
CPUQuota=200%
MemoryMax=4G
TasksMax=512
# In each member unit's [Service] section
Slice=myapp.slice
# Default slices: -.slice (root) -> system.slice, user.slice, machine.slice
systemd-cgls # cgroup tree with the processes in it
systemd-cgtop # live per-cgroup CPU/memory/IO
systemctl status myapp.slice # Slice membership and aggregate usage
# Read the effective value systemd actually applied
systemctl show myapp.service -p MemoryMax -p CPUQuota -p Slice
Gotcha: a limit only takes effect if the controller is delegated down to the unit's cgroup. systemd enables controllers on demand, so
MemoryMax=on a deeply nested unit can read back asinfinityif a parent slice never enabled the memory controller.systemctl show -p MemoryMaxis the source of truth, not the unit file.All of this assumes the unified (cgroup v2) hierarchy, the build-time default since systemd 243. systemd 256 refuses to boot under cgroup v1 without
SYSTEMD_CGROUP_ENABLE_LEGACY_FORCE=1on the kernel command line, and 258 removed cgroup v1 support outright. On a v1 box the directive names are the same but the semantics differ — checksystemd-analyze --versionfordefault-hierarchy=unifiedbefore trusting a v2 answer.
Security Directives
A reasonable starting sandbox for a network service that writes only to its own state directory:
[Service]
DynamicUser=yes
StateDirectory=myapp
NoNewPrivileges=yes
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
AmbientCapabilities=CAP_NET_BIND_SERVICE
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
SystemCallFilter=@system-service
SystemCallFilter=~@privileged
| Directive | Effect |
|---|---|
ProtectSystem=strict |
The whole filesystem is read-only except /dev, /proc, /sys and anything in ReadWritePaths= |
ProtectHome=yes |
/home, /root and /run/user are empty and inaccessible |
PrivateTmp=yes |
Private /tmp and /var/tmp, discarded on stop |
PrivateDevices=yes |
Minimal /dev — no raw disks, no /dev/kmem |
NoNewPrivileges=yes |
No setuid binary can raise privileges. Free, and almost never breaks anything |
CapabilityBoundingSet= |
The only capabilities the service may ever hold. Empty means none |
ProtectKernel*= |
Block writes to /proc/sys, /sys, and module loading |
RestrictAddressFamilies= |
Socket families the service may create |
SystemCallFilter= |
seccomp allow-list (@set) or deny-list (~@set) |
Verify the result rather than assuming it:
systemd-analyze security myapp.service # Score a loaded unit
systemd-analyze security --offline=true ./myapp.service # Score a file, no manager needed
Adding sandboxing to a running service breaks it more often than not — usually
ProtectSystem=strictwithout the matchingReadWritePaths=, orPrivateNetwork=yeson something that resolves DNS. Add directives a few at a time and restart between each batch.
For the full sandboxing surface — the systemd-analyze security workflow, credentials, IP and
namespace filtering, and the order to apply directives in without breaking the service — see the
companion systemd Service Hardening sheet.
Readiness, Watchdogs, and Failure Handling
Restart= brings a dead process back. Readiness protocols and watchdogs decide when it counts
as alive and how long a hung one is tolerated.
Key Concepts
- Readiness: with
Type=simple/exec, "started" means "spawned". Onlynotify,notify-reload,dbus, andforkingreport actual readiness, which is whatAfter=ordering depends on. - Watchdog: the service must ping systemd every
WatchdogSec=; miss one and systemd kills and (withRestart=on-watchdogoron-failure) restarts it. Catches hangs that a liveness-freeRestart=never sees. - Rate limiting:
StartLimitBurst=starts insideStartLimitIntervalSec=and systemd stops trying. Both live in[Unit], not[Service].
Readiness Notification
[Service]
Type=notify
# Only the main process may send (default)
NotifyAccess=main
ExecStart=/usr/bin/myapp
The service calls sd_notify(0, "READY=1") once it is actually serving. From Python, either use
the systemd-python bindings or write the datagram directly:
import os
import socket
def notify(state: str) -> None:
"""Send a state string to systemd's notification socket.
No-ops when NOTIFY_SOCKET is unset (i.e. not running under systemd), so the
same code runs unchanged in a container or a dev shell.
"""
addr = os.environ.get("NOTIFY_SOCKET")
if not addr:
return
if addr.startswith("@"): # Abstract namespace socket
addr = "\0" + addr[1:]
with socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM) as sock:
sock.connect(addr)
sock.sendall(state.encode("utf-8"))
notify("READY=1") # After the listener is bound
notify("STATUS=Serving on :8080") # Shows in `systemctl status`
notify("RELOADING=1\nMONOTONIC_USEC=…") # Type=notify-reload only
NotifyAccess=alllets any process in the cgroup send state changes, including a child that has no business doing so. Keep it atmainunless a supervisor forks the real worker.
Watchdogs
[Service]
Type=notify
ExecStart=/usr/bin/myapp
# Ping deadline
WatchdogSec=30
# Signal on miss (default; gives a core dump)
WatchdogSignal=SIGABRT
# Or on-failure, which also covers watchdog kills
Restart=on-watchdog
systemd exports WATCHDOG_USEC (microseconds) to the service. Ping at half that interval —
the deadline is not a suggestion:
# Ping at half the deadline. Keep this a float: WatchdogSec=1 gives
# WATCHDOG_USEC=1000000, and integer division would floor the interval to 0 --
# which reads as "no watchdog configured" and silently stops the pings.
usec = int(os.environ.get("WATCHDOG_USEC", "0"))
interval = usec / 2_000_000 if usec else 0.0 # microseconds -> seconds, halved
if interval:
notify("WATCHDOG=1") # Repeat every `interval` seconds
Ping from the loop that does the real work, not a bare timer thread. A watchdog fed by a thread that keeps running while the request loop is wedged reports health it cannot observe.
Restart Strategy
[Unit]
# Window for counting starts
StartLimitIntervalSec=300
# Give up after 5 starts in that window
StartLimitBurst=5
# Run this unit when we enter failed state
OnFailure=notify-admin@%n.service
# 249+; the mirror of OnFailure
OnSuccess=cleanup.service
[Service]
# no | always | on-success | on-failure |
Restart=on-failure
# on-abnormal | on-abort | on-watchdog
# Fixed delay between attempts
RestartSec=5
# 254+: ramp RestartSec -> RestartMaxDelaySec
RestartSteps=5
# over this many steps (exponential backoff)
RestartMaxDelaySec=60
| Value | Restarts on |
|---|---|
no |
Never (default) |
on-success |
Clean exit only |
on-failure |
Non-zero exit, signal, timeout, or watchdog — the usual choice |
on-abnormal |
Signal, timeout, or watchdog, but not a non-zero exit code |
on-abort |
Uncaught signal only |
on-watchdog |
Watchdog timeout only |
always |
Any exit, clean or not |
Footgun:
Restart=alwaysplus a config error is an infinite crash loop that fills the journal and, with noStartLimitBurst=, never gives up.Restart=on-failurewith a start limit is almost always what you want. When a unit does hit its limit it stays failed untilsystemctl reset-failed— a restart alone will not clear it.
The OnFailure= handler receives the failed unit's name via %n, so one templated notifier
serves every unit:
# /etc/systemd/system/notify-admin@.service
[Unit]
Description=Alert on failure of %I
[Service]
Type=oneshot
ExecStart=/usr/local/bin/alert "unit %I failed on %H"
Dependencies
Control service start order and relationships between units.
Key Concepts
flowchart LR
subgraph "Dependency Types"
A[Wants] -->|Weak| B[Target]
C[Requires] -->|Strong| B
D[After] -->|Order| B
E[Before] -->|Order| B
end
- Requires: Hard dependency; failure stops dependent unit
- Wants: Soft dependency; failure doesn't affect dependent unit
- After/Before: Ordering only; no activation dependency
- BindsTo: Like Requires, but also stops when dependency stops
Dependency Directives
[Unit]
# Ordering (when to start)
# Start after these units
After=network.target postgresql.service
# Start before this unit
Before=httpd.service
# Activation dependencies
# Hard dependency (fail if missing)
Requires=postgresql.service
# Soft dependency (continue if missing)
Wants=redis.service
# Stop when dependency stops
BindsTo=docker.service
# Conflict handling
# Cannot run simultaneously
Conflicts=sendmail.service
# Conditional execution
# Only start if path exists
ConditionPathExists=/etc/myapp/config
# Start if path does NOT exist
ConditionPathExists=!/etc/myapp/disable
# Start if file has content
ConditionFileNotEmpty=/etc/myapp/config
# Fail if path missing
AssertPathExists=/data
Common Dependency Patterns
# Web application with database
[Unit]
Description=Web Application
After=network.target postgresql.service redis.service
Requires=postgresql.service
Wants=redis.service
# Service requiring network
[Unit]
Description=Network Service
After=network-online.target
Wants=network-online.target
# Mount-dependent service
[Unit]
Description=Data Processing Service
After=data.mount
RequiresMountsFor=/data
Dependency Inspection
# Show unit dependencies
systemctl list-dependencies nginx.service
# Show reverse dependencies
systemctl list-dependencies --reverse nginx.service
# Show ordering dependencies
systemctl list-dependencies --after nginx.service
systemctl list-dependencies --before nginx.service
Targets and Boot Process
Targets group units and define system states, replacing traditional runlevels.
Key Concepts
flowchart TD
A[sysinit.target] --> B[basic.target]
B --> C[network.target]
C --> D[network-online.target]
B --> E[multi-user.target]
E --> F[graphical.target]
G[rescue.target] --> A
H[emergency.target]
style E fill:#90EE90
style F fill:#87CEEB
Common Targets
| Target | Description | Runlevel Equivalent |
|---|---|---|
poweroff.target |
System shutdown | 0 |
rescue.target |
Single-user mode | 1 |
multi-user.target |
Multi-user, no GUI | 3 |
graphical.target |
Multi-user with GUI | 5 |
reboot.target |
System reboot | 6 |
emergency.target |
Emergency shell | - |
Target Management
# View current target
systemctl get-default
# Set default target
sudo systemctl set-default multi-user.target
# Switch to target immediately
sudo systemctl isolate rescue.target
# List all targets
systemctl list-units --type=target
# Show target dependencies
systemctl list-dependencies graphical.target
Creating Custom Targets
# /etc/systemd/system/myapp.target
[Unit]
Description=My Application Stack
Requires=multi-user.target
After=multi-user.target
AllowIsolate=yes
[Install]
WantedBy=multi-user.target
# Services belonging to custom target
[Install]
WantedBy=myapp.target
Boot Process Overview
# View boot process
systemd-analyze # Boot time summary
systemd-analyze blame # Time per unit
systemd-analyze critical-chain # Critical path
systemd-analyze plot > boot.svg # Visual boot chart
# Boot targets
systemd-analyze verify myservice.service # Validate unit file
Timer Units for Scheduling
Timer units replace cron for scheduling tasks with better logging and dependency management.
Key Concepts
- Timers activate corresponding
.serviceunits - Support both real-time (calendar) and monotonic (relative) timers
- Persistent timers catch up on missed runs
Timer Types
flowchart LR
subgraph "Timer Types"
A[Realtime] --> B[OnCalendar]
C[Monotonic] --> D[OnBootSec]
C --> E[OnUnitActiveSec]
C --> F[OnStartupSec]
end
Calendar Timer Examples
# /etc/systemd/system/backup.timer
[Unit]
Description=Daily Backup Timer
[Timer]
# Calendar expressions
# Every day at midnight
OnCalendar=daily
# Every Monday at midnight
OnCalendar=weekly
# Every day at 04:00
OnCalendar=*-*-* 04:00:00
# Weekdays at 09:00
OnCalendar=Mon..Fri *-*-* 09:00:00
# First of every month
OnCalendar=*-*-01 00:00:00
# Every 15 minutes
OnCalendar=*:0/15
# Run if missed
Persistent=true
# Random delay up to 5 min
RandomizedDelaySec=300
# Timer accuracy
AccuracySec=1s
[Install]
WantedBy=timers.target
Monotonic Timer Examples
# /etc/systemd/system/cleanup.timer
[Unit]
Description=Periodic Cleanup Timer
[Timer]
# 5 min after boot
OnBootSec=5min
# 1 hour after last activation
OnUnitActiveSec=1h
# 10 min after systemd start
OnStartupSec=10min
[Install]
WantedBy=timers.target
Service for Timer
# /etc/systemd/system/backup.service
[Unit]
Description=Daily Backup Service
[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup.sh
User=backup
Nice=19
IOSchedulingClass=idle
Timer Management
# List active timers
systemctl list-timers
systemctl list-timers --all
# Enable and start timer
sudo systemctl enable backup.timer
sudo systemctl start backup.timer
# Check timer status
systemctl status backup.timer
# Test calendar expression
systemd-analyze calendar "Mon..Fri *-*-* 09:00:00"
systemd-analyze calendar --iterations=5 "daily"
# Manually trigger the service
sudo systemctl start backup.service
Calendar Expression Reference
| Expression | Description |
|---|---|
minutely |
Every minute |
hourly |
Every hour at :00 |
daily |
Every day at 00:00 |
weekly |
Every Monday at 00:00 |
monthly |
First of month at 00:00 |
*:0/15 |
Every 15 minutes |
*-*-* 04:00:00 |
Daily at 04:00 |
Mon,Wed,Fri *-*-* 10:00:00 |
Specific days at 10:00 |
*-*-1,15 00:00:00 |
1st and 15th of month |
Socket Activation
systemd opens the listening socket, then hands it to the service. The service never binds anything itself.
sequenceDiagram
participant C as Client
participant S as systemd (PID 1)
participant A as myapp.service
S->>S: Bind myapp.socket at boot
C->>S: connect() on port 8080
Note over S: Kernel queues the connection
S->>A: Start service, pass fd 3
A->>A: sd_listen_fds() then accept()
A-->>C: Response
Note over C,A: Later connections go straight to the running service
Three things this buys you: services start on first use rather than at boot; a restart drops no
pending connections, because PID 1 keeps the listening socket and the kernel queues arrivals in
its backlog while the service is down; and a service can serve port 443 without ever holding
CAP_NET_BIND_SERVICE, because PID 1 did the binding.
That second one is narrower than it sounds. Connections the service has already
accept()ed are its own file descriptors and die with it — socket activation is not a substitute for draining established sessions before shutdown. What it removes is the connection-refused window between stop and start.
Socket Unit Structure
# /etc/systemd/system/myapp.socket - name must match myapp.service
[Unit]
Description=Socket for myapp
[Socket]
# TCP; also ListenDatagram=, ListenFIFO=
ListenStream=0.0.0.0:8080
# A unit may listen on several addresses
ListenStream=/run/myapp/api.sock
# Default: one service, all connections
Accept=no
Backlog=1024
# Owner of a filesystem socket
SocketUser=www-data
SocketMode=0660
# Exported in $LISTEN_FDNAMES
FileDescriptorName=api
[Install]
WantedBy=sockets.target
# Enable the socket, NOT the service - the socket pulls the service in
sudo systemctl enable --now myapp.socket
systemctl list-sockets # Sockets, their units, and activation state
systemctl status myapp.socket
Accept=no vs Accept=yes
Accept=no (default) |
Accept=yes |
|
|---|---|---|
| Service instances | One, handling every connection | One per connection |
| Unit named | myapp.service |
myapp@.service (template) |
| Socket reaches the app as | fd 3 onwards, via $LISTEN_FDS |
stdin/stdout, inetd-style |
| Use for | Anything long-lived | Rare, low-rate protocols |
Accept=yesforks a process per connection. It is the inetd model and it is a denial-of-service vector on a busy port — default toAccept=no.
Picking Up the Socket
systemd passes LISTEN_FDS (how many), LISTEN_PID (which process they are for), and
LISTEN_FDNAMES (the FileDescriptorName= labels). Descriptors start at fd 3:
import os
import socket
SD_LISTEN_FDS_START = 3
def listen_fds() -> list[socket.socket]:
# Return the sockets systemd passed to this process, or [] if there are none.
# LISTEN_PID guards against a forked child inheriting the variables and
# claiming descriptors it does not own.
try:
if int(os.environ.get("LISTEN_PID", 0)) != os.getpid():
return []
count = int(os.environ.get("LISTEN_FDS", 0))
except ValueError:
# Malformed environment - treat as "no sockets" rather than crashing.
return []
return [socket.socket(fileno=SD_LISTEN_FDS_START + i) for i in range(count)]
socks = listen_fds()
if socks:
server = socks[0] # Already bound and listening
else:
server = socket.create_server(("0.0.0.0", 8080)) # Standalone fallback
Unset
LISTEN_PID/LISTEN_FDSafter reading them if you fork, or the child will think the sockets are its own. The C helpersd_listen_fds(1, ...)does this for you; most language bindings expose the same "unset environment" flag.
Drop-in Overrides
Never edit a packaged unit in /usr/lib/systemd/system/ — the next upgrade overwrites it. Layer
a drop-in instead.
# Creates /etc/systemd/system/myapp.service.d/override.conf, then daemon-reloads
sudo systemctl edit myapp.service
# Named drop-in, so separate concerns stay in separate files (253+)
sudo systemctl edit --drop-in=hardening.conf myapp.service
# Copy the whole unit into /etc for wholesale replacement
sudo systemctl edit --full myapp.service
# Temporary override in /run, gone at reboot
sudo systemctl edit --runtime myapp.service
# Throw away every local override and go back to the packaged unit
sudo systemctl revert myapp.service
Seeing What Actually Applies
systemctl cat myapp.service # Merged view, with each source file named
systemctl show myapp.service # Every resolved property, post-merge
systemctl show myapp.service -p ExecStart -p MemoryMax
systemd-delta --type=extended # Every drop-in on the system
sudo systemctl daemon-reload # Required after any hand-edited unit file
Precedence, lowest to highest: /usr/lib/systemd/system/ (packages) ->
/run/systemd/system/ (runtime) -> /etc/systemd/system/ (admin). /etc outranks /run —
confirm with systemd-analyze unit-paths, which prints the search path highest-priority first.
So a systemctl edit --runtime override does not beat a persistent one of the same name in
/etc; it beats the packaged unit only. Drop-ins from all three locations are merged, and within
a unit's .d/ directory files apply in lexical order, which is why they conventionally start
10-, 20-.
The List-Directive Reset Footgun
Most directives are last-one-wins. ExecStart=, Environment=, ReadWritePaths=,
AssertPathExists= and friends are lists — a drop-in appends to them. To replace rather than
append, assign the empty value first:
# /etc/systemd/system/myapp.service.d/10-command.conf
[Service]
# Clear the inherited list
ExecStart=
ExecStart=/usr/bin/myapp --config /etc/myapp/prod.yaml
Without the empty ExecStart=, systemd refuses to load the unit at all:
myapp.service: Service has more than one ExecStart= setting,
which is only allowed for Type=oneshot services. Refusing.
Environment= fails more quietly — it accumulates, and the last definition of a given key wins,
so a stale value in the base unit can survive an override that looked correct.
Dependencies are the exception, and there is no workaround.
After=,Before=,Requires=,Wants=,BindsTo=and the rest cannot be reset: systemd.unit(5) states that "dependencies can only be added in drop-ins. If you want to remove dependencies, you have to override the entire unit." An emptyAfter=in a drop-in is not an error and not a reset — it simply does nothing. Usesystemctl edit --full(orsystemctl maskplus your own unit) to drop one.
# Validate a unit and its drop-ins together before reloading
systemd-analyze verify /etc/systemd/system/myapp.service
sudo systemctl mask myapp.service # Symlink to /dev/null: cannot be started at all
sudo systemctl unmask myapp.service
User Services and Lingering
Every logged-in user gets their own systemd --user manager. Units live in
~/.config/systemd/user/ and run as that user, with no root anywhere in the path.
systemctl --user daemon-reload
systemctl --user enable --now myapp.service
systemctl --user status myapp.service
journalctl --user -u myapp.service -f
systemctl --user list-units --type=service
# ~/.config/systemd/user/myapp.service - note the [Install] target
[Unit]
Description=My user service
[Service]
# %h = the user's home directory
ExecStart=%h/.local/bin/myapp
Restart=on-failure
[Install]
# NOT multi-user.target
WantedBy=default.target
Lingering
By default the user manager starts at login and is torn down with the last session — a user service stops when you log out, and never starts at boot.
sudo loginctl enable-linger mike # Start at boot, survive logout
loginctl show-user mike -p Linger # Linger=yes
sudo loginctl disable-linger mike
This is the single most common "my user timer never fires" cause. Enabling the unit is not enough; without lingering there is no manager running to fire it.
Differences from system units worth knowing:
User=andGroup=are meaningless — you are already unprivileged. The namespacing directives (ProtectSystem=,ProtectHome=,PrivateTmp=,PrivateDevices=) also do nothing in a user service, because filesystem namespacing is privileged; addPrivateUsers=yesand most of them start working. Resource control (MemoryMax=,CPUQuota=) works as normal.systemctl --userneedsXDG_RUNTIME_DIRandDBUS_SESSION_BUS_ADDRESS. A baresudo -u mike systemctl --user ...fails with "Failed to connect to bus"; usemachinectl shell mike@orsystemd-run --machine=mike@ --userinstead.- Rootless Podman Quadlets are user units, so the same lingering rule governs whether your containers come back after a reboot.
Transient Units with systemd-run
Run a command under systemd's supervision without writing a unit file. Ideal for one-off jobs that need a resource cap, a timeout, or journal capture.
# One-off supervised command; output goes to the journal
sudo systemd-run --unit=reindex /usr/local/bin/reindex.sh
journalctl -u reindex -f
# Wait for it and get its exit status back
sudo systemd-run --wait --collect --unit=reindex /usr/local/bin/reindex.sh
# Attach stdin/stdout instead of the journal
sudo systemd-run --pipe --wait /usr/bin/du -sh /var/log
# Cap a heavy job so it cannot take the box down
sudo systemd-run --property=MemoryMax=2G --property=CPUQuota=50% \
--property=IOWeight=10 --unit=backup /usr/local/bin/backup.sh
# Transient timers - no .timer file needed
sudo systemd-run --on-active=90 /usr/bin/systemctl restart nginx
sudo systemd-run --on-calendar='*-*-* 03:00:00' --unit=nightly /usr/local/bin/nightly.sh
# Put an interactive build under a resource cap (scope, not service)
systemd-run --user --scope --property=MemoryMax=4G -- make -j8
| Flag | Effect |
|---|---|
--unit=NAME |
Name the transient unit instead of taking run-<pid> |
--scope |
Register the calling process's children rather than forking a service |
--wait |
Block until it finishes and propagate the exit status |
--collect (-G) |
Unload the unit afterwards, even on failure |
--pipe (-P) |
Pass stdin/stdout/stderr through instead of journalling |
--property=K=V (-p) |
Any [Service] directive, including the sandboxing ones |
--user |
Run under your own user manager |
--scoperuns the command in the caller's context — noExecStart=, no restart, no sandboxing. It is for putting an existing process tree (a build, a shell) into a cgroup. Everything else wants the default service mode.
Trying a hardening directive before committing it to a unit file is the killer use:
sudo systemd-run --pipe --wait -p ProtectSystem=strict -p PrivateTmp=yes \
-p CapabilityBoundingSet= /usr/local/bin/myapp --self-test
Troubleshooting
Tools and techniques for diagnosing systemd service issues.
Key Concepts
- journalctl: Centralised logging for all systemd units
- systemctl status: Quick overview of unit state
- systemd-analyze: Boot and performance analysis
Checking Service Status
# Basic status
systemctl status nginx.service
systemctl status nginx # .service suffix optional
# Detailed status
systemctl show nginx.service # All properties
systemctl show nginx -p MainPID # Specific property
systemctl show nginx -p ActiveState,SubState
# Check if active/enabled
systemctl is-active nginx
systemctl is-enabled nginx
systemctl is-failed nginx
Viewing Logs with journalctl
# Service logs
journalctl -u nginx.service # All logs for unit
journalctl -u nginx -f # Follow logs (like tail -f)
journalctl -u nginx --since today # Today's logs
journalctl -u nginx --since "1 hour ago"
journalctl -u nginx -n 50 # Last 50 lines
journalctl -u nginx -p err # Errors only
# Priority levels: emerg, alert, crit, err, warning, notice, info, debug
journalctl -u nginx -p warning # Warning and above
# Output formats
journalctl -u nginx -o json # JSON output
journalctl -u nginx -o json-pretty # Pretty JSON
journalctl -u nginx -o cat # Plain text, no metadata
# Boot-specific logs
journalctl -b # Current boot
journalctl -b -1 # Previous boot
journalctl --list-boots # List all boots
# Kernel messages
journalctl -k # Kernel messages only
journalctl -k -p err # Kernel errors
Finding Failed Units
# List failed units
systemctl --failed
systemctl list-units --state=failed
# Reset failed state
sudo systemctl reset-failed
sudo systemctl reset-failed nginx.service
# Find units in specific states
systemctl list-units --state=inactive
systemctl list-units --state=activating
Debugging Service Start Failures
# Verbose unit start
sudo systemctl start nginx.service --no-block
journalctl -u nginx -f
# Check syntax
systemd-analyze verify /etc/systemd/system/myapp.service
# Test service manually
sudo -u appuser /usr/bin/myapp --config /etc/myapp/config.yaml
# Check dependencies
systemctl list-dependencies nginx.service
systemctl list-dependencies --reverse nginx.service
# Environment issues
systemctl show nginx -p Environment
systemctl cat nginx.service # Show unit file contents
Common Diagnostic Commands
# Reload configuration after changes
sudo systemctl daemon-reload
# Check disk space for logs
journalctl --disk-usage
# Vacuum old logs
sudo journalctl --vacuum-time=7d # Keep 7 days
sudo journalctl --vacuum-size=500M # Keep 500MB
# System state overview
systemctl status # Overall system status
systemctl list-units # All loaded units
systemctl list-unit-files # All installed units
# Check resource usage
systemd-cgtop # Top for cgroups
systemctl status myapp --no-pager -l # Full output, no pager
Quick Reference
Essential Commands
| Command | Description |
|---|---|
systemctl start unit |
Start a unit |
systemctl stop unit |
Stop a unit |
systemctl restart unit |
Restart a unit |
systemctl reload unit |
Reload configuration |
systemctl enable unit |
Enable at boot |
systemctl disable unit |
Disable at boot |
systemctl status unit |
Show unit status |
systemctl daemon-reload |
Reload unit files |
systemctl mask unit |
Prevent unit from starting |
systemctl unmask unit |
Remove mask |
systemctl cat unit |
Merged unit file plus every drop-in |
systemctl edit unit |
Create/edit a drop-in override |
systemctl revert unit |
Discard local overrides |
systemctl show unit -p X |
Resolved value of property X |
systemctl reset-failed unit |
Clear a failed state and its start-limit counter |
systemctl list-sockets |
Socket units and what they activate |
systemctl --user ... |
Operate on your own user manager |
Inspection and Analysis
| Command | Description |
|---|---|
systemd-analyze verify unit.service |
Validate syntax, directive names, and dependencies |
systemd-analyze security unit.service |
Score the sandbox; --offline=true works on a file |
systemd-analyze calendar 'Mon *-*-* 02:00' |
Normalise and show the next elapse of a timer expression |
systemd-analyze syscall-filter @mount |
List the syscalls in a SystemCallFilter= set |
systemd-analyze blame |
Per-unit boot time, slowest first |
systemd-analyze critical-chain |
The time-critical path through boot |
systemd-analyze cat-config systemd/journald.conf |
Merged view of a config file and its drop-ins |
systemd-delta |
Every unit overridden on the system |
systemd-escape -p --suffix=mount /mnt/data |
Compute the unit name for a path |
systemd-cgls / systemd-cgtop |
cgroup tree / live per-cgroup resource usage |
Common journalctl Flags
| Flag | Description |
|---|---|
-u unit |
Show logs for unit |
-f |
Follow (tail) logs |
-n N |
Show last N lines |
-p priority |
Filter by priority |
-b |
Current boot only |
--since "time" |
Logs since time |
-o format |
Output format |
Unit File Locations
| Location | Purpose |
|---|---|
/etc/systemd/system/ |
Local admin units (highest priority) |
/run/systemd/system/ |
Runtime units |
/usr/lib/systemd/system/ |
Package-installed units |
~/.config/systemd/user/ |
User units |
Common Issues and Solutions
Service Fails to Start
# Check logs for errors
journalctl -u myapp -n 100 --no-pager
# Verify unit file syntax
systemd-analyze verify /etc/systemd/system/myapp.service
# Check ExecStart path and permissions
ls -la /usr/bin/myapp
sudo -u appuser /usr/bin/myapp --help
# Ensure dependencies are running
systemctl status postgresql.service
Service Keeps Restarting
# Check restart limits
systemctl show myapp -p StartLimitBurst,StartLimitIntervalSec
# View recent failures
journalctl -u myapp --since "10 min ago"
# Temporarily disable restart for debugging
# In [Service]: Restart=no
sudo systemctl daemon-reload
sudo systemctl restart myapp
Timer Not Triggering
# Check timer status
systemctl status backup.timer
# Verify calendar expression
systemd-analyze calendar "*-*-* 04:00:00" --iterations=3
# Ensure timer is enabled and started
sudo systemctl enable --now backup.timer
# Check if service exists
systemctl cat backup.service
Permission Denied Errors
# Check User/Group directives
systemctl show myapp -p User,Group
# Verify file permissions
ls -la /opt/myapp/
ls -la /var/lib/myapp/
# Check SELinux/AppArmor
getenforce # SELinux status
ausearch -m AVC -ts recent # Recent denials
aa-status # AppArmor status
Unit Not Found After Creation
# Reload systemd configuration
sudo systemctl daemon-reload
# Check file permissions
ls -la /etc/systemd/system/myapp.service
# Verify unit file name matches
systemctl cat myapp.service
Logs Not Appearing
# Check if journald is running
systemctl status systemd-journald
# Verify logging configuration
cat /etc/systemd/journald.conf
# Check disk space
df -h /var/log/journal
# Restart journald
sudo systemctl restart systemd-journald
Drop-in Override Ignored
# Confirm the drop-in is being read at all, and from where
systemctl cat myapp.service
# Compare the resolved value against what you wrote
systemctl show myapp.service -p ExecStart
# Reload after any hand-edited file
sudo systemctl daemon-reload
Fix: the three usual causes are a missing daemon-reload; a .d directory whose name omits
the unit suffix (myapp.d/, not myapp.service.d/); and a file that does not end in .conf,
which systemd silently ignores. For list-valued directives, remember the empty-assignment reset.
Unit Fails After Adding Sandboxing
# Which directive actually broke it
journalctl -u myapp.service -b --no-pager | tail -30
# Bisect by running the binary under the suspect directives only
sudo systemd-run --pipe --wait -p ProtectSystem=strict -p PrivateDevices=yes /usr/bin/myapp
Fix: ProtectSystem=strict without a matching ReadWritePaths= is the usual culprit — the
service writes somewhere it can no longer reach. PrivateNetwork=yes breaks anything that talks
to the network, including DNS. A code=exited, status=203/EXEC means the binary could not be
executed at all — often NoExecPaths= or a RootDirectory= that hides it.
User Service Does Not Start at Boot
loginctl show-user "$USER" -p Linger # Linger=no means no manager at boot
sudo loginctl enable-linger "$USER"
systemctl --user is-enabled myapp.service
Fix: enable lingering. Also check the [Install] section targets default.target — a user
unit installed to multi-user.target enables without error and never starts.
Socket Activated Service Never Starts
systemctl list-sockets # Is the socket even listening?
systemctl status myapp.socket
ss -lntp | grep 8080 # Confirm the kernel has the port
Fix: enable the socket, not the service (systemctl enable --now myapp.socket). The unit
names must match (myapp.socket → myapp.service) unless the socket sets Service=. If the
service starts and immediately exits, it is probably binding its own port instead of picking up
$LISTEN_FDS.
Related Topics
- systemd Service Hardening — the sandboxing directives,
systemd-analyze security, and the order to apply them in - Linux CLI — the surrounding shell fluency, including
journalctlin day-to-day use - Linux Storage Management —
.mountunits,x-systemd.automount, andsystemd-escape - Linux Performance Analysis — reading the cgroup accounting that
MemoryMax=andCPUQuota=drive - launchctl — the macOS equivalent, and where the two models genuinely differ
- Podman Quadlets and systemd — containers declared as units, and the dependency graph Quadlet generates
- SELinux / AppArmor — the MAC layer that sits underneath systemd's own sandboxing