Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

systemd

Init system and service manager for Linux: units, timers, sockets, drop-in overrides, resource control, and journald.

systemd

Init system and service manager for Linux, managing system processes, services, and boot targets.

Overview

systemd is the default init system for most modern Linux distributions. It manages the complete lifecycle of system services, handles dependencies between units, and provides extensive logging through journald. Units are the fundamental building blocks, with services, timers, mounts, and targets being the most commonly used types.

systemd Architecturesystemd PID 1Service UnitsTimer UnitsMount UnitsTarget UnitsSocket Unitsjournaldmulti-user.targetgraphical.targetnetwork.targetsystemd Architecturesystemd PID 1Service UnitsTimer UnitsMount UnitsTarget UnitsSocket Unitsjournaldmulti-user.targetgraphical.targetnetwork.target

Unit File Structure

Unit files define how systemd manages services, timers, mounts, and other resources. They use INI-style syntax with sections denoted by square brackets.

Key Concepts

  • Unit files location: /etc/systemd/system/ (admin), /usr/lib/systemd/system/ (packages)
  • Override files: /etc/systemd/system/myapp.service.d/*.conf (directory carries the full unit name)
  • User units: ~/.config/systemd/user/

Comments must be on their own line. Unit files are not shell or INI-with-trailing-comments: a # after a value is part of the value. PrivateTmp=yes # sandbox silently parses as the boolean "yes # sandbox", fails, and is dropped — you get a service with no sandboxing and only a Failed to parse boolean value, ignoring line buried in the journal. On ExecStart= it is worse: the comment text becomes extra argv entries. Put every comment on a line of its own, and run systemd-analyze verify before trusting a unit you hand-edited.

Service Unit Structure

# /etc/systemd/system/myapp.service
[Unit]
Description=My Application Service
Documentation=https://example.com/docs
After=network.target postgresql.service
Requires=postgresql.service

[Service]
Type=simple
User=appuser
Group=appgroup
WorkingDirectory=/opt/myapp
Environment=NODE_ENV=production
EnvironmentFile=/etc/myapp/env
ExecStartPre=/usr/bin/myapp-check
ExecStart=/usr/bin/myapp --config /etc/myapp/config.yaml
ExecReload=/bin/kill -HUP $MAINPID
# No ExecStop= needed: systemd sends SIGTERM to the main process, then SIGKILL
# after TimeoutStopSec=. Only set ExecStop= if the daemon needs a custom drain
# command; `ExecStop=/bin/kill -TERM $MAINPID` is a cargo-culted no-op.
Restart=on-failure
RestartSec=5
TimeoutStartSec=30
TimeoutStopSec=30

[Install]
WantedBy=multi-user.target

Timer Unit Structure

# /etc/systemd/system/backup.timer
[Unit]
Description=Daily Backup Timer

[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true
RandomizedDelaySec=300

[Install]
WantedBy=timers.target

Mount Unit Structure

# /etc/systemd/system/data.mount
[Unit]
Description=Data Volume Mount

[Mount]
What=/dev/sdb1
Where=/data
Type=ext4
Options=defaults,noatime

[Install]
WantedBy=multi-user.target

The filename is not free-form. A mount unit's name must be the escaped form of its Where= path, or systemd refuses to load it with Where= setting doesn't match unit name. Compute it rather than guessing: systemd-escape -p --suffix=mount /srv/data gives srv-data.mount. The same rule governs .automount, .swap, and .device units.

Service Types

Type Considered "started" when Use for
simple systemd forks the process (immediately) Default; foreground processes that need no ordering guarantee
exec The binary has been execve()d successfully Better default than simple — catches a missing/unexecutable ExecStart at start time
forking The parent exits after forking the child Traditional daemons that background themselves
oneshot The process exits Scripts and setup steps; pair with RemainAfterExit=yes
notify The process sends READY=1 via sd_notify Services that must be ready before dependants start
notify-reload As notify, plus RELOADING=1/READY=1 around reloads Services where systemctl reload must block until the reload completes
dbus The configured BusName= appears on the bus D-Bus activated services
idle As simple, but delayed until other jobs are dispatched Cosmetic only — keeps console output tidy at boot

Choosing a type: prefer exec over simple unless you have a reason not to — it is the same model but reports a start failure rather than a "success" followed immediately by a crash. Only notify/notify-reload/dbus/forking give real readiness ordering: with simple and exec, After= says nothing about whether the dependency is actually serving, only that it was spawned. exec requires systemd 240+, notify-reload requires 253+.

Common Directives

Essential directives for configuring service behaviour and execution context.

Key Concepts

  • Execution directives: Control how the service runs
  • Resource directives: Limit CPU, memory, and I/O
  • Security directives: Sandbox and restrict service capabilities

Execution Directives

[Service]
Type=exec
ExecStartPre=/usr/bin/app check
ExecStart=/usr/bin/app start
ExecStartPost=/usr/bin/notify-start
ExecReload=/bin/kill -HUP $MAINPID
ExecStop=/usr/bin/app stop
ExecStopPost=/usr/bin/cleanup
PIDFile=/run/app.pid
RemainAfterExit=yes
Directive Effect
ExecStartPre= Runs before ExecStart=; a non-zero exit aborts the start. Repeatable
ExecStart= The main command. Exactly one, except for Type=oneshot
ExecStartPost= Runs after the service is considered started
ExecReload= What systemctl reload runs
ExecStop= Custom stop command; omit it and systemd sends SIGTERM
ExecStopPost= Cleanup, run whether the stop succeeded or the service crashed
ExecCondition= 243+; a non-zero exit skips the unit without marking it failed
PIDFile= Where a Type=forking daemon writes its PID. Absolute path under /run
RemainAfterExit= Keep a oneshot unit "active" after its command exits

Prefix an Exec*= path with - to ignore a non-zero exit (ExecStartPre=-/usr/bin/optional), with + to run it with full privileges regardless of User= and the sandboxing, and with @ to set argv[0] separately. Only ExecStart= runs a shell if you ask for one — ExecStart=/bin/sh -c '...' — because systemd does no shell expansion of its own: globs, pipes, && and $(...) in an ExecStart= are passed through as literal arguments.

User and Permission Directives

[Service]
User=appuser
Group=appgroup
SupplementaryGroups=ssl-cert
WorkingDirectory=/opt/app
StateDirectory=myapp
CacheDirectory=myapp
LogsDirectory=myapp
RuntimeDirectory=myapp
Directive Effect
User= / Group= Drop privileges to this account. The account must already exist
DynamicUser=yes systemd allocates a transient UID for the lifetime of the service — no account to create, manage, or forget to remove
SupplementaryGroups= Extra groups, e.g. ssl-cert for a key the service must read
WorkingDirectory= cd here first; ~ means the User='s home
StateDirectory= Creates and chowns /var/lib/myapp, and makes it writable under ProtectSystem=strict
CacheDirectory= Likewise for /var/cache/myapp
LogsDirectory= Likewise for /var/log/myapp
RuntimeDirectory= Likewise for /run/myapp, removed when the service stops
RootDirectory= chroot into this path before executing

The *Directory= directives are the ones people miss. They create the directory with the right owner and mode, add it to the sandbox's writable set automatically, and let you turn on ProtectSystem=strict without hunting for the paths the service needs. Use them instead of mkdir in an ExecStartPre=.

Environment Directives

[Service]
Environment=HOME=/opt/app
Environment="KEY=value with spaces"
EnvironmentFile=/etc/myapp/env
EnvironmentFile=-/etc/myapp/optional
PassEnvironment=SSH_AUTH_SOCK
UnsetEnvironment=SECRET_KEY
  • A - prefix on EnvironmentFile= makes a missing file non-fatal.
  • EnvironmentFile= is not a shell script: no export, no command substitution, and quoting follows systemd's own rules. FOO=$BAR sets the literal string $BAR.
  • Environment= is a list, so a drop-in appends to it. Later definitions of the same key win.
  • Secrets in Environment= are world-readable via systemctl show and /proc/<pid>/environ. Use LoadCredential= instead — see the hardening sheet.

Resource Limits

[Service]
CPUQuota=50%
CPUWeight=100
MemoryMax=1G
MemoryHigh=800M
MemorySwapMax=0
IOWeight=100
IOReadBandwidthMax=/dev/sda 10M
IOWriteBandwidthMax=/dev/sda 5M
TasksMax=100
LimitNOFILE=65536
LimitNPROC=4096
Directive Effect
CPUQuota= Hard ceiling as a percentage of one core — 200% is two cores
CPUWeight= Relative share under contention only (1–10000, default 100)
MemoryMax= Hard cap; exceeding it invokes the OOM killer on the unit
MemoryHigh= Soft cap; the kernel throttles and reclaims above it instead of killing
MemorySwapMax=0 No swap for this unit
IOWeight= Relative I/O share (1–10000)
IO*BandwidthMax= Absolute per-device ceiling; needs a real block device path
TasksMax= Max processes and threads in the unit's cgroup
LimitNOFILE=, LimitNPROC= Per-process setrlimit() ceilings, inherited by children

Set MemoryHigh= below MemoryMax= and you get back-pressure before the kill. MemoryMax= on its own turns a slow leak into an abrupt SIGKILL with nothing useful in the journal.

Slices and the Resource Hierarchy

Every unit lands in a slice — a cgroup node that its resource limits are applied under. Setting limits on a slice budgets a whole group of services at once.

# /etc/systemd/system/myapp.slice
[Unit]
Description=Budget for all myapp services

[Slice]
# 2 cores across every unit in the slice
CPUQuota=200%
MemoryMax=4G
TasksMax=512
# In each member unit's [Service] section
Slice=myapp.slice
# Default slices: -.slice (root) -> system.slice, user.slice, machine.slice
systemd-cgls                            # cgroup tree with the processes in it
systemd-cgtop                           # live per-cgroup CPU/memory/IO
systemctl status myapp.slice            # Slice membership and aggregate usage

# Read the effective value systemd actually applied
systemctl show myapp.service -p MemoryMax -p CPUQuota -p Slice

Gotcha: a limit only takes effect if the controller is delegated down to the unit's cgroup. systemd enables controllers on demand, so MemoryMax= on a deeply nested unit can read back as infinity if a parent slice never enabled the memory controller. systemctl show -p MemoryMax is the source of truth, not the unit file.

All of this assumes the unified (cgroup v2) hierarchy, the build-time default since systemd 243. systemd 256 refuses to boot under cgroup v1 without SYSTEMD_CGROUP_ENABLE_LEGACY_FORCE=1 on the kernel command line, and 258 removed cgroup v1 support outright. On a v1 box the directive names are the same but the semantics differ — check systemd-analyze --version for default-hierarchy=unified before trusting a v2 answer.

Security Directives

A reasonable starting sandbox for a network service that writes only to its own state directory:

[Service]
DynamicUser=yes
StateDirectory=myapp
NoNewPrivileges=yes
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
AmbientCapabilities=CAP_NET_BIND_SERVICE
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
SystemCallFilter=@system-service
SystemCallFilter=~@privileged
Directive Effect
ProtectSystem=strict The whole filesystem is read-only except /dev, /proc, /sys and anything in ReadWritePaths=
ProtectHome=yes /home, /root and /run/user are empty and inaccessible
PrivateTmp=yes Private /tmp and /var/tmp, discarded on stop
PrivateDevices=yes Minimal /dev — no raw disks, no /dev/kmem
NoNewPrivileges=yes No setuid binary can raise privileges. Free, and almost never breaks anything
CapabilityBoundingSet= The only capabilities the service may ever hold. Empty means none
ProtectKernel*= Block writes to /proc/sys, /sys, and module loading
RestrictAddressFamilies= Socket families the service may create
SystemCallFilter= seccomp allow-list (@set) or deny-list (~@set)

Verify the result rather than assuming it:

systemd-analyze security myapp.service         # Score a loaded unit
systemd-analyze security --offline=true ./myapp.service   # Score a file, no manager needed

Adding sandboxing to a running service breaks it more often than not — usually ProtectSystem=strict without the matching ReadWritePaths=, or PrivateNetwork=yes on something that resolves DNS. Add directives a few at a time and restart between each batch.

For the full sandboxing surface — the systemd-analyze security workflow, credentials, IP and namespace filtering, and the order to apply directives in without breaking the service — see the companion systemd Service Hardening sheet.

Readiness, Watchdogs, and Failure Handling

Restart= brings a dead process back. Readiness protocols and watchdogs decide when it counts as alive and how long a hung one is tolerated.

Key Concepts

  • Readiness: with Type=simple/exec, "started" means "spawned". Only notify, notify-reload, dbus, and forking report actual readiness, which is what After= ordering depends on.
  • Watchdog: the service must ping systemd every WatchdogSec=; miss one and systemd kills and (with Restart=on-watchdog or on-failure) restarts it. Catches hangs that a liveness-free Restart= never sees.
  • Rate limiting: StartLimitBurst= starts inside StartLimitIntervalSec= and systemd stops trying. Both live in [Unit], not [Service].

Readiness Notification

[Service]
Type=notify
# Only the main process may send (default)
NotifyAccess=main
ExecStart=/usr/bin/myapp

The service calls sd_notify(0, "READY=1") once it is actually serving. From Python, either use the systemd-python bindings or write the datagram directly:

import os
import socket

def notify(state: str) -> None:
    """Send a state string to systemd's notification socket.

    No-ops when NOTIFY_SOCKET is unset (i.e. not running under systemd), so the
    same code runs unchanged in a container or a dev shell.
    """
    addr = os.environ.get("NOTIFY_SOCKET")
    if not addr:
        return
    if addr.startswith("@"):            # Abstract namespace socket
        addr = "\0" + addr[1:]
    with socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM) as sock:
        sock.connect(addr)
        sock.sendall(state.encode("utf-8"))

notify("READY=1")                       # After the listener is bound
notify("STATUS=Serving on :8080")       # Shows in `systemctl status`
notify("RELOADING=1\nMONOTONIC_USEC=…")  # Type=notify-reload only

NotifyAccess=all lets any process in the cgroup send state changes, including a child that has no business doing so. Keep it at main unless a supervisor forks the real worker.

Watchdogs

[Service]
Type=notify
ExecStart=/usr/bin/myapp
# Ping deadline
WatchdogSec=30
# Signal on miss (default; gives a core dump)
WatchdogSignal=SIGABRT
# Or on-failure, which also covers watchdog kills
Restart=on-watchdog

systemd exports WATCHDOG_USEC (microseconds) to the service. Ping at half that interval — the deadline is not a suggestion:

# Ping at half the deadline. Keep this a float: WatchdogSec=1 gives
# WATCHDOG_USEC=1000000, and integer division would floor the interval to 0 --
# which reads as "no watchdog configured" and silently stops the pings.
usec = int(os.environ.get("WATCHDOG_USEC", "0"))
interval = usec / 2_000_000 if usec else 0.0   # microseconds -> seconds, halved
if interval:
    notify("WATCHDOG=1")                # Repeat every `interval` seconds

Ping from the loop that does the real work, not a bare timer thread. A watchdog fed by a thread that keeps running while the request loop is wedged reports health it cannot observe.

Restart Strategy

[Unit]
# Window for counting starts
StartLimitIntervalSec=300
# Give up after 5 starts in that window
StartLimitBurst=5
# Run this unit when we enter failed state
OnFailure=notify-admin@%n.service
# 249+; the mirror of OnFailure
OnSuccess=cleanup.service

[Service]
# no | always | on-success | on-failure |
Restart=on-failure
                                        # on-abnormal | on-abort | on-watchdog
# Fixed delay between attempts
RestartSec=5
# 254+: ramp RestartSec -> RestartMaxDelaySec
RestartSteps=5
# over this many steps (exponential backoff)
RestartMaxDelaySec=60
Value Restarts on
no Never (default)
on-success Clean exit only
on-failure Non-zero exit, signal, timeout, or watchdog — the usual choice
on-abnormal Signal, timeout, or watchdog, but not a non-zero exit code
on-abort Uncaught signal only
on-watchdog Watchdog timeout only
always Any exit, clean or not

Footgun: Restart=always plus a config error is an infinite crash loop that fills the journal and, with no StartLimitBurst=, never gives up. Restart=on-failure with a start limit is almost always what you want. When a unit does hit its limit it stays failed until systemctl reset-failed — a restart alone will not clear it.

The OnFailure= handler receives the failed unit's name via %n, so one templated notifier serves every unit:

# /etc/systemd/system/notify-admin@.service
[Unit]
Description=Alert on failure of %I

[Service]
Type=oneshot
ExecStart=/usr/local/bin/alert "unit %I failed on %H"

Dependencies

Control service start order and relationships between units.

Key Concepts

Dependency TypesWeakStrongOrderOrderWantsTargetRequiresAfterBeforeDependency TypesWeakStrongOrderOrderWantsTargetRequiresAfterBefore
  • Requires: Hard dependency; failure stops dependent unit
  • Wants: Soft dependency; failure doesn't affect dependent unit
  • After/Before: Ordering only; no activation dependency
  • BindsTo: Like Requires, but also stops when dependency stops

Dependency Directives

[Unit]
# Ordering (when to start)
# Start after these units
After=network.target postgresql.service
# Start before this unit
Before=httpd.service

# Activation dependencies
# Hard dependency (fail if missing)
Requires=postgresql.service
# Soft dependency (continue if missing)
Wants=redis.service
# Stop when dependency stops
BindsTo=docker.service

# Conflict handling
# Cannot run simultaneously
Conflicts=sendmail.service

# Conditional execution
# Only start if path exists
ConditionPathExists=/etc/myapp/config
# Start if path does NOT exist
ConditionPathExists=!/etc/myapp/disable
# Start if file has content
ConditionFileNotEmpty=/etc/myapp/config
# Fail if path missing
AssertPathExists=/data

Common Dependency Patterns

# Web application with database
[Unit]
Description=Web Application
After=network.target postgresql.service redis.service
Requires=postgresql.service
Wants=redis.service

# Service requiring network
[Unit]
Description=Network Service
After=network-online.target
Wants=network-online.target

# Mount-dependent service
[Unit]
Description=Data Processing Service
After=data.mount
RequiresMountsFor=/data

Dependency Inspection

# Show unit dependencies
systemctl list-dependencies nginx.service

# Show reverse dependencies
systemctl list-dependencies --reverse nginx.service

# Show ordering dependencies
systemctl list-dependencies --after nginx.service
systemctl list-dependencies --before nginx.service

Targets and Boot Process

Targets group units and define system states, replacing traditional runlevels.

Key Concepts

sysinit.targetbasic.targetnetwork.targetnetwork-online.targetmulti-user.targetgraphical.targetrescue.targetemergency.targetsysinit.targetbasic.targetnetwork.targetnetwork-online.targetmulti-user.targetgraphical.targetrescue.targetemergency.target

Common Targets

Target Description Runlevel Equivalent
poweroff.target System shutdown 0
rescue.target Single-user mode 1
multi-user.target Multi-user, no GUI 3
graphical.target Multi-user with GUI 5
reboot.target System reboot 6
emergency.target Emergency shell -

Target Management

# View current target
systemctl get-default

# Set default target
sudo systemctl set-default multi-user.target

# Switch to target immediately
sudo systemctl isolate rescue.target

# List all targets
systemctl list-units --type=target

# Show target dependencies
systemctl list-dependencies graphical.target

Creating Custom Targets

# /etc/systemd/system/myapp.target
[Unit]
Description=My Application Stack
Requires=multi-user.target
After=multi-user.target
AllowIsolate=yes

[Install]
WantedBy=multi-user.target
# Services belonging to custom target
[Install]
WantedBy=myapp.target

Boot Process Overview

# View boot process
systemd-analyze                         # Boot time summary
systemd-analyze blame                   # Time per unit
systemd-analyze critical-chain          # Critical path
systemd-analyze plot > boot.svg         # Visual boot chart

# Boot targets
systemd-analyze verify myservice.service  # Validate unit file

Timer Units for Scheduling

Timer units replace cron for scheduling tasks with better logging and dependency management.

Key Concepts

  • Timers activate corresponding .service units
  • Support both real-time (calendar) and monotonic (relative) timers
  • Persistent timers catch up on missed runs

Timer Types

Timer TypesRealtimeOnCalendarMonotonicOnBootSecOnUnitActiveSecOnStartupSecTimer TypesRealtimeOnCalendarMonotonicOnBootSecOnUnitActiveSecOnStartupSec

Calendar Timer Examples

# /etc/systemd/system/backup.timer
[Unit]
Description=Daily Backup Timer

[Timer]
# Calendar expressions
# Every day at midnight
OnCalendar=daily
# Every Monday at midnight
OnCalendar=weekly
# Every day at 04:00
OnCalendar=*-*-* 04:00:00
# Weekdays at 09:00
OnCalendar=Mon..Fri *-*-* 09:00:00
# First of every month
OnCalendar=*-*-01 00:00:00
# Every 15 minutes
OnCalendar=*:0/15

# Run if missed
Persistent=true
# Random delay up to 5 min
RandomizedDelaySec=300
# Timer accuracy
AccuracySec=1s

[Install]
WantedBy=timers.target

Monotonic Timer Examples

# /etc/systemd/system/cleanup.timer
[Unit]
Description=Periodic Cleanup Timer

[Timer]
# 5 min after boot
OnBootSec=5min
# 1 hour after last activation
OnUnitActiveSec=1h
# 10 min after systemd start
OnStartupSec=10min

[Install]
WantedBy=timers.target

Service for Timer

# /etc/systemd/system/backup.service
[Unit]
Description=Daily Backup Service

[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup.sh
User=backup
Nice=19
IOSchedulingClass=idle

Timer Management

# List active timers
systemctl list-timers
systemctl list-timers --all

# Enable and start timer
sudo systemctl enable backup.timer
sudo systemctl start backup.timer

# Check timer status
systemctl status backup.timer

# Test calendar expression
systemd-analyze calendar "Mon..Fri *-*-* 09:00:00"
systemd-analyze calendar --iterations=5 "daily"

# Manually trigger the service
sudo systemctl start backup.service

Calendar Expression Reference

Expression Description
minutely Every minute
hourly Every hour at :00
daily Every day at 00:00
weekly Every Monday at 00:00
monthly First of month at 00:00
*:0/15 Every 15 minutes
*-*-* 04:00:00 Daily at 04:00
Mon,Wed,Fri *-*-* 10:00:00 Specific days at 10:00
*-*-1,15 00:00:00 1st and 15th of month

Socket Activation

systemd opens the listening socket, then hands it to the service. The service never binds anything itself.

myapp.servicesystemd (PID 1)Clientmyapp.servicesystemd (PID 1)ClientKernel queues the connectionLater connections go straight to the running serviceBind myapp.socket at bootconnect() on port 8080Start service, pass fd 3sd_listen_fds() then accept()Responsemyapp.servicesystemd (PID 1)Clientmyapp.servicesystemd (PID 1)ClientKernel queues the connectionLater connections go straight to the running serviceBind myapp.socket at bootconnect() on port 8080Start service, pass fd 3sd_listen_fds() then accept()Response

Three things this buys you: services start on first use rather than at boot; a restart drops no pending connections, because PID 1 keeps the listening socket and the kernel queues arrivals in its backlog while the service is down; and a service can serve port 443 without ever holding CAP_NET_BIND_SERVICE, because PID 1 did the binding.

That second one is narrower than it sounds. Connections the service has already accept()ed are its own file descriptors and die with it — socket activation is not a substitute for draining established sessions before shutdown. What it removes is the connection-refused window between stop and start.

Socket Unit Structure

# /etc/systemd/system/myapp.socket - name must match myapp.service
[Unit]
Description=Socket for myapp

[Socket]
# TCP; also ListenDatagram=, ListenFIFO=
ListenStream=0.0.0.0:8080
# A unit may listen on several addresses
ListenStream=/run/myapp/api.sock
# Default: one service, all connections
Accept=no
Backlog=1024
# Owner of a filesystem socket
SocketUser=www-data
SocketMode=0660
# Exported in $LISTEN_FDNAMES
FileDescriptorName=api

[Install]
WantedBy=sockets.target
# Enable the socket, NOT the service - the socket pulls the service in
sudo systemctl enable --now myapp.socket
systemctl list-sockets                  # Sockets, their units, and activation state
systemctl status myapp.socket

Accept=no vs Accept=yes

Accept=no (default) Accept=yes
Service instances One, handling every connection One per connection
Unit named myapp.service myapp@.service (template)
Socket reaches the app as fd 3 onwards, via $LISTEN_FDS stdin/stdout, inetd-style
Use for Anything long-lived Rare, low-rate protocols

Accept=yes forks a process per connection. It is the inetd model and it is a denial-of-service vector on a busy port — default to Accept=no.

Picking Up the Socket

systemd passes LISTEN_FDS (how many), LISTEN_PID (which process they are for), and LISTEN_FDNAMES (the FileDescriptorName= labels). Descriptors start at fd 3:

import os
import socket

SD_LISTEN_FDS_START = 3


def listen_fds() -> list[socket.socket]:
    # Return the sockets systemd passed to this process, or [] if there are none.
    # LISTEN_PID guards against a forked child inheriting the variables and
    # claiming descriptors it does not own.
    try:
        if int(os.environ.get("LISTEN_PID", 0)) != os.getpid():
            return []
        count = int(os.environ.get("LISTEN_FDS", 0))
    except ValueError:
        # Malformed environment - treat as "no sockets" rather than crashing.
        return []
    return [socket.socket(fileno=SD_LISTEN_FDS_START + i) for i in range(count)]


socks = listen_fds()
if socks:
    server = socks[0]                             # Already bound and listening
else:
    server = socket.create_server(("0.0.0.0", 8080))  # Standalone fallback

Unset LISTEN_PID/LISTEN_FDS after reading them if you fork, or the child will think the sockets are its own. The C helper sd_listen_fds(1, ...) does this for you; most language bindings expose the same "unset environment" flag.

Drop-in Overrides

Never edit a packaged unit in /usr/lib/systemd/system/ — the next upgrade overwrites it. Layer a drop-in instead.

# Creates /etc/systemd/system/myapp.service.d/override.conf, then daemon-reloads
sudo systemctl edit myapp.service

# Named drop-in, so separate concerns stay in separate files (253+)
sudo systemctl edit --drop-in=hardening.conf myapp.service

# Copy the whole unit into /etc for wholesale replacement
sudo systemctl edit --full myapp.service

# Temporary override in /run, gone at reboot
sudo systemctl edit --runtime myapp.service

# Throw away every local override and go back to the packaged unit
sudo systemctl revert myapp.service

Seeing What Actually Applies

systemctl cat myapp.service              # Merged view, with each source file named
systemctl show myapp.service             # Every resolved property, post-merge
systemctl show myapp.service -p ExecStart -p MemoryMax
systemd-delta --type=extended            # Every drop-in on the system
sudo systemctl daemon-reload             # Required after any hand-edited unit file

Precedence, lowest to highest: /usr/lib/systemd/system/ (packages) -> /run/systemd/system/ (runtime) -> /etc/systemd/system/ (admin). /etc outranks /run — confirm with systemd-analyze unit-paths, which prints the search path highest-priority first. So a systemctl edit --runtime override does not beat a persistent one of the same name in /etc; it beats the packaged unit only. Drop-ins from all three locations are merged, and within a unit's .d/ directory files apply in lexical order, which is why they conventionally start 10-, 20-.

The List-Directive Reset Footgun

Most directives are last-one-wins. ExecStart=, Environment=, ReadWritePaths=, AssertPathExists= and friends are lists — a drop-in appends to them. To replace rather than append, assign the empty value first:

# /etc/systemd/system/myapp.service.d/10-command.conf
[Service]
# Clear the inherited list
ExecStart=
ExecStart=/usr/bin/myapp --config /etc/myapp/prod.yaml

Without the empty ExecStart=, systemd refuses to load the unit at all:

myapp.service: Service has more than one ExecStart= setting,
which is only allowed for Type=oneshot services. Refusing.

Environment= fails more quietly — it accumulates, and the last definition of a given key wins, so a stale value in the base unit can survive an override that looked correct.

Dependencies are the exception, and there is no workaround. After=, Before=, Requires=, Wants=, BindsTo= and the rest cannot be reset: systemd.unit(5) states that "dependencies can only be added in drop-ins. If you want to remove dependencies, you have to override the entire unit." An empty After= in a drop-in is not an error and not a reset — it simply does nothing. Use systemctl edit --full (or systemctl mask plus your own unit) to drop one.

# Validate a unit and its drop-ins together before reloading
systemd-analyze verify /etc/systemd/system/myapp.service
sudo systemctl mask myapp.service        # Symlink to /dev/null: cannot be started at all
sudo systemctl unmask myapp.service

User Services and Lingering

Every logged-in user gets their own systemd --user manager. Units live in ~/.config/systemd/user/ and run as that user, with no root anywhere in the path.

systemctl --user daemon-reload
systemctl --user enable --now myapp.service
systemctl --user status myapp.service
journalctl --user -u myapp.service -f
systemctl --user list-units --type=service
# ~/.config/systemd/user/myapp.service - note the [Install] target
[Unit]
Description=My user service

[Service]
# %h = the user's home directory
ExecStart=%h/.local/bin/myapp
Restart=on-failure

[Install]
# NOT multi-user.target
WantedBy=default.target

Lingering

By default the user manager starts at login and is torn down with the last session — a user service stops when you log out, and never starts at boot.

sudo loginctl enable-linger mike         # Start at boot, survive logout
loginctl show-user mike -p Linger        # Linger=yes
sudo loginctl disable-linger mike

This is the single most common "my user timer never fires" cause. Enabling the unit is not enough; without lingering there is no manager running to fire it.

Differences from system units worth knowing:

  • User= and Group= are meaningless — you are already unprivileged. The namespacing directives (ProtectSystem=, ProtectHome=, PrivateTmp=, PrivateDevices=) also do nothing in a user service, because filesystem namespacing is privileged; add PrivateUsers=yes and most of them start working. Resource control (MemoryMax=, CPUQuota=) works as normal.
  • systemctl --user needs XDG_RUNTIME_DIR and DBUS_SESSION_BUS_ADDRESS. A bare sudo -u mike systemctl --user ... fails with "Failed to connect to bus"; use machinectl shell mike@ or systemd-run --machine=mike@ --user instead.
  • Rootless Podman Quadlets are user units, so the same lingering rule governs whether your containers come back after a reboot.

Transient Units with systemd-run

Run a command under systemd's supervision without writing a unit file. Ideal for one-off jobs that need a resource cap, a timeout, or journal capture.

# One-off supervised command; output goes to the journal
sudo systemd-run --unit=reindex /usr/local/bin/reindex.sh
journalctl -u reindex -f

# Wait for it and get its exit status back
sudo systemd-run --wait --collect --unit=reindex /usr/local/bin/reindex.sh

# Attach stdin/stdout instead of the journal
sudo systemd-run --pipe --wait /usr/bin/du -sh /var/log

# Cap a heavy job so it cannot take the box down
sudo systemd-run --property=MemoryMax=2G --property=CPUQuota=50% \
    --property=IOWeight=10 --unit=backup /usr/local/bin/backup.sh

# Transient timers - no .timer file needed
sudo systemd-run --on-active=90 /usr/bin/systemctl restart nginx
sudo systemd-run --on-calendar='*-*-* 03:00:00' --unit=nightly /usr/local/bin/nightly.sh

# Put an interactive build under a resource cap (scope, not service)
systemd-run --user --scope --property=MemoryMax=4G -- make -j8
Flag Effect
--unit=NAME Name the transient unit instead of taking run-<pid>
--scope Register the calling process's children rather than forking a service
--wait Block until it finishes and propagate the exit status
--collect (-G) Unload the unit afterwards, even on failure
--pipe (-P) Pass stdin/stdout/stderr through instead of journalling
--property=K=V (-p) Any [Service] directive, including the sandboxing ones
--user Run under your own user manager

--scope runs the command in the caller's context — no ExecStart=, no restart, no sandboxing. It is for putting an existing process tree (a build, a shell) into a cgroup. Everything else wants the default service mode.

Trying a hardening directive before committing it to a unit file is the killer use:

sudo systemd-run --pipe --wait -p ProtectSystem=strict -p PrivateTmp=yes \
    -p CapabilityBoundingSet= /usr/local/bin/myapp --self-test

Troubleshooting

Tools and techniques for diagnosing systemd service issues.

Key Concepts

  • journalctl: Centralised logging for all systemd units
  • systemctl status: Quick overview of unit state
  • systemd-analyze: Boot and performance analysis

Checking Service Status

# Basic status
systemctl status nginx.service
systemctl status nginx                  # .service suffix optional

# Detailed status
systemctl show nginx.service            # All properties
systemctl show nginx -p MainPID         # Specific property
systemctl show nginx -p ActiveState,SubState

# Check if active/enabled
systemctl is-active nginx
systemctl is-enabled nginx
systemctl is-failed nginx

Viewing Logs with journalctl

# Service logs
journalctl -u nginx.service             # All logs for unit
journalctl -u nginx -f                  # Follow logs (like tail -f)
journalctl -u nginx --since today       # Today's logs
journalctl -u nginx --since "1 hour ago"
journalctl -u nginx -n 50               # Last 50 lines
journalctl -u nginx -p err              # Errors only

# Priority levels: emerg, alert, crit, err, warning, notice, info, debug
journalctl -u nginx -p warning          # Warning and above

# Output formats
journalctl -u nginx -o json             # JSON output
journalctl -u nginx -o json-pretty      # Pretty JSON
journalctl -u nginx -o cat              # Plain text, no metadata

# Boot-specific logs
journalctl -b                           # Current boot
journalctl -b -1                        # Previous boot
journalctl --list-boots                 # List all boots

# Kernel messages
journalctl -k                           # Kernel messages only
journalctl -k -p err                    # Kernel errors

Finding Failed Units

# List failed units
systemctl --failed
systemctl list-units --state=failed

# Reset failed state
sudo systemctl reset-failed
sudo systemctl reset-failed nginx.service

# Find units in specific states
systemctl list-units --state=inactive
systemctl list-units --state=activating

Debugging Service Start Failures

# Verbose unit start
sudo systemctl start nginx.service --no-block
journalctl -u nginx -f

# Check syntax
systemd-analyze verify /etc/systemd/system/myapp.service

# Test service manually
sudo -u appuser /usr/bin/myapp --config /etc/myapp/config.yaml

# Check dependencies
systemctl list-dependencies nginx.service
systemctl list-dependencies --reverse nginx.service

# Environment issues
systemctl show nginx -p Environment
systemctl cat nginx.service             # Show unit file contents

Common Diagnostic Commands

# Reload configuration after changes
sudo systemctl daemon-reload

# Check disk space for logs
journalctl --disk-usage

# Vacuum old logs
sudo journalctl --vacuum-time=7d        # Keep 7 days
sudo journalctl --vacuum-size=500M      # Keep 500MB

# System state overview
systemctl status                        # Overall system status
systemctl list-units                    # All loaded units
systemctl list-unit-files               # All installed units

# Check resource usage
systemd-cgtop                           # Top for cgroups
systemctl status myapp --no-pager -l    # Full output, no pager

Quick Reference

Essential Commands

Command Description
systemctl start unit Start a unit
systemctl stop unit Stop a unit
systemctl restart unit Restart a unit
systemctl reload unit Reload configuration
systemctl enable unit Enable at boot
systemctl disable unit Disable at boot
systemctl status unit Show unit status
systemctl daemon-reload Reload unit files
systemctl mask unit Prevent unit from starting
systemctl unmask unit Remove mask
systemctl cat unit Merged unit file plus every drop-in
systemctl edit unit Create/edit a drop-in override
systemctl revert unit Discard local overrides
systemctl show unit -p X Resolved value of property X
systemctl reset-failed unit Clear a failed state and its start-limit counter
systemctl list-sockets Socket units and what they activate
systemctl --user ... Operate on your own user manager

Inspection and Analysis

Command Description
systemd-analyze verify unit.service Validate syntax, directive names, and dependencies
systemd-analyze security unit.service Score the sandbox; --offline=true works on a file
systemd-analyze calendar 'Mon *-*-* 02:00' Normalise and show the next elapse of a timer expression
systemd-analyze syscall-filter @mount List the syscalls in a SystemCallFilter= set
systemd-analyze blame Per-unit boot time, slowest first
systemd-analyze critical-chain The time-critical path through boot
systemd-analyze cat-config systemd/journald.conf Merged view of a config file and its drop-ins
systemd-delta Every unit overridden on the system
systemd-escape -p --suffix=mount /mnt/data Compute the unit name for a path
systemd-cgls / systemd-cgtop cgroup tree / live per-cgroup resource usage

Common journalctl Flags

Flag Description
-u unit Show logs for unit
-f Follow (tail) logs
-n N Show last N lines
-p priority Filter by priority
-b Current boot only
--since "time" Logs since time
-o format Output format

Unit File Locations

Location Purpose
/etc/systemd/system/ Local admin units (highest priority)
/run/systemd/system/ Runtime units
/usr/lib/systemd/system/ Package-installed units
~/.config/systemd/user/ User units

Common Issues and Solutions

Service Fails to Start

# Check logs for errors
journalctl -u myapp -n 100 --no-pager

# Verify unit file syntax
systemd-analyze verify /etc/systemd/system/myapp.service

# Check ExecStart path and permissions
ls -la /usr/bin/myapp
sudo -u appuser /usr/bin/myapp --help

# Ensure dependencies are running
systemctl status postgresql.service

Service Keeps Restarting

# Check restart limits
systemctl show myapp -p StartLimitBurst,StartLimitIntervalSec

# View recent failures
journalctl -u myapp --since "10 min ago"

# Temporarily disable restart for debugging
# In [Service]: Restart=no
sudo systemctl daemon-reload
sudo systemctl restart myapp

Timer Not Triggering

# Check timer status
systemctl status backup.timer

# Verify calendar expression
systemd-analyze calendar "*-*-* 04:00:00" --iterations=3

# Ensure timer is enabled and started
sudo systemctl enable --now backup.timer

# Check if service exists
systemctl cat backup.service

Permission Denied Errors

# Check User/Group directives
systemctl show myapp -p User,Group

# Verify file permissions
ls -la /opt/myapp/
ls -la /var/lib/myapp/

# Check SELinux/AppArmor
getenforce                              # SELinux status
ausearch -m AVC -ts recent              # Recent denials
aa-status                               # AppArmor status

Unit Not Found After Creation

# Reload systemd configuration
sudo systemctl daemon-reload

# Check file permissions
ls -la /etc/systemd/system/myapp.service

# Verify unit file name matches
systemctl cat myapp.service

Logs Not Appearing

# Check if journald is running
systemctl status systemd-journald

# Verify logging configuration
cat /etc/systemd/journald.conf

# Check disk space
df -h /var/log/journal

# Restart journald
sudo systemctl restart systemd-journald

Drop-in Override Ignored

# Confirm the drop-in is being read at all, and from where
systemctl cat myapp.service

# Compare the resolved value against what you wrote
systemctl show myapp.service -p ExecStart

# Reload after any hand-edited file
sudo systemctl daemon-reload

Fix: the three usual causes are a missing daemon-reload; a .d directory whose name omits the unit suffix (myapp.d/, not myapp.service.d/); and a file that does not end in .conf, which systemd silently ignores. For list-valued directives, remember the empty-assignment reset.

Unit Fails After Adding Sandboxing

# Which directive actually broke it
journalctl -u myapp.service -b --no-pager | tail -30

# Bisect by running the binary under the suspect directives only
sudo systemd-run --pipe --wait -p ProtectSystem=strict -p PrivateDevices=yes /usr/bin/myapp

Fix: ProtectSystem=strict without a matching ReadWritePaths= is the usual culprit — the service writes somewhere it can no longer reach. PrivateNetwork=yes breaks anything that talks to the network, including DNS. A code=exited, status=203/EXEC means the binary could not be executed at all — often NoExecPaths= or a RootDirectory= that hides it.

User Service Does Not Start at Boot

loginctl show-user "$USER" -p Linger      # Linger=no means no manager at boot
sudo loginctl enable-linger "$USER"
systemctl --user is-enabled myapp.service

Fix: enable lingering. Also check the [Install] section targets default.target — a user unit installed to multi-user.target enables without error and never starts.

Socket Activated Service Never Starts

systemctl list-sockets                    # Is the socket even listening?
systemctl status myapp.socket
ss -lntp | grep 8080                      # Confirm the kernel has the port

Fix: enable the socket, not the service (systemctl enable --now myapp.socket). The unit names must match (myapp.socket → myapp.service) unless the socket sets Service=. If the service starts and immediately exits, it is probably binding its own port instead of picking up $LISTEN_FDS.

Related Topics

  • systemd Service Hardening — the sandboxing directives, systemd-analyze security, and the order to apply them in
  • Linux CLI — the surrounding shell fluency, including journalctl in day-to-day use
  • Linux Storage Management — .mount units, x-systemd.automount, and systemd-escape
  • Linux Performance Analysis — reading the cgroup accounting that MemoryMax= and CPUQuota= drive
  • launchctl — the macOS equivalent, and where the two models genuinely differ
  • Podman Quadlets and systemd — containers declared as units, and the dependency graph Quadlet generates
  • SELinux / AppArmor — the MAC layer that sits underneath systemd's own sandboxing