Clair
Running Clair v4 as a service: indexer, matcher, and notifier, the config file, clairctl, updaters, and the Quay integration.
Clair
A vulnerability scanner you operate as a service: Clair indexes container manifests into a Postgres store, then re-matches them against a continuously updated vulnerability database.
Overview
Clair is Red Hat's open-source container scanner and the engine behind Quay's security tab. It is architecturally unlike Trivy or Grype: those are single binaries you run in CI, whereas Clair is a long-lived HTTP service backed by PostgreSQL. You do not "run Clair on an image" — you POST a manifest to it and later ask for a report.
That difference is the whole decision. Clair costs you a database, an updater budget, and an operational surface. In exchange you get continuous re-matching: layers are indexed once, and every subsequent vulnerability-database update re-evaluates every image you have ever submitted, without refetching a byte. For a registry holding tens of thousands of tags, that is the only model that scales.
Clair v4 delegates all analysis to the ClairCore library; Clair itself is the service wrapper — HTTP transport, config, auth, notifications, and the Postgres schema.
v4 is a complete rewrite. Most Clair material online is v2-era and does not apply. If you see
clair-scanner,analyze-local-images,quay.io/coreos/clair, apostgres:top-level config key, orclairctl analyze, it is v2 and none of it works against v4. The current image isquay.io/projectquay/clair, the config is the one documented below, andclairctlhas an entirely different command set.
Everything in this sheet was checked against Clair v4.9.0 (claircore v1.5.48), the current release.
graph TB
Client["Client (Quay, clairctl, CI)"]
LB["Layer 7 load balancer (path routing)"]
subgraph Clair["Clair services"]
I["Indexer<br/>/indexer/api/v1"]
M["Matcher<br/>/matcher/api/v1"]
N["Notifier<br/>/notifier/api/v1"]
end
DB[("PostgreSQL")]
Reg["Container registry<br/>(layer blobs)"]
Vuln["Vulnerability sources<br/>(secdb, OSV, VEX, ...)"]
Hook["Webhook / AMQP / STOMP"]
Client --> LB
LB --> I
LB --> M
LB --> N
I -->|"fetch layers"| Reg
I --> DB
M --> DB
N --> DB
M -->|"updaters"| Vuln
M -->|"IndexReport"| I
N --> M
N --> Hook
In combo mode all three services run in one OS process and the load balancer disappears; the arrows between them become in-process calls.
The v4 Data Model
Four objects, and the separation between the middle two is the entire point of the design.
| Object | Produced by | Meaning |
|---|---|---|
| Manifest | The client | A digest plus a list of layers, each with a fetchable URI and headers |
| IndexReport | Indexer | What is in the image: packages, distribution, repositories, and which layer introduced each |
| VulnerabilityReport | Matcher | An IndexReport joined against the current vulnerability database |
| Notification | Notifier | "This manifest's affected status changed" — a pointer, not a report |
A Clair manifest is not an OCI manifest. Clair never talks to a registry API; the client resolves registry auth and hands Clair pre-authorised blob URLs:
{
"hash": "sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e",
"layers": [
{
"hash": "sha256:25f1d6b1951ac8eb3740558fe94cb83d377bdadf95fd9f98b50d2e1b96130471",
"uri": "https://production.cloudfront.docker.com/registry-v2/.../data?Expires=...&Signature=...",
"headers": {
"Referer": ["https://index.docker.io/v2/library/alpine/blobs/sha256:25f1d6b1..."]
}
}
]
}
Both clairctl manifest and Quay's indexing worker produce exactly this shape. Quay additionally sets Accept: application/gzip and a signed download header with a bounded validity.
Why indexing and matching are separate
Manifests and layers are content-addressed, so Clair indexes each layer once across the whole corpus. If four thousand images share a ubuntu:24.04 base, that base is fetched and scanned once. Matching then runs over the stored IndexReport, which means:
- A new CVE published this morning re-scores every image in the database with no network I/O against the registry.
- The matcher API can be called as often as you like and always answers against the newest data.
- Re-indexing is only needed when Clair's own internal scanners change, which clients detect by watching the
index_stateendpoint's ETag.
sequenceDiagram
participant C as Client
participant I as Indexer
participant R as Registry
participant M as Matcher
participant U as Updaters
C->>I: POST /indexer/api/v1/index_report {manifest}
I->>R: GET layer blobs (uri + headers)
R-->>I: layer tarballs
I->>I: package, dist, repo scanners
I-->>C: 201 Created + IndexReport + ETag
Note over U,M: hours later, independent of any client
U->>M: new advisories ingested
C->>M: GET /matcher/api/v1/vulnerability_report/{digest}
M-->>C: 200 VulnerabilityReport (fresh, no refetch)
Deployment Modes
The mode is set by the -mode flag or CLAIR_MODE, and is not in the config file. One config file serves every node type; each process reads only the stanzas its mode needs.
| Mode | Runs | Requires |
|---|---|---|
combo |
Indexer, matcher, and notifier in one process | All three config blocks present |
indexer |
Indexer only | indexer.connstring |
matcher |
Matcher only | matcher.connstring and matcher.indexer_addr |
notifier |
Notifier only | notifier.connstring, indexer_addr, and matcher_addr |
flowchart TB
subgraph Combo["Combo mode"]
P["One process: indexer + matcher + notifier"]
P --> CDB[("One database")]
end
subgraph Dist["Distributed mode"]
LB2["Layer 7 LB, path-prefix routing"]
LB2 --> IX["indexer pods"]
LB2 --> MX["matcher pods"]
LB2 --> NX["notifier pods"]
IX --> DB4[("indexer db")]
MX --> DB5[("matcher db")]
NX --> DB6[("notifier db")]
end
Combo is the right default. Reach for distributed only when you need to scale indexing and matching asymmetrically — indexing is I/O- and CPU-bound on layer decompression, matching is database-bound. Distributed mode requires a layer 7 load balancer doing path-prefix routing (/indexer/, /matcher/, /notifier/), because the paths are the only thing distinguishing the services. On Kubernetes that is a Service plus Ingress per role.
The services never share tables even in combo mode, so a single combo process can point each stanza at its own database. Do that when one workload's connection or I/O profile is starving the others.
Database requirements
- PostgreSQL. Verified here against 17.11; Clair uses
pgxv5 and nothing exotic. - Set
migrations: trueon each stanza that owns a database, or Clair will not create its schema. - Every service opens its own pool. In combo mode against a single database, three pools plus the advisory-lock connections come out of one
max_connectionsbudget — size it accordingly, and preferpool_max_connsin the connection string over the deprecatedmatcher.max_conn_pool.
Running Clair
The official image is quay.io/projectquay/clair. latest tracks the development branch — pin a version tag.
# Image facts (podman inspect quay.io/projectquay/clair:4.9.0)
# User nobody:nobody
# Entrypoint /usr/bin/clair
# WorkingDir /run
# Env CLAIR_CONF=/config/config.yaml
# CLAIR_MODE=combo
# SSL_CERT_DIR=/etc/ssl/certs:/etc/pki/tls/certs:/var/run/certs
Because it runs as nobody, any host path Clair must write (an updater export, for instance) needs to be group- or world-writable, or mounted with a matching user namespace.
# 1. Postgres
podman network create clairnet
podman run -d --name clair-db --network clairnet \
-e POSTGRES_USER=clair -e POSTGRES_PASSWORD=clair -e POSTGRES_DB=clair \
docker.io/library/postgres:17
# 2. Clair, combo mode
podman run -d --name clair --network clairnet -p 6060:6060 -p 6061:6061 \
-v "$PWD:/config:ro,Z" \
-e CLAIR_MODE=combo -e CLAIR_CONF=/config/config.yaml \
quay.io/projectquay/clair:4.9.0
# 3. Confirm it came up
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:6061/healthz # 200
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:6061/readyz # 200
Startup log, trimmed to the lines that matter:
{"level":"info","component":"main","version":"v4.9.0 (claircore v1.5.48)","message":"ready"}
{"level":"info","component":"introspection/New","endpoint":"/metrics","server":":6061","message":"configuring prometheus"}
{"level":"info","component":"notifier/postgres/Init","message":"performing notifier migrations"}
{"level":"info","component":"libvuln/updates/Manager.Run","total":46,"batchSize":10,"message":"running updaters"}
{"level":"info","component":"libvuln/updates/Manager.Run","retention":10,"message":"GC started"}
{"level":"info","component":"libvuln/updates/Manager.Start","interval":"6h0m0s","message":"starting background updates"}
The total on the running updaters line is the single most useful number in that log: it is how many updaters your sets selection actually resolved to.
Running the binary directly is identical:
clair -conf ./config.yaml -mode combo
The Config File
One YAML (or JSON) document, shared by every node. Below is a realistic combo-mode config with the keys worth knowing annotated.
---
# Where the public API listens. Default ":6060".
http_listen_addr: ":6060"
# Metrics, /healthz and /readyz. Keep this off the public ingress.
introspection_addr: ":6061"
# debug-color | debug | info | warn | error | fatal | panic
log_level: info
indexer:
connstring: "host=clair-db port=5432 user=clair password=clair dbname=clair sslmode=disable"
# Poll interval (seconds) for the per-manifest scan lock. Default 1.
scanlock_retry: 10
# Layers scanned in parallel per manifest. Small values raise latency.
layer_scan_concurrency: 2
# MUST be true on a fresh database or Clair aborts at startup.
migrations: true
# 0 auto-sizes from CPU count; negative means unlimited. Excess requests get 429.
index_report_request_concurrency: 0
# Block outbound Internet from the layer fetcher. Private ranges still allowed.
airgap: false
matcher:
connstring: "host=clair-db port=5432 user=clair password=clair dbname=clair sslmode=disable"
migrations: true
# How often updaters run. Default 6h; anything lower is linted as aggressive.
period: 6h
# Cache-Control max-age hinted to clients. Defaults to `period`.
cache_age: 6h
# Retained update operations per vulnerability database. Default 10;
# <0 disables GC, <2 silently becomes 10. Notifications need at least 2.
update_retention: 10
# Set true when vulnerability data arrives by import instead (air-gap).
disable_updaters: false
# Skip CVSS and similar enrichment. Added in 4.9.0.
disable_enrichment: false
# Distributed mode only; ignored in combo.
# indexer_addr: "http://clair-indexer:6060/"
# Restrict which matchers run. nil (omitted) means "all defaults".
# matchers:
# names: ["alpine-matcher", "debian-matcher", "ubuntu-matcher"]
# Restrict which updater sets run. nil (omitted) means "all defaults".
updaters:
sets:
- alpine
- debian
- ubuntu
- rhel-vex
- osv
- clair.cvss
config:
ubuntu:
ignore_distributions: ["cosmic"]
notifier:
connstring: "host=clair-db port=5432 user=clair password=clair dbname=clair sslmode=disable"
migrations: true
# Both default to their own minimum: poll 6h, delivery 1h.
poll_interval: 6h
delivery_interval: 1h
# One notification per manifest instead of one per vulnerability.
disable_summary: false
webhook:
target: "https://hooks.internal.example/clair"
# Trailing slash matters; Clair lints its absence.
callback: "https://clair.internal.example/notifier/api/v1/notification/"
headers:
X-Clair-Env: ["prod"]
auth:
psk:
# base64 of the raw HMAC key
key: "Ya0C5EzF+5V7crTs8LCh+x+PqRn+Ol/Q0yEhLPdP+jQ="
# Accepted JWT issuers. An empty list accepts any issuer.
iss: ["clairctl", "quay"]
metrics:
name: prometheus
# trace: see Observability, below
Defaults worth memorising
| Key | Default | Note |
|---|---|---|
http_listen_addr |
:6060 |
|
indexer.scanlock_retry |
1 (second) |
Name is a historical accident |
indexer.index_report_request_concurrency |
0 → auto-size |
Over-limit requests get 429 |
matcher.period |
6h |
Lint warns below this |
matcher.cache_age |
= matcher.period |
Emitted as Cache-Control: max-age |
matcher.update_retention |
10 |
<0 disables GC; <2 silently becomes 10 |
notifier.poll_interval |
6h |
Also the minimum; raised from 5m in 4.9.0 |
notifier.delivery_interval |
1h |
Also the minimum; raised from 1m in 4.9.0 |
metrics.prometheus.endpoint |
/metrics |
On introspection_addr |
TMPDIR |
/var/tmp |
Where layers are staged during indexing |
Values under one minute for the two notifier intervals are silently replaced with the default rather than rejected.
Validating a config
clairctl check-config resolves the file (including drop-ins) and prints the merged result. It does not apply application defaults, so it shows what you wrote, not what Clair will run.
clairctl check-config -o yaml config.yaml # -o json is the default
WRN some values do no round-trip the yaml encoder correctly -- make sure to consult the documentation
http_listen_addr: :6060
log_level: 0
matcher:
period: 21600000000000 # durations become nanoseconds; harmless
update_retention: 0 # pre-defaults, not the effective value
Since 4.7.0 unknown keys are a hard error, which makes the command a genuine typo-catcher:
ERR error="error decoding config \"config.yaml\": unknown field \"conn_string\""
Footgun: the
config.yaml.sampleshipped in the Clair repository does not load. It still containsmatcher.updater_sets(replaced by the top-levelupdaters.sets) andnotifier.amqp.exchange.durable(the key isdurability), and either one is now fatal:unknown field "updater_sets". It also setstrace.name: jaeger, deprecated since 4.8.0, and 1m/5m notifier intervals, below the 4.9.0 minimums. Treat it as a sketch, not a template, and put it throughcheck-configbefore you deploy it.
Drop-ins
From 4.7.0, config.yaml.d/*.yaml (RFC 7386 merge) and config.yaml.d/*.yaml-patch (RFC 6902 patch) are loaded in lexical order after the root file. Extensions must match the root file's format. Merge semantics on lists are replace-not-append, so use a patch document when you mean to append. A final zz-validate.yaml-patch containing test operations makes Clair refuse to start when an expected value is missing — a cheap guard for a templated deployment.
The boot-time linter
Clair lints the config on every start and logs the findings without failing. These are real and useful:
{"lint":"automatically sizing number of concurrent requests (at $.indexer.index_report_request_concurrency)"}
{"lint":"small values will limit resource utilization and increase latency (at $.indexer.layer_scan_concurrency)"}
{"lint":"interval is very fast: may result in increased workload (at $.notifier.delivery_interval)"}
{"lint":"URL should end in a \"/\" (at $.notifier.webhook.callback)"}
Others fire for an aggressive matcher.period ("updater period is very aggressive: most sources are updated daily"), a negative update_retention ("update garbage collection is off" — 0 is silently promoted to 10 and lints nothing), and a set max_conn_pool ("this parameter will be ignored in a future release"). trace.name: jaeger is the exception: the deprecation warning exists in the source, but Trace implements no validate method, so its lint is never reached and 4.9.0 starts silently on a Jaeger exporter.
Updaters
Updaters fetch and parse upstream vulnerability data. By default they run inside the matcher process on matcher.period, and an initial run fires at startup.
updaters.sets selects which run. Omitting the key (or null) runs all defaults; an empty list runs none. Measured against Clair 4.9.0 / claircore 1.5.48, each set expands to this many concrete updaters:
| Set | Updaters | Covers |
|---|---|---|
alpine |
46 | Alpine secdb, main + community, per release branch |
aws |
3 | Amazon Linux 1/2/2023 |
clair.cvss |
1 | Not a source — the CVSS enricher that attaches scores |
debian |
1 | The Debian security tracker JSON |
oracle |
20 | Oracle Linux ELSA, per release |
osv |
32 | OSV.dev — language ecosystems (PyPI, npm, Go, Maven, crates, …) |
photon |
3 | VMware Photon OS |
rhel-vex |
1 | Red Hat VEX (replaces the old OVAL feed) |
suse |
7 | SLES and openSUSE |
ubuntu |
7 | Ubuntu security tracker, per supported release |
Footgun:
rhelandrhccare still listed in the upstream config reference but are not registered in current claircore. Naming either yields zero updaters. Worse, an unrecognised set name produces no error and no warning — just{"message":"running updaters","total":0}buried in the log, and later an empty vulnerability database. Userhel-vexfor Red Hat content, and grep the startup log for thetotalfield whenever you changesets.
The matchers are selected separately by matchers.names. The current default set is 13: alpine-matcher, aws-matcher, debian-matcher, gobin, java-maven, oracle, photon, python, rhel, rhel-container-matcher, ruby-gem, suse, ubuntu-matcher. (The upstream reference omits ruby-gem.) Note the asymmetry: rhel is a valid matcher name even though it is no longer a valid updater set name.
Per-updater configuration
updaters.config passes an arbitrary object to a named set's constructor. Names can be generated dynamically, so read the logs to confirm the exact key.
updaters:
sets: [rhel-vex, ubuntu]
config:
ubuntu:
ignore_distributions: ["cosmic"]
Turning updaters off
matcher:
disable_updaters: true
Use this when vulnerability data arrives by another route — an offline import, or a dedicated updater node writing to a shared matcher database. It is also the correct setting on every matcher replica but one in a distributed deployment, to stop N pods fetching the same feeds.
Air-Gapped Operation
This is where Clair genuinely differentiates. clairctl can run updaters somewhere with Internet access, serialise the result, and import it into a cluster that has none.
flowchart LR
subgraph Online["Connected environment"]
WS["Workstation or CI job"]
SRC["Vulnerability sources"]
SRC --> WS
WS --> F["updates.json.gz"]
end
subgraph Transfer["Site transfer policy"]
F --> X["Sneakernet / one-way / review"]
end
subgraph Offline["Air-gapped cluster"]
X --> W["Internal web server"]
W --> CTL["clairctl import-updaters"]
CTL --> DB[("Clair matcher db")]
DB --> MM["Matcher"]
end
# On a connected host. Needs the config (for `updaters`) but NOT the database.
clairctl -c config.yaml export-updaters updates.json.gz
# Move it across the boundary however site policy allows.
scp updates.json.gz internal-webserver:/var/www/
# Inside the cluster. This one DOES need the matcher database.
clairctl -c config.yaml import-updaters http://web.svc/updates.json.gz
Compression is inferred from the extension (.gz, .zst); --gzip/--zstd force it. import-updaters also accepts - for stdin and an HTTP(S) URI directly.
Measured on the alpine set alone (46 updaters):
export-updaters 17s wall, 1.5 MB gzipped, 86 MB uncompressed
import-updaters 2.7s
Imports are idempotent — each record carries the source's fingerprint, and a re-import of unchanged data logs fingerprint match, skipping per updater rather than rewriting rows.
The matcher side must be told to stop reaching out:
matcher:
disable_updaters: true
indexer:
airgap: true # blocks the layer fetcher's egress; private ranges still allowed
indexer.airgap matters because the indexer, not the client, fetches layer blobs. With it set, Clair will only pull from RFC 1918 / ULA addresses — i.e. your internal registry.
Budget for the full default updater set being very much larger than the Alpine figures above. OSV alone is 32 updaters covering every major language ecosystem, and it dominates both the export size and the database.
clairctl
The reference client. It ships inside the Clair image (--entrypoint clairctl) and installs standalone with go install github.com/quay/clair/v4/cmd/clairctl@latest.
clairctl v4.9.0 (claircore v1.5.48)
COMMANDS:
manifest print a clair manifest for the named container
report request vulnerability reports for the named containers
export-updaters run updaters and export results
import-updaters import updates
delete deletes index reports for given manifest digests
check-config print a fully-resolved clair config
Advanced:
admin run administrator task (pre / post / oneoff upgrade tasks)
GLOBAL OPTIONS:
-D print debugging logs
-q quieter log output
--config value, -c value clair configuration file (default: "config.yaml") [$CLAIR_CONF]
--issuer value, --iss jwt "issuer" for authenticated requests (default: "clairctl")
--config, --issuer and -D are global flags and must precede the subcommand; --host and --out belong to report. Upstream's own getting-started page shows clairctl report --config …, which on 4.9.0 just prints the report usage and exits. check-config takes its files positionally, not via -c. The environment variables CLAIR_CONF and CLAIR_API substitute for -c and --host respectively.
# Print the manifest clairctl would submit — registry auth resolved, blob URLs signed
clairctl manifest docker.io/library/alpine:3.20
# Full round trip: build manifest, index it, fetch and print the report
clairctl -c config.yaml report --host http://clair:6060/ docker.io/library/alpine:3.14.0
# Machine-readable; also accepts xml
clairctl -c config.yaml report -o json --host http://clair:6060/ myregistry.example/app:1.4
# Several images, don't stop on the first failure; skip already-indexed manifests
clairctl -c config.yaml report -k --novel --host http://clair:6060/ app:1.4 app:1.5 app:1.6
# Forget a manifest (frees indexer rows)
clairctl -c config.yaml delete --host http://clair:6060/ sha256:1775bebe...
Real output against a deliberately stale image:
alpine:3.14.0 found libssl1.1 1.1.1k-r0 CVE-2023-0286 (fixed: 1.1.1t-r0)
alpine:3.14.0 found libssl1.1 1.1.1k-r0 CVE-2022-0778 (fixed: 1.1.1n-r0)
alpine:3.14.0 found ssl_client 1.33.1-r2 CVE-2022-28391 (fixed: 1.33.1-r7)
alpine:3.14.0 found zlib 1.2.11-r3 CVE-2018-25032 (fixed: 1.2.12-r0)
alpine:3.14.0 found apk-tools 2.12.5-r1 CVE-2021-36159 (fixed: 2.12.6-r0)
A clean image prints alpine:3.20 ok and exits zero. clairctl report has no severity gate and no failing exit code — it is a reporting client, not a CI gate. If you want a build to fail on a threshold, parse -o json yourself or use Trivy/Grype for that job.
Authentication
clairctl reads auth.psk from the config file and signs a JWT itself; the --issuer value (default clairctl) must appear in the server's auth.psk.iss list. That is the only reason report needs -c at all — drop the config and you get 401s.
The HTTP API
| Method | Path | Purpose |
|---|---|---|
POST |
/indexer/api/v1/index_report |
Submit a manifest for indexing |
GET |
/indexer/api/v1/index_report/{digest} |
Retrieve a stored IndexReport |
DELETE |
/indexer/api/v1/index_report/{digest} |
Delete one IndexReport |
DELETE |
/indexer/api/v1/index_report |
Bulk delete (digests in the body) |
GET |
/indexer/api/v1/index_state |
Indexer configuration state; ETag drives re-indexing |
GET |
/matcher/api/v1/vulnerability_report/{digest} |
The report you actually want |
GET |
/notifier/api/v1/notification/{id} |
Paged notification set |
DELETE |
/notifier/api/v1/notification/{id} |
Acknowledge and free the set |
GET |
/openapi/v1 |
The OpenAPI document (also authenticated) |
GET |
/healthz, /readyz, /metrics |
On introspection_addr, not the API port |
Paths under /indexer/api/v1/internal/ and /matcher/api/v1/internal/ (affected_manifest, update_diff, update_operation) exist for the notifier's own use. They are not in the OpenAPI spec and any ingress should block them.
A real curl flow
# Mint a PSK JWT (HS256 over the base64-decoded key, iss must be in auth.psk.iss)
TOKEN=$(python3 - "$CLAIR_PSK" <<'EOF'
import base64, hashlib, hmac, json, sys, time
key = base64.b64decode(sys.argv[1])
b64 = lambda b: base64.urlsafe_b64encode(b).rstrip(b"=").decode()
now = int(time.time())
h = b64(json.dumps({"alg":"HS256","typ":"JWT"},separators=(",",":")).encode())
p = b64(json.dumps({"iss":"clairctl","iat":now,"nbf":now-30,"exp":now+3600},separators=(",",":")).encode())
print(f"{h}.{p}." + b64(hmac.new(key, f"{h}.{p}".encode(), hashlib.sha256).digest()))
EOF
)
# 1. What state is the indexer in? The ETag changes when scanners change.
curl -s -H "Authorization: Bearer $TOKEN" \
http://localhost:6060/indexer/api/v1/index_state
# HTTP/1.1 200 OK
# Content-Type: application/vnd.clair.index_state.v1+json
# Etag: "cb31df8269833698891b35a63251a81a"
# {"state":"cb31df8269833698891b35a63251a81a"}
# 2. Build the manifest and submit it
clairctl manifest docker.io/library/alpine:3.20 > manifest.json
curl -s -D- -X POST -H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" -d @manifest.json \
http://localhost:6060/indexer/api/v1/index_report -o index_report.json
# HTTP/1.1 201 Created
# Etag: "cb31df8269833698891b35a63251a81a"
# Link: </indexer/api/v1/index_report/sha256:c64c687c...>; rel="https://projectquay.io/clair/v1/index_report"
# Link: </matcher/api/v1/vulnerability_report/sha256:c64c687c...>; rel="https://projectquay.io/clair/v1/vulnerability_report"
# Location: /indexer/api/v1/index_report/sha256:c64c687c...
# 3. Ask the matcher
DIGEST=sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e
curl -s -D- -H "Authorization: Bearer $TOKEN" \
"http://localhost:6060/matcher/api/v1/vulnerability_report/$DIGEST"
# HTTP/1.1 200 OK
# Cache-Control: max-age=21600
Indexing is synchronous: the POST blocks until the layers are fetched and scanned, then returns the finished IndexReport with "state": "IndexFinished". There is no polling loop. The two Link headers tell a client exactly where to go next without string-building, and the Etag is the indexer state to compare against later.
Status codes to handle:
| Code | Meaning |
|---|---|
201 |
Manifest indexed (returned for a re-submission too — content addressing means it is cheap) |
200 |
Report retrieved |
202 |
The matcher has no vulnerability data yet. Empty body |
204 |
Delete succeeded |
400 |
Malformed manifest |
401 |
Missing, malformed, expired, or wrong-issuer JWT |
404 |
No IndexReport for that digest — index it first |
429 |
index_report_request_concurrency exceeded |
Report shapes
An IndexReport, abbreviated:
{
"manifest_hash": "sha256:c64c687c...", "state": "IndexFinished", "success": true, "err": "",
"packages": { "26": { "name": "ssl_client", "version": "1.36.1-r31", "kind": "binary",
"source": { "name": "busybox", "version": "1.36.1-r31", "kind": "source" },
"arch": "x86_64" } },
"distributions": { "1": { "did": "alpine", "version": "3.20", "pretty_name": "Alpine Linux v3.20" } },
"repository": {},
"environments": { "26": [ { "package_db": "lib/apk/db/installed",
"introduced_in": "sha256:5843afab...", "distribution_id": "1" } ] }
}
environments is the part people miss: it records which layer introduced each package, which is what lets you tell a base-image problem from an application one.
A VulnerabilityReport adds vulnerabilities (keyed by id) and package_vulnerabilities (package id → list of vulnerability ids):
{
"id": "97535",
"updater": "alpine-main-v3.14-updater",
"name": "CVE-2022-2097",
"links": "https://security.alpinelinux.org/vuln/CVE-2022-2097",
"severity": "",
"normalized_severity": "Unknown",
"package": { "name": "openssl", "kind": "source" },
"distribution": { "did": "alpine", "version_id": "3.14" },
"fixed_in_version": "1.1.1q-r0"
}
normalized_severity is Clair's cross-source scale (Unknown, Negligible, Low, Medium, High, Critical). Alpine's secdb carries no severity at all, so Alpine findings come back Unknown unless the clair.cvss enricher is enabled to attach NVD scores. Do not build a severity gate on Alpine data without it.
Authentication
v4 handles auth itself — the v2-era jwtproxy is gone. The only mechanism is a pre-shared key validating HS256 JWTs.
auth:
psk:
key: "Ya0C5EzF+5V7crTs8LCh+x+PqRn+Ol/Q0yEhLPdP+jQ=" # base64 of the raw key
iss: ["clairctl", "quay"] # empty list accepts any issuer
# Generate a key
head -c 32 /dev/urandom | base64
With auth configured, every API path returns 401 without a valid bearer token — including /openapi/v1. The introspection port (/healthz, /readyz, /metrics) is not authenticated, which is the reason to keep it on a separate listener and off any public ingress.
Who presents what:
| Caller | Issuer | Lifetime |
|---|---|---|
clairctl |
clairctl (override with --issuer) |
Short-lived, minted per invocation |
| Quay | quay |
5 minutes |
| Clair's own notifier (outbound webhook) | clair-notifier |
60 seconds |
The last one is easy to miss and worth exploiting: Clair signs its own outbound webhook deliveries with the same PSK, so your receiver can verify the call is genuinely from Clair instead of trusting the source address.
Notifier
The notifier watches for update operations that change the affected status of an already-indexed manifest, and tells you about it. A notification is a pointer, not a report — on receipt, go and re-fetch the VulnerabilityReport.
Delivery is by webhook, AMQP, or STOMP; if more than one is configured, preference order is webhook, then AMQP, then STOMP.
AMQP and STOMP delivery are deprecated as of v4.9.0 and will be removed. Upstream's guidance is to write a webhook-to-broker transducer. New deployments should use the webhook.
notifier.webhook.signedis deprecated too.
notifier:
connstring: "..."
migrations: true
poll_interval: 6h # how often to check the matcher for update operations
delivery_interval: 1h # how often to retry outstanding deliveries
disable_summary: false # false = one notification per manifest, most severe wins
webhook:
target: "https://hooks.internal.example/clair"
callback: "https://clair.internal.example/notifier/api/v1/notification/"
headers:
X-Clair-Env: ["prod"]
The POST body is deliberately tiny:
{ "notification_id": "…uuid…", "callback": "https://clair.internal.example/notifier/api/v1/notification/…uuid…" }
An actual delivery, captured from a listener:
POST / HTTP/1.1
Host: hook:8080
User-Agent: clair/v4.9.0 (claircore v1.5.48)
Content-Type: application/json
Authorization: Bearer eyJhbGciOiJIUzI1NiJ9… # iss "clair-notifier", 60s lifetime
X-Clair-Env: lab
Follow the callback to page through the set:
curl -H "Authorization: Bearer $TOKEN" \
"https://clair.example/notifier/api/v1/notification/$ID?page_size=1000"
# { "page": { "size": 1000, "next": "…" }, "notifications": [ … ] }
# Repeat with ?next=<page.next> until `next` is absent. page_size must be
# repeated on every request or paging goes wrong.
# Acknowledge — otherwise the notifier holds the rows until its own expiry
curl -X DELETE -H "Authorization: Bearer $TOKEN" \
"https://clair.example/notifier/api/v1/notification/$ID"
Each notification carries id, manifest, reason (added, removed, changed), and a vulnerability summary with name, severity, fixed_in_version, package, distribution and links. With summarisation on (the default) you get the most severe change per manifest, not one message per CVE.
Broker delivery (deprecated)
For an existing AMQP or STOMP deployment: both take uris (a priority-ordered list), a callback URL, a tls block, and a direct boolean. With direct: false the broker receives the same {notification_id, callback} pointer as the webhook; with direct: true the notifications themselves are published, rollup capping how many ride in one message (0 behaves as 1).
notifier:
amqp: # AMQP 0.x only — RabbitMQ, not ActiveMQ
uris: ["amqps://user:pass@broker:5671/vhost"]
exchange: { name: "clair", type: "direct", durability: true, auto_delete: false }
routing_key: "notifications"
direct: true
rollup: 5
Clair declares nothing on the broker: exchanges, queues, and bindings must exist before the deliverer starts, or every attempt fails. AMQP 1.x brokers were only ever reachable via the STOMP deliverer.
Two operational notes:
- The notifier is silently disabled when no delivery mechanism is configured. It logs
notifier disabledwithreason: "no delivery mechanisms configured"at info level and everything else looks healthy. - In distributed mode the notifier requires both
indexer_addrandmatcher_addr; it reaches the internalupdate_diffandaffected_manifestendpoints on those services. NOTIFIER_TEST_MODE=1(any value) makes the notifier emit synthetic notifications onpoll_interval, which is the sane way to develop a receiver.
The Quay Integration
Clair's primary real-world deployment is behind Quay, which polls its own manifest table and drives Clair's API on a worker loop. The registry side — repository notifications, the security tab, robot accounts — is covered in the Container Registries sheet; this is the Clair side of the same wire.
Quay's configuration keys (from quay/quay config.py):
| Key | Default | Meaning |
|---|---|---|
FEATURE_SECURITY_SCANNER |
false |
Master switch |
FEATURE_SECURITY_NOTIFICATIONS |
false |
Enables vulnerability_found repository events |
SECURITY_SCANNER_V4_ENDPOINT |
null |
e.g. http://clair:6060 |
SECURITY_SCANNER_V4_PSK |
null |
base64 key; must equal Clair's auth.psk.key |
SECURITY_SCANNER_INDEXING_INTERVAL |
30 |
Seconds between worker passes |
SECURITY_SCANNER_V4_REINDEX_THRESHOLD |
300 |
Minimum seconds before re-indexing a manifest |
SECURITY_SCANNER_V4_INDEX_MAX_LAYER_SIZE |
null |
e.g. 8G; larger layers are skipped |
SECURITY_SCANNER_MAX_SCAN_RETRIES |
5 |
Give up on a manifest after this many failures |
SECURITY_SCANNER_V4_MANIFEST_CLEANUP |
true |
DELETE index reports for manifests removed from Quay |
Quay signs every request HS256 with iss: "quay" and a five-minute expiry, so Clair needs quay in auth.psk.iss. It stores the Etag returned from POST /index_report and compares it against GET /index_state to decide when Clair's internal scanners have changed and a re-index is warranted.
Crucially, Quay hands Clair a signed, time-bounded blob download URL plus headers for each layer, exactly like clairctl manifest does. Clair therefore needs network reachability to wherever Quay's blobs live (its storage backend or CDN), not Quay's registry API, and no registry credentials of its own. For a foreign-layer image Quay passes the upstream URL through verbatim, which is a case where indexer.airgap: true will legitimately block indexing.
Deploying the pair on Kubernetes: run Clair in combo mode with its own Postgres, put both behind cluster-internal Services, and keep introspection_addr off the Route. The Quay Operator does exactly this.
Observability
Metrics are Prometheus-format on the introspection port at /metrics, controlled by metrics.name: prometheus (path overridable with metrics.prometheus.endpoint). Since 4.9.0 OTLP export is also supported for both metrics and traces.
curl -s http://localhost:6061/metrics | grep '^clair_cmd_version_info'
# clair_cmd_version_info{claircore_version="v1.5.48",goversion="go1.24.11",
# revision="f6a412cc (2025-12-10T10:22:03-08:00)",version="v4.9.0"} 1
Two prefixes matter. clair_http_* covers the transport — per-handler in-flight gauges, request duration and response size, labelled by the exact route (/indexer/api/v1/index_report, /matcher/api/v1/vulnerability_report/:digest, …). claircore_* covers the engine — per-query duration histograms for every database method (claircore_indexer_registerscanners_duration_seconds, and so on), which is where you look when indexing slows down. The exact metric set is explicitly not API and changes between releases; an up-to-date Grafana dashboard lives in contrib/openshift/grafana in the Clair repository.
/healthz and /readyz both return 200 once the HTTP server is up. Note the startup log line no health check configured; unconditionally reporting OK — these are liveness signals for the process, not a statement about database or updater health. Alert on clair_http_* error rates and on updater freshness (query update_operation in the database), not on /healthz.
Tracing is OpenTelemetry:
trace:
name: otlp # otlp | sentry | jaeger (deprecated since 4.8.0)
probability: 0.05
otlp:
grpc:
endpoint: "otel-collector:4317" # default localhost:4317; http default :4318
insecure: true
Only one of otlp.http or otlp.grpc should be present. Jaeger's own project has moved to OTLP ingestion, and Clair may refuse name: jaeger in a future release. Do not wait for a warning: upstream's config reference says a Jaeger configuration "will print a warning", but 4.9.0 prints nothing and configures the exporter regardless. The lint exists in config/introspection.go and is unreachable, because Trace implements no validate method for the config walker to call.
Sizing and Operational Notes
Database growth is driven by updaters, not by images. Measured on a fresh Postgres 17 with three indexed manifests, adding one set at a time:
# `alpine` set only (46 updaters) — first run completes in ~15s
database 66 MB
vuln 44 MB 117,899 rows
uo_vuln 12 MB (update-operation → vulnerability join)
update_operation 96 kB
everything else < 100 kB (indexreport, manifest, layer, package, dist, …)
# plus the `debian` set — ONE updater, ~2 minutes, and it more than doubles the store
database 184 MB
vuln 236,604 rows (118,163 of them debian/updater)
Thirty-two tables in total, and vuln plus uo_vuln are effectively all of it. Both scale with the number of updater sets and with update_retention, and the per-set cost varies by two orders of magnitude — a single Debian updater outweighs all 46 Alpine ones. Enabling the full default set, OSV in particular (32 updaters covering every major language ecosystem), is another step change again. Provision generously, keep autovacuum healthy, and only raise update_retention above the default 10 if you genuinely need to diff far back.
Updater load is bursty and outbound. Every matcher.period the matcher fetches from a dozen upstream hosts (secdb.alpinelinux.org, security-tracker.debian.org, Red Hat's VEX bucket, OSV's GCS bucket, …). Egress firewall rules and any HTTP proxy (HTTPS_PROXY is honoured, as are the other Go standard-library variables) must allow them, or updaters fail quietly and the database goes stale. In a distributed deployment leave updaters enabled on exactly one matcher.
The layer-fetch path is the part that surprises people. Clair — not the client — downloads layer blobs, using the URI and headers in the manifest. Consequences:
- Presigned URLs expire. A manifest built minutes ago may fail to index; regenerate it rather than retrying.
- Clair needs egress to the registry's blob storage, which is frequently a different host (and a different firewall rule) from the registry API.
- Layers are staged on disk under
TMPDIR, default/var/tmp. Size it at roughly twice the largest uncompressed layer in your corpus. There is no config key for this; set the environment variable. In a container this must be a writable volume, not the read-only root. layer_scan_concurrencymultiplies that disk requirement.
Connection limits. Combo mode against one database opens three pools plus lock connections from a single process. Multiply by replica count before setting Postgres max_connections, and prefer pool_max_conns= in the connection string over matcher.max_conn_pool, which is deprecated.
Clair versus Trivy and Grype
Brief, because the detailed table lives in the Grype and Syft sheet.
| Clair | Trivy | Grype | |
|---|---|---|---|
| Shape | HTTP service | Single binary | Single binary |
| State | PostgreSQL | Local cache | Local cache |
| Invocation | API, driven by a registry or scheduler | CLI, in a pipeline | CLI, over an SBOM |
| Re-matching | Automatic, continuous, no refetch | Re-run the scan | Re-match a stored SBOM |
| CI gating | No exit-code gate | --exit-code on severity |
--fail-on severity |
| Air-gap | First-class export/import | Offline DB bundle | Offline DB bundle |
Clair wins in three narrow but real places. Registry-integrated continuous scanning: push once, and every future advisory re-scores the image with no further layer traffic — nothing else does that without you rebuilding the plumbing. Air-gapped operation: export-updaters/import-updaters is a designed workflow with fingerprint-based idempotency, not a tarball of a cache directory. Multi-tenant scale: content-addressed layer dedup across a large corpus, and a queryable database of every finding rather than N JSON files.
Everywhere else — a pipeline gate, a developer laptop, IaC or secret scanning, SBOM generation, filesystem and repository targets — pick Trivy or Grype. Running a Postgres cluster to fail a build is not a trade worth making. Many shops sensibly run both: Clair behind the registry, Trivy in CI.
Quick Reference
| Task | Command |
|---|---|
| Start combo | clair -conf config.yaml -mode combo |
| Start via container | -e CLAIR_MODE=combo -e CLAIR_CONF=/config/config.yaml |
| Validate config | clairctl check-config -o yaml config.yaml |
| Build a manifest | clairctl manifest <image> |
| Scan an image | clairctl -c config.yaml report --host http://clair:6060/ <image> |
| JSON output | clairctl -c config.yaml report -o json --host … <image> |
| Forget a manifest | clairctl -c config.yaml delete --host … <digest> |
| Export updates | clairctl -c config.yaml export-updaters updates.json.gz |
| Import updates | clairctl -c config.yaml import-updaters updates.json.gz |
| Health | curl :6061/healthz, curl :6061/readyz |
| Metrics | curl :6061/metrics |
| Generate a PSK | head -c 32 /dev/urandom | base64 |
| Setting | Value |
|---|---|
| API port | 6060 (http_listen_addr, default :6060) |
| Introspection port | Whatever introspection_addr says; unauthenticated |
| Image | quay.io/projectquay/clair:4.9.0 (runs as nobody) |
| Config env | CLAIR_CONF, CLAIR_MODE; client uses CLAIR_API, CLAIR_CONF |
| Layer staging | TMPDIR, default /var/tmp |
| Updater sets | alpine aws clair.cvss debian oracle osv photon rhel-vex suse ubuntu |
| Modes | combo indexer matcher notifier |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
GET vulnerability_report returns 202 with an empty body |
The matcher has no vulnerability data — updaters have not completed a first run | Wait for running updaters / GC completed in the log, or import updates. clairctl surfaces this as unexpected return status: 202 |
Reports come back empty but the API returns 200 |
An unrecognised updaters.sets entry, e.g. rhel or rhcc |
Names are matched silently. Check the startup log for "message":"running updaters","total":0 and use rhel-vex |
unknown field "updater_sets" at startup |
Copied config.yaml.sample; the key moved to top-level updaters.sets in 4.7.0, and unknown keys are now fatal |
Move it, and validate with clairctl check-config before deploying |
relation "scanner" does not exist (SQLSTATE 42P01) |
migrations not enabled on a fresh database |
Set migrations: true on every stanza that owns a database |
Combo mode loops on dial unix /tmp/.s.PGSQL.5432 |
A required stanza is missing, so its connstring is empty and defaults to a local socket | Combo mode needs indexer, matcher, and notifier blocks, even if unused |
failed to validate config: matcher mode requires a remote Indexer address |
Distributed matcher without matcher.indexer_addr |
Set it; likewise indexer_addr and matcher_addr on the notifier |
bad mode "scanner": unknown mode "scanner" |
Typo in CLAIR_MODE |
One of combo, indexer, matcher, notifier |
| Indexing fails with a fetch error on a private registry | Clair, not the client, pulls layers; the presigned URL expired or Clair cannot reach blob storage | Regenerate the manifest; open egress to the blob host, which is often not the registry API host |
| Indexing fails only in an air-gapped cluster | indexer.airgap: true blocks non-private addresses, including foreign layers |
Mirror the layer internally, or unset airgap for that path |
| Indexer fails with no space left on device | Layers stage under TMPDIR (/var/tmp), often the read-only container root or a tiny emptyDir |
Mount a volume and set TMPDIR; budget twice the largest uncompressed layer |
429 from the indexer under load |
index_report_request_concurrency reached |
Raise it, set -1 for unlimited, or back off in the client |
| Postgres refusing connections in combo mode | Three pools per process plus lock connections, multiplied by replicas | Raise max_connections, or set pool_max_conns= in the connstring |
| Notifier silently does nothing | No delivery mechanism configured | Look for notifier disabled, reason: "no delivery mechanisms configured"; add a webhook block |
| Notifier configured but nothing arrives | poll_interval/delivery_interval below 1 minute are silently replaced by the 6h/1h defaults |
Set realistic values; use NOTIFIER_TEST_MODE=1 to force synthetic notifications |
All Alpine findings show Unknown severity |
Alpine's secdb carries no severity data | Enable the clair.cvss updater set to attach NVD scores |
| Every image looks vulnerable after an upgrade | Clair's internal scanners changed, so old IndexReports are stale | Compare index_state ETags and re-index; Quay does this automatically |
| Updaters stop working behind a proxy | Go's standard variables are honoured but must be set | Set HTTPS_PROXY / SSL_CERT_DIR in the environment, not the config file |
401 on every request including /openapi/v1 |
auth.psk is set and the token's iss is not in auth.psk.iss |
Add the issuer (clairctl, quay), or pass --issuer |
Related Topics
The following topics complement this cheatsheet and would be valuable additions:
- Trivy - The CLI counterpart: severity gating, SBOM and VEX, IaC and secret scanning, and the CI ergonomics Clair deliberately lacks
- Grype and Syft - SBOM-first scanning and the detailed comparison table across all three scanners
- Container Registries - Quay and Harbor from the registry side: repository notifications, robot accounts, and the security tab Clair feeds
- PostgreSQL - Connection pooling, autovacuum, and sizing for the store that dominates Clair's operational cost
- Container Security - Runtime hardening, rootless isolation, and the supply-chain controls that scanning alone cannot provide
- Prometheus - Recording rules and alerting for the
clair_http_*andclaircore_*series, and the updater-freshness alert worth writing