Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Clair

Running Clair v4 as a service: indexer, matcher, and notifier, the config file, clairctl, updaters, and the Quay integration.

Clair

A vulnerability scanner you operate as a service: Clair indexes container manifests into a Postgres store, then re-matches them against a continuously updated vulnerability database.

Overview

Clair is Red Hat's open-source container scanner and the engine behind Quay's security tab. It is architecturally unlike Trivy or Grype: those are single binaries you run in CI, whereas Clair is a long-lived HTTP service backed by PostgreSQL. You do not "run Clair on an image" — you POST a manifest to it and later ask for a report.

That difference is the whole decision. Clair costs you a database, an updater budget, and an operational surface. In exchange you get continuous re-matching: layers are indexed once, and every subsequent vulnerability-database update re-evaluates every image you have ever submitted, without refetching a byte. For a registry holding tens of thousands of tags, that is the only model that scales.

Clair v4 delegates all analysis to the ClairCore library; Clair itself is the service wrapper — HTTP transport, config, auth, notifications, and the Postgres schema.

v4 is a complete rewrite. Most Clair material online is v2-era and does not apply. If you see clair-scanner, analyze-local-images, quay.io/coreos/clair, a postgres: top-level config key, or clairctl analyze, it is v2 and none of it works against v4. The current image is quay.io/projectquay/clair, the config is the one documented below, and clairctl has an entirely different command set.

Everything in this sheet was checked against Clair v4.9.0 (claircore v1.5.48), the current release.

Clair servicesfetch layersupdatersIndexReportClient (Quay,clairctl, CI)Layer 7 loadbalancer (pathrouting)Indexer/indexer/api/v1Matcher/matcher/api/v1Notifier/notifier/api/v1PostgreSQLContainer registry(layer blobs)Vulnerabilitysources(secdb, OSV, VEX,...)Webhook / AMQP /STOMPClair servicesfetch layersupdatersIndexReportClient (Quay,clairctl, CI)Layer 7 loadbalancer (pathrouting)Indexer/indexer/api/v1Matcher/matcher/api/v1Notifier/notifier/api/v1PostgreSQLContainer registry(layer blobs)Vulnerabilitysources(secdb, OSV, VEX,...)Webhook / AMQP /STOMP

In combo mode all three services run in one OS process and the load balancer disappears; the arrows between them become in-process calls.


The v4 Data Model

Four objects, and the separation between the middle two is the entire point of the design.

Object Produced by Meaning
Manifest The client A digest plus a list of layers, each with a fetchable URI and headers
IndexReport Indexer What is in the image: packages, distribution, repositories, and which layer introduced each
VulnerabilityReport Matcher An IndexReport joined against the current vulnerability database
Notification Notifier "This manifest's affected status changed" — a pointer, not a report

A Clair manifest is not an OCI manifest. Clair never talks to a registry API; the client resolves registry auth and hands Clair pre-authorised blob URLs:

{
  "hash": "sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e",
  "layers": [
    {
      "hash": "sha256:25f1d6b1951ac8eb3740558fe94cb83d377bdadf95fd9f98b50d2e1b96130471",
      "uri": "https://production.cloudfront.docker.com/registry-v2/.../data?Expires=...&Signature=...",
      "headers": {
        "Referer": ["https://index.docker.io/v2/library/alpine/blobs/sha256:25f1d6b1..."]
      }
    }
  ]
}

Both clairctl manifest and Quay's indexing worker produce exactly this shape. Quay additionally sets Accept: application/gzip and a signed download header with a bounded validity.

Why indexing and matching are separate

Manifests and layers are content-addressed, so Clair indexes each layer once across the whole corpus. If four thousand images share a ubuntu:24.04 base, that base is fetched and scanned once. Matching then runs over the stored IndexReport, which means:

  • A new CVE published this morning re-scores every image in the database with no network I/O against the registry.
  • The matcher API can be called as often as you like and always answers against the newest data.
  • Re-indexing is only needed when Clair's own internal scanners change, which clients detect by watching the index_state endpoint's ETag.
UpdatersMatcherRegistryIndexerClientUpdatersMatcherRegistryIndexerClienthours later, independent of any clientPOST /indexer/api/v1/index_report {manifest}GET layer blobs (uri + headers)layer tarballspackage, dist, repo scanners201 Created + IndexReport + ETagnew advisories ingestedGET /matcher/api/v1/vulnerability_report/{digest}200 VulnerabilityReport (fresh, no refetch)UpdatersMatcherRegistryIndexerClientUpdatersMatcherRegistryIndexerClienthours later, independent of any clientPOST /indexer/api/v1/index_report {manifest}GET layer blobs (uri + headers)layer tarballspackage, dist, repo scanners201 Created + IndexReport + ETagnew advisories ingestedGET /matcher/api/v1/vulnerability_report/{digest}200 VulnerabilityReport (fresh, no refetch)

Deployment Modes

The mode is set by the -mode flag or CLAIR_MODE, and is not in the config file. One config file serves every node type; each process reads only the stanzas its mode needs.

Mode Runs Requires
combo Indexer, matcher, and notifier in one process All three config blocks present
indexer Indexer only indexer.connstring
matcher Matcher only matcher.connstring and matcher.indexer_addr
notifier Notifier only notifier.connstring, indexer_addr, and matcher_addr
Distributed modeLayer 7 LB,path-prefix routingindexer podsmatcher podsnotifier podsindexer dbmatcher dbnotifier dbCombo modeOne process: indexer+ matcher + notifierOne databaseDistributed modeLayer 7 LB,path-prefix routingindexer podsmatcher podsnotifier podsindexer dbmatcher dbnotifier dbCombo modeOne process: indexer+ matcher + notifierOne database

Combo is the right default. Reach for distributed only when you need to scale indexing and matching asymmetrically — indexing is I/O- and CPU-bound on layer decompression, matching is database-bound. Distributed mode requires a layer 7 load balancer doing path-prefix routing (/indexer/, /matcher/, /notifier/), because the paths are the only thing distinguishing the services. On Kubernetes that is a Service plus Ingress per role.

The services never share tables even in combo mode, so a single combo process can point each stanza at its own database. Do that when one workload's connection or I/O profile is starving the others.

Database requirements

  • PostgreSQL. Verified here against 17.11; Clair uses pgx v5 and nothing exotic.
  • Set migrations: true on each stanza that owns a database, or Clair will not create its schema.
  • Every service opens its own pool. In combo mode against a single database, three pools plus the advisory-lock connections come out of one max_connections budget — size it accordingly, and prefer pool_max_conns in the connection string over the deprecated matcher.max_conn_pool.

Running Clair

The official image is quay.io/projectquay/clair. latest tracks the development branch — pin a version tag.

# Image facts (podman inspect quay.io/projectquay/clair:4.9.0)
#   User        nobody:nobody
#   Entrypoint  /usr/bin/clair
#   WorkingDir  /run
#   Env         CLAIR_CONF=/config/config.yaml
#               CLAIR_MODE=combo
#               SSL_CERT_DIR=/etc/ssl/certs:/etc/pki/tls/certs:/var/run/certs

Because it runs as nobody, any host path Clair must write (an updater export, for instance) needs to be group- or world-writable, or mounted with a matching user namespace.

# 1. Postgres
podman network create clairnet
podman run -d --name clair-db --network clairnet \
  -e POSTGRES_USER=clair -e POSTGRES_PASSWORD=clair -e POSTGRES_DB=clair \
  docker.io/library/postgres:17

# 2. Clair, combo mode
podman run -d --name clair --network clairnet -p 6060:6060 -p 6061:6061 \
  -v "$PWD:/config:ro,Z" \
  -e CLAIR_MODE=combo -e CLAIR_CONF=/config/config.yaml \
  quay.io/projectquay/clair:4.9.0

# 3. Confirm it came up
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:6061/healthz   # 200
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:6061/readyz    # 200

Startup log, trimmed to the lines that matter:

{"level":"info","component":"main","version":"v4.9.0 (claircore v1.5.48)","message":"ready"}
{"level":"info","component":"introspection/New","endpoint":"/metrics","server":":6061","message":"configuring prometheus"}
{"level":"info","component":"notifier/postgres/Init","message":"performing notifier migrations"}
{"level":"info","component":"libvuln/updates/Manager.Run","total":46,"batchSize":10,"message":"running updaters"}
{"level":"info","component":"libvuln/updates/Manager.Run","retention":10,"message":"GC started"}
{"level":"info","component":"libvuln/updates/Manager.Start","interval":"6h0m0s","message":"starting background updates"}

The total on the running updaters line is the single most useful number in that log: it is how many updaters your sets selection actually resolved to.

Running the binary directly is identical:

clair -conf ./config.yaml -mode combo

The Config File

One YAML (or JSON) document, shared by every node. Below is a realistic combo-mode config with the keys worth knowing annotated.

---
# Where the public API listens. Default ":6060".
http_listen_addr: ":6060"
# Metrics, /healthz and /readyz. Keep this off the public ingress.
introspection_addr: ":6061"
# debug-color | debug | info | warn | error | fatal | panic
log_level: info

indexer:
  connstring: "host=clair-db port=5432 user=clair password=clair dbname=clair sslmode=disable"
  # Poll interval (seconds) for the per-manifest scan lock. Default 1.
  scanlock_retry: 10
  # Layers scanned in parallel per manifest. Small values raise latency.
  layer_scan_concurrency: 2
  # MUST be true on a fresh database or Clair aborts at startup.
  migrations: true
  # 0 auto-sizes from CPU count; negative means unlimited. Excess requests get 429.
  index_report_request_concurrency: 0
  # Block outbound Internet from the layer fetcher. Private ranges still allowed.
  airgap: false

matcher:
  connstring: "host=clair-db port=5432 user=clair password=clair dbname=clair sslmode=disable"
  migrations: true
  # How often updaters run. Default 6h; anything lower is linted as aggressive.
  period: 6h
  # Cache-Control max-age hinted to clients. Defaults to `period`.
  cache_age: 6h
  # Retained update operations per vulnerability database. Default 10;
  # <0 disables GC, <2 silently becomes 10. Notifications need at least 2.
  update_retention: 10
  # Set true when vulnerability data arrives by import instead (air-gap).
  disable_updaters: false
  # Skip CVSS and similar enrichment. Added in 4.9.0.
  disable_enrichment: false
  # Distributed mode only; ignored in combo.
  # indexer_addr: "http://clair-indexer:6060/"

# Restrict which matchers run. nil (omitted) means "all defaults".
# matchers:
#   names: ["alpine-matcher", "debian-matcher", "ubuntu-matcher"]

# Restrict which updater sets run. nil (omitted) means "all defaults".
updaters:
  sets:
    - alpine
    - debian
    - ubuntu
    - rhel-vex
    - osv
    - clair.cvss
  config:
    ubuntu:
      ignore_distributions: ["cosmic"]

notifier:
  connstring: "host=clair-db port=5432 user=clair password=clair dbname=clair sslmode=disable"
  migrations: true
  # Both default to their own minimum: poll 6h, delivery 1h.
  poll_interval: 6h
  delivery_interval: 1h
  # One notification per manifest instead of one per vulnerability.
  disable_summary: false
  webhook:
    target: "https://hooks.internal.example/clair"
    # Trailing slash matters; Clair lints its absence.
    callback: "https://clair.internal.example/notifier/api/v1/notification/"
    headers:
      X-Clair-Env: ["prod"]

auth:
  psk:
    # base64 of the raw HMAC key
    key: "Ya0C5EzF+5V7crTs8LCh+x+PqRn+Ol/Q0yEhLPdP+jQ="
    # Accepted JWT issuers. An empty list accepts any issuer.
    iss: ["clairctl", "quay"]

metrics:
  name: prometheus
# trace: see Observability, below

Defaults worth memorising

Key Default Note
http_listen_addr :6060
indexer.scanlock_retry 1 (second) Name is a historical accident
indexer.index_report_request_concurrency 0 → auto-size Over-limit requests get 429
matcher.period 6h Lint warns below this
matcher.cache_age = matcher.period Emitted as Cache-Control: max-age
matcher.update_retention 10 <0 disables GC; <2 silently becomes 10
notifier.poll_interval 6h Also the minimum; raised from 5m in 4.9.0
notifier.delivery_interval 1h Also the minimum; raised from 1m in 4.9.0
metrics.prometheus.endpoint /metrics On introspection_addr
TMPDIR /var/tmp Where layers are staged during indexing

Values under one minute for the two notifier intervals are silently replaced with the default rather than rejected.

Validating a config

clairctl check-config resolves the file (including drop-ins) and prints the merged result. It does not apply application defaults, so it shows what you wrote, not what Clair will run.

clairctl check-config -o yaml config.yaml    # -o json is the default
WRN some values do no round-trip the yaml encoder correctly -- make sure to consult the documentation
http_listen_addr: :6060
log_level: 0
matcher:
  period: 21600000000000     # durations become nanoseconds; harmless
  update_retention: 0        # pre-defaults, not the effective value

Since 4.7.0 unknown keys are a hard error, which makes the command a genuine typo-catcher:

ERR error="error decoding config \"config.yaml\": unknown field \"conn_string\""

Footgun: the config.yaml.sample shipped in the Clair repository does not load. It still contains matcher.updater_sets (replaced by the top-level updaters.sets) and notifier.amqp.exchange.durable (the key is durability), and either one is now fatal: unknown field "updater_sets". It also sets trace.name: jaeger, deprecated since 4.8.0, and 1m/5m notifier intervals, below the 4.9.0 minimums. Treat it as a sketch, not a template, and put it through check-config before you deploy it.

Drop-ins

From 4.7.0, config.yaml.d/*.yaml (RFC 7386 merge) and config.yaml.d/*.yaml-patch (RFC 6902 patch) are loaded in lexical order after the root file. Extensions must match the root file's format. Merge semantics on lists are replace-not-append, so use a patch document when you mean to append. A final zz-validate.yaml-patch containing test operations makes Clair refuse to start when an expected value is missing — a cheap guard for a templated deployment.

The boot-time linter

Clair lints the config on every start and logs the findings without failing. These are real and useful:

{"lint":"automatically sizing number of concurrent requests (at $.indexer.index_report_request_concurrency)"}
{"lint":"small values will limit resource utilization and increase latency (at $.indexer.layer_scan_concurrency)"}
{"lint":"interval is very fast: may result in increased workload (at $.notifier.delivery_interval)"}
{"lint":"URL should end in a \"/\" (at $.notifier.webhook.callback)"}

Others fire for an aggressive matcher.period ("updater period is very aggressive: most sources are updated daily"), a negative update_retention ("update garbage collection is off" — 0 is silently promoted to 10 and lints nothing), and a set max_conn_pool ("this parameter will be ignored in a future release"). trace.name: jaeger is the exception: the deprecation warning exists in the source, but Trace implements no validate method, so its lint is never reached and 4.9.0 starts silently on a Jaeger exporter.


Updaters

Updaters fetch and parse upstream vulnerability data. By default they run inside the matcher process on matcher.period, and an initial run fires at startup.

updaters.sets selects which run. Omitting the key (or null) runs all defaults; an empty list runs none. Measured against Clair 4.9.0 / claircore 1.5.48, each set expands to this many concrete updaters:

Set Updaters Covers
alpine 46 Alpine secdb, main + community, per release branch
aws 3 Amazon Linux 1/2/2023
clair.cvss 1 Not a source — the CVSS enricher that attaches scores
debian 1 The Debian security tracker JSON
oracle 20 Oracle Linux ELSA, per release
osv 32 OSV.dev — language ecosystems (PyPI, npm, Go, Maven, crates, …)
photon 3 VMware Photon OS
rhel-vex 1 Red Hat VEX (replaces the old OVAL feed)
suse 7 SLES and openSUSE
ubuntu 7 Ubuntu security tracker, per supported release

Footgun: rhel and rhcc are still listed in the upstream config reference but are not registered in current claircore. Naming either yields zero updaters. Worse, an unrecognised set name produces no error and no warning — just {"message":"running updaters","total":0} buried in the log, and later an empty vulnerability database. Use rhel-vex for Red Hat content, and grep the startup log for the total field whenever you change sets.

The matchers are selected separately by matchers.names. The current default set is 13: alpine-matcher, aws-matcher, debian-matcher, gobin, java-maven, oracle, photon, python, rhel, rhel-container-matcher, ruby-gem, suse, ubuntu-matcher. (The upstream reference omits ruby-gem.) Note the asymmetry: rhel is a valid matcher name even though it is no longer a valid updater set name.

Per-updater configuration

updaters.config passes an arbitrary object to a named set's constructor. Names can be generated dynamically, so read the logs to confirm the exact key.

updaters:
  sets: [rhel-vex, ubuntu]
  config:
    ubuntu:
      ignore_distributions: ["cosmic"]

Turning updaters off

matcher:
  disable_updaters: true

Use this when vulnerability data arrives by another route — an offline import, or a dedicated updater node writing to a shared matcher database. It is also the correct setting on every matcher replica but one in a distributed deployment, to stop N pods fetching the same feeds.


Air-Gapped Operation

This is where Clair genuinely differentiates. clairctl can run updaters somewhere with Internet access, serialise the result, and import it into a cluster that has none.

Air-gapped clusterSite transfer policyConnected environmentWorkstation or CIjobVulnerabilitysourcesupdates.json.gzSneakernet / one-way/ reviewInternal web serverclairctlimport-updatersClair matcher dbMatcherAir-gapped clusterSite transfer policyConnected environmentWorkstation or CIjobVulnerabilitysourcesupdates.json.gzSneakernet / one-way/ reviewInternal web serverclairctlimport-updatersClair matcher dbMatcher
# On a connected host. Needs the config (for `updaters`) but NOT the database.
clairctl -c config.yaml export-updaters updates.json.gz

# Move it across the boundary however site policy allows.
scp updates.json.gz internal-webserver:/var/www/

# Inside the cluster. This one DOES need the matcher database.
clairctl -c config.yaml import-updaters http://web.svc/updates.json.gz

Compression is inferred from the extension (.gz, .zst); --gzip/--zstd force it. import-updaters also accepts - for stdin and an HTTP(S) URI directly.

Measured on the alpine set alone (46 updaters):

export-updaters  17s wall, 1.5 MB gzipped, 86 MB uncompressed
import-updaters  2.7s

Imports are idempotent — each record carries the source's fingerprint, and a re-import of unchanged data logs fingerprint match, skipping per updater rather than rewriting rows.

The matcher side must be told to stop reaching out:

matcher:
  disable_updaters: true
indexer:
  airgap: true      # blocks the layer fetcher's egress; private ranges still allowed

indexer.airgap matters because the indexer, not the client, fetches layer blobs. With it set, Clair will only pull from RFC 1918 / ULA addresses — i.e. your internal registry.

Budget for the full default updater set being very much larger than the Alpine figures above. OSV alone is 32 updaters covering every major language ecosystem, and it dominates both the export size and the database.


clairctl

The reference client. It ships inside the Clair image (--entrypoint clairctl) and installs standalone with go install github.com/quay/clair/v4/cmd/clairctl@latest.

clairctl v4.9.0 (claircore v1.5.48)

COMMANDS:
   manifest         print a clair manifest for the named container
   report           request vulnerability reports for the named containers
   export-updaters  run updaters and export results
   import-updaters  import updates
   delete           deletes index reports for given manifest digests
   check-config     print a fully-resolved clair config
   Advanced:
     admin          run administrator task (pre / post / oneoff upgrade tasks)

GLOBAL OPTIONS:
   -D                        print debugging logs
   -q                        quieter log output
   --config value, -c value  clair configuration file (default: "config.yaml") [$CLAIR_CONF]
   --issuer value, --iss     jwt "issuer" for authenticated requests (default: "clairctl")

--config, --issuer and -D are global flags and must precede the subcommand; --host and --out belong to report. Upstream's own getting-started page shows clairctl report --config …, which on 4.9.0 just prints the report usage and exits. check-config takes its files positionally, not via -c. The environment variables CLAIR_CONF and CLAIR_API substitute for -c and --host respectively.

# Print the manifest clairctl would submit — registry auth resolved, blob URLs signed
clairctl manifest docker.io/library/alpine:3.20

# Full round trip: build manifest, index it, fetch and print the report
clairctl -c config.yaml report --host http://clair:6060/ docker.io/library/alpine:3.14.0

# Machine-readable; also accepts xml
clairctl -c config.yaml report -o json --host http://clair:6060/ myregistry.example/app:1.4

# Several images, don't stop on the first failure; skip already-indexed manifests
clairctl -c config.yaml report -k --novel --host http://clair:6060/ app:1.4 app:1.5 app:1.6

# Forget a manifest (frees indexer rows)
clairctl -c config.yaml delete --host http://clair:6060/ sha256:1775bebe...

Real output against a deliberately stale image:

alpine:3.14.0 found libssl1.1    1.1.1k-r0 CVE-2023-0286  (fixed: 1.1.1t-r0)
alpine:3.14.0 found libssl1.1    1.1.1k-r0 CVE-2022-0778  (fixed: 1.1.1n-r0)
alpine:3.14.0 found ssl_client   1.33.1-r2 CVE-2022-28391 (fixed: 1.33.1-r7)
alpine:3.14.0 found zlib         1.2.11-r3 CVE-2018-25032 (fixed: 1.2.12-r0)
alpine:3.14.0 found apk-tools    2.12.5-r1 CVE-2021-36159 (fixed: 2.12.6-r0)

A clean image prints alpine:3.20 ok and exits zero. clairctl report has no severity gate and no failing exit code — it is a reporting client, not a CI gate. If you want a build to fail on a threshold, parse -o json yourself or use Trivy/Grype for that job.

Authentication

clairctl reads auth.psk from the config file and signs a JWT itself; the --issuer value (default clairctl) must appear in the server's auth.psk.iss list. That is the only reason report needs -c at all — drop the config and you get 401s.


The HTTP API

Method Path Purpose
POST /indexer/api/v1/index_report Submit a manifest for indexing
GET /indexer/api/v1/index_report/{digest} Retrieve a stored IndexReport
DELETE /indexer/api/v1/index_report/{digest} Delete one IndexReport
DELETE /indexer/api/v1/index_report Bulk delete (digests in the body)
GET /indexer/api/v1/index_state Indexer configuration state; ETag drives re-indexing
GET /matcher/api/v1/vulnerability_report/{digest} The report you actually want
GET /notifier/api/v1/notification/{id} Paged notification set
DELETE /notifier/api/v1/notification/{id} Acknowledge and free the set
GET /openapi/v1 The OpenAPI document (also authenticated)
GET /healthz, /readyz, /metrics On introspection_addr, not the API port

Paths under /indexer/api/v1/internal/ and /matcher/api/v1/internal/ (affected_manifest, update_diff, update_operation) exist for the notifier's own use. They are not in the OpenAPI spec and any ingress should block them.

A real curl flow

# Mint a PSK JWT (HS256 over the base64-decoded key, iss must be in auth.psk.iss)
TOKEN=$(python3 - "$CLAIR_PSK" <<'EOF'
import base64, hashlib, hmac, json, sys, time
key = base64.b64decode(sys.argv[1])
b64 = lambda b: base64.urlsafe_b64encode(b).rstrip(b"=").decode()
now = int(time.time())
h = b64(json.dumps({"alg":"HS256","typ":"JWT"},separators=(",",":")).encode())
p = b64(json.dumps({"iss":"clairctl","iat":now,"nbf":now-30,"exp":now+3600},separators=(",",":")).encode())
print(f"{h}.{p}." + b64(hmac.new(key, f"{h}.{p}".encode(), hashlib.sha256).digest()))
EOF
)

# 1. What state is the indexer in? The ETag changes when scanners change.
curl -s -H "Authorization: Bearer $TOKEN" \
  http://localhost:6060/indexer/api/v1/index_state
# HTTP/1.1 200 OK
# Content-Type: application/vnd.clair.index_state.v1+json
# Etag: "cb31df8269833698891b35a63251a81a"
# {"state":"cb31df8269833698891b35a63251a81a"}

# 2. Build the manifest and submit it
clairctl manifest docker.io/library/alpine:3.20 > manifest.json
curl -s -D- -X POST -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" -d @manifest.json \
  http://localhost:6060/indexer/api/v1/index_report -o index_report.json
# HTTP/1.1 201 Created
# Etag: "cb31df8269833698891b35a63251a81a"
# Link: </indexer/api/v1/index_report/sha256:c64c687c...>; rel="https://projectquay.io/clair/v1/index_report"
# Link: </matcher/api/v1/vulnerability_report/sha256:c64c687c...>; rel="https://projectquay.io/clair/v1/vulnerability_report"
# Location: /indexer/api/v1/index_report/sha256:c64c687c...

# 3. Ask the matcher
DIGEST=sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e
curl -s -D- -H "Authorization: Bearer $TOKEN" \
  "http://localhost:6060/matcher/api/v1/vulnerability_report/$DIGEST"
# HTTP/1.1 200 OK
# Cache-Control: max-age=21600

Indexing is synchronous: the POST blocks until the layers are fetched and scanned, then returns the finished IndexReport with "state": "IndexFinished". There is no polling loop. The two Link headers tell a client exactly where to go next without string-building, and the Etag is the indexer state to compare against later.

Status codes to handle:

Code Meaning
201 Manifest indexed (returned for a re-submission too — content addressing means it is cheap)
200 Report retrieved
202 The matcher has no vulnerability data yet. Empty body
204 Delete succeeded
400 Malformed manifest
401 Missing, malformed, expired, or wrong-issuer JWT
404 No IndexReport for that digest — index it first
429 index_report_request_concurrency exceeded

Report shapes

An IndexReport, abbreviated:

{
  "manifest_hash": "sha256:c64c687c...", "state": "IndexFinished", "success": true, "err": "",
  "packages": { "26": { "name": "ssl_client", "version": "1.36.1-r31", "kind": "binary",
                        "source": { "name": "busybox", "version": "1.36.1-r31", "kind": "source" },
                        "arch": "x86_64" } },
  "distributions": { "1": { "did": "alpine", "version": "3.20", "pretty_name": "Alpine Linux v3.20" } },
  "repository": {},
  "environments": { "26": [ { "package_db": "lib/apk/db/installed",
                              "introduced_in": "sha256:5843afab...", "distribution_id": "1" } ] }
}

environments is the part people miss: it records which layer introduced each package, which is what lets you tell a base-image problem from an application one.

A VulnerabilityReport adds vulnerabilities (keyed by id) and package_vulnerabilities (package id → list of vulnerability ids):

{
  "id": "97535",
  "updater": "alpine-main-v3.14-updater",
  "name": "CVE-2022-2097",
  "links": "https://security.alpinelinux.org/vuln/CVE-2022-2097",
  "severity": "",
  "normalized_severity": "Unknown",
  "package": { "name": "openssl", "kind": "source" },
  "distribution": { "did": "alpine", "version_id": "3.14" },
  "fixed_in_version": "1.1.1q-r0"
}

normalized_severity is Clair's cross-source scale (Unknown, Negligible, Low, Medium, High, Critical). Alpine's secdb carries no severity at all, so Alpine findings come back Unknown unless the clair.cvss enricher is enabled to attach NVD scores. Do not build a severity gate on Alpine data without it.


Authentication

v4 handles auth itself — the v2-era jwtproxy is gone. The only mechanism is a pre-shared key validating HS256 JWTs.

auth:
  psk:
    key: "Ya0C5EzF+5V7crTs8LCh+x+PqRn+Ol/Q0yEhLPdP+jQ="   # base64 of the raw key
    iss: ["clairctl", "quay"]                              # empty list accepts any issuer
# Generate a key
head -c 32 /dev/urandom | base64

With auth configured, every API path returns 401 without a valid bearer token — including /openapi/v1. The introspection port (/healthz, /readyz, /metrics) is not authenticated, which is the reason to keep it on a separate listener and off any public ingress.

Who presents what:

Caller Issuer Lifetime
clairctl clairctl (override with --issuer) Short-lived, minted per invocation
Quay quay 5 minutes
Clair's own notifier (outbound webhook) clair-notifier 60 seconds

The last one is easy to miss and worth exploiting: Clair signs its own outbound webhook deliveries with the same PSK, so your receiver can verify the call is genuinely from Clair instead of trusting the source address.


Notifier

The notifier watches for update operations that change the affected status of an already-indexed manifest, and tells you about it. A notification is a pointer, not a report — on receipt, go and re-fetch the VulnerabilityReport.

Delivery is by webhook, AMQP, or STOMP; if more than one is configured, preference order is webhook, then AMQP, then STOMP.

AMQP and STOMP delivery are deprecated as of v4.9.0 and will be removed. Upstream's guidance is to write a webhook-to-broker transducer. New deployments should use the webhook. notifier.webhook.signed is deprecated too.

notifier:
  connstring: "..."
  migrations: true
  poll_interval: 6h        # how often to check the matcher for update operations
  delivery_interval: 1h    # how often to retry outstanding deliveries
  disable_summary: false   # false = one notification per manifest, most severe wins
  webhook:
    target: "https://hooks.internal.example/clair"
    callback: "https://clair.internal.example/notifier/api/v1/notification/"
    headers:
      X-Clair-Env: ["prod"]

The POST body is deliberately tiny:

{ "notification_id": "…uuid…", "callback": "https://clair.internal.example/notifier/api/v1/notification/…uuid…" }

An actual delivery, captured from a listener:

POST / HTTP/1.1
Host: hook:8080
User-Agent: clair/v4.9.0 (claircore v1.5.48)
Content-Type: application/json
Authorization: Bearer eyJhbGciOiJIUzI1NiJ9…      # iss "clair-notifier", 60s lifetime
X-Clair-Env: lab

Follow the callback to page through the set:

curl -H "Authorization: Bearer $TOKEN" \
  "https://clair.example/notifier/api/v1/notification/$ID?page_size=1000"
# { "page": { "size": 1000, "next": "…" }, "notifications": [ … ] }

# Repeat with ?next=<page.next> until `next` is absent. page_size must be
# repeated on every request or paging goes wrong.

# Acknowledge — otherwise the notifier holds the rows until its own expiry
curl -X DELETE -H "Authorization: Bearer $TOKEN" \
  "https://clair.example/notifier/api/v1/notification/$ID"

Each notification carries id, manifest, reason (added, removed, changed), and a vulnerability summary with name, severity, fixed_in_version, package, distribution and links. With summarisation on (the default) you get the most severe change per manifest, not one message per CVE.

Broker delivery (deprecated)

For an existing AMQP or STOMP deployment: both take uris (a priority-ordered list), a callback URL, a tls block, and a direct boolean. With direct: false the broker receives the same {notification_id, callback} pointer as the webhook; with direct: true the notifications themselves are published, rollup capping how many ride in one message (0 behaves as 1).

notifier:
  amqp:                                    # AMQP 0.x only — RabbitMQ, not ActiveMQ
    uris: ["amqps://user:pass@broker:5671/vhost"]
    exchange: { name: "clair", type: "direct", durability: true, auto_delete: false }
    routing_key: "notifications"
    direct: true
    rollup: 5

Clair declares nothing on the broker: exchanges, queues, and bindings must exist before the deliverer starts, or every attempt fails. AMQP 1.x brokers were only ever reachable via the STOMP deliverer.

Two operational notes:

  • The notifier is silently disabled when no delivery mechanism is configured. It logs notifier disabled with reason: "no delivery mechanisms configured" at info level and everything else looks healthy.
  • In distributed mode the notifier requires both indexer_addr and matcher_addr; it reaches the internal update_diff and affected_manifest endpoints on those services.
  • NOTIFIER_TEST_MODE=1 (any value) makes the notifier emit synthetic notifications on poll_interval, which is the sane way to develop a receiver.

The Quay Integration

Clair's primary real-world deployment is behind Quay, which polls its own manifest table and drives Clair's API on a worker loop. The registry side — repository notifications, the security tab, robot accounts — is covered in the Container Registries sheet; this is the Clair side of the same wire.

Quay's configuration keys (from quay/quay config.py):

Key Default Meaning
FEATURE_SECURITY_SCANNER false Master switch
FEATURE_SECURITY_NOTIFICATIONS false Enables vulnerability_found repository events
SECURITY_SCANNER_V4_ENDPOINT null e.g. http://clair:6060
SECURITY_SCANNER_V4_PSK null base64 key; must equal Clair's auth.psk.key
SECURITY_SCANNER_INDEXING_INTERVAL 30 Seconds between worker passes
SECURITY_SCANNER_V4_REINDEX_THRESHOLD 300 Minimum seconds before re-indexing a manifest
SECURITY_SCANNER_V4_INDEX_MAX_LAYER_SIZE null e.g. 8G; larger layers are skipped
SECURITY_SCANNER_MAX_SCAN_RETRIES 5 Give up on a manifest after this many failures
SECURITY_SCANNER_V4_MANIFEST_CLEANUP true DELETE index reports for manifests removed from Quay

Quay signs every request HS256 with iss: "quay" and a five-minute expiry, so Clair needs quay in auth.psk.iss. It stores the Etag returned from POST /index_report and compares it against GET /index_state to decide when Clair's internal scanners have changed and a re-index is warranted.

Crucially, Quay hands Clair a signed, time-bounded blob download URL plus headers for each layer, exactly like clairctl manifest does. Clair therefore needs network reachability to wherever Quay's blobs live (its storage backend or CDN), not Quay's registry API, and no registry credentials of its own. For a foreign-layer image Quay passes the upstream URL through verbatim, which is a case where indexer.airgap: true will legitimately block indexing.

Deploying the pair on Kubernetes: run Clair in combo mode with its own Postgres, put both behind cluster-internal Services, and keep introspection_addr off the Route. The Quay Operator does exactly this.


Observability

Metrics are Prometheus-format on the introspection port at /metrics, controlled by metrics.name: prometheus (path overridable with metrics.prometheus.endpoint). Since 4.9.0 OTLP export is also supported for both metrics and traces.

curl -s http://localhost:6061/metrics | grep '^clair_cmd_version_info'
# clair_cmd_version_info{claircore_version="v1.5.48",goversion="go1.24.11",
#   revision="f6a412cc (2025-12-10T10:22:03-08:00)",version="v4.9.0"} 1

Two prefixes matter. clair_http_* covers the transport — per-handler in-flight gauges, request duration and response size, labelled by the exact route (/indexer/api/v1/index_report, /matcher/api/v1/vulnerability_report/:digest, …). claircore_* covers the engine — per-query duration histograms for every database method (claircore_indexer_registerscanners_duration_seconds, and so on), which is where you look when indexing slows down. The exact metric set is explicitly not API and changes between releases; an up-to-date Grafana dashboard lives in contrib/openshift/grafana in the Clair repository.

/healthz and /readyz both return 200 once the HTTP server is up. Note the startup log line no health check configured; unconditionally reporting OK — these are liveness signals for the process, not a statement about database or updater health. Alert on clair_http_* error rates and on updater freshness (query update_operation in the database), not on /healthz.

Tracing is OpenTelemetry:

trace:
  name: otlp             # otlp | sentry | jaeger (deprecated since 4.8.0)
  probability: 0.05
  otlp:
    grpc:
      endpoint: "otel-collector:4317"   # default localhost:4317; http default :4318
      insecure: true

Only one of otlp.http or otlp.grpc should be present. Jaeger's own project has moved to OTLP ingestion, and Clair may refuse name: jaeger in a future release. Do not wait for a warning: upstream's config reference says a Jaeger configuration "will print a warning", but 4.9.0 prints nothing and configures the exporter regardless. The lint exists in config/introspection.go and is unreachable, because Trace implements no validate method for the config walker to call.


Sizing and Operational Notes

Database growth is driven by updaters, not by images. Measured on a fresh Postgres 17 with three indexed manifests, adding one set at a time:

# `alpine` set only (46 updaters) — first run completes in ~15s
database         66 MB
vuln             44 MB      117,899 rows
uo_vuln          12 MB      (update-operation → vulnerability join)
update_operation 96 kB
everything else  < 100 kB   (indexreport, manifest, layer, package, dist, …)

# plus the `debian` set — ONE updater, ~2 minutes, and it more than doubles the store
database        184 MB
vuln                        236,604 rows (118,163 of them debian/updater)

Thirty-two tables in total, and vuln plus uo_vuln are effectively all of it. Both scale with the number of updater sets and with update_retention, and the per-set cost varies by two orders of magnitude — a single Debian updater outweighs all 46 Alpine ones. Enabling the full default set, OSV in particular (32 updaters covering every major language ecosystem), is another step change again. Provision generously, keep autovacuum healthy, and only raise update_retention above the default 10 if you genuinely need to diff far back.

Updater load is bursty and outbound. Every matcher.period the matcher fetches from a dozen upstream hosts (secdb.alpinelinux.org, security-tracker.debian.org, Red Hat's VEX bucket, OSV's GCS bucket, …). Egress firewall rules and any HTTP proxy (HTTPS_PROXY is honoured, as are the other Go standard-library variables) must allow them, or updaters fail quietly and the database goes stale. In a distributed deployment leave updaters enabled on exactly one matcher.

The layer-fetch path is the part that surprises people. Clair — not the client — downloads layer blobs, using the URI and headers in the manifest. Consequences:

  • Presigned URLs expire. A manifest built minutes ago may fail to index; regenerate it rather than retrying.
  • Clair needs egress to the registry's blob storage, which is frequently a different host (and a different firewall rule) from the registry API.
  • Layers are staged on disk under TMPDIR, default /var/tmp. Size it at roughly twice the largest uncompressed layer in your corpus. There is no config key for this; set the environment variable. In a container this must be a writable volume, not the read-only root.
  • layer_scan_concurrency multiplies that disk requirement.

Connection limits. Combo mode against one database opens three pools plus lock connections from a single process. Multiply by replica count before setting Postgres max_connections, and prefer pool_max_conns= in the connection string over matcher.max_conn_pool, which is deprecated.


Clair versus Trivy and Grype

Brief, because the detailed table lives in the Grype and Syft sheet.

Clair Trivy Grype
Shape HTTP service Single binary Single binary
State PostgreSQL Local cache Local cache
Invocation API, driven by a registry or scheduler CLI, in a pipeline CLI, over an SBOM
Re-matching Automatic, continuous, no refetch Re-run the scan Re-match a stored SBOM
CI gating No exit-code gate --exit-code on severity --fail-on severity
Air-gap First-class export/import Offline DB bundle Offline DB bundle

Clair wins in three narrow but real places. Registry-integrated continuous scanning: push once, and every future advisory re-scores the image with no further layer traffic — nothing else does that without you rebuilding the plumbing. Air-gapped operation: export-updaters/import-updaters is a designed workflow with fingerprint-based idempotency, not a tarball of a cache directory. Multi-tenant scale: content-addressed layer dedup across a large corpus, and a queryable database of every finding rather than N JSON files.

Everywhere else — a pipeline gate, a developer laptop, IaC or secret scanning, SBOM generation, filesystem and repository targets — pick Trivy or Grype. Running a Postgres cluster to fail a build is not a trade worth making. Many shops sensibly run both: Clair behind the registry, Trivy in CI.


Quick Reference

Task Command
Start combo clair -conf config.yaml -mode combo
Start via container -e CLAIR_MODE=combo -e CLAIR_CONF=/config/config.yaml
Validate config clairctl check-config -o yaml config.yaml
Build a manifest clairctl manifest <image>
Scan an image clairctl -c config.yaml report --host http://clair:6060/ <image>
JSON output clairctl -c config.yaml report -o json --host … <image>
Forget a manifest clairctl -c config.yaml delete --host … <digest>
Export updates clairctl -c config.yaml export-updaters updates.json.gz
Import updates clairctl -c config.yaml import-updaters updates.json.gz
Health curl :6061/healthz, curl :6061/readyz
Metrics curl :6061/metrics
Generate a PSK head -c 32 /dev/urandom | base64
Setting Value
API port 6060 (http_listen_addr, default :6060)
Introspection port Whatever introspection_addr says; unauthenticated
Image quay.io/projectquay/clair:4.9.0 (runs as nobody)
Config env CLAIR_CONF, CLAIR_MODE; client uses CLAIR_API, CLAIR_CONF
Layer staging TMPDIR, default /var/tmp
Updater sets alpine aws clair.cvss debian oracle osv photon rhel-vex suse ubuntu
Modes combo indexer matcher notifier

Common Issues and Solutions

Issue Cause Solution
GET vulnerability_report returns 202 with an empty body The matcher has no vulnerability data — updaters have not completed a first run Wait for running updaters / GC completed in the log, or import updates. clairctl surfaces this as unexpected return status: 202
Reports come back empty but the API returns 200 An unrecognised updaters.sets entry, e.g. rhel or rhcc Names are matched silently. Check the startup log for "message":"running updaters","total":0 and use rhel-vex
unknown field "updater_sets" at startup Copied config.yaml.sample; the key moved to top-level updaters.sets in 4.7.0, and unknown keys are now fatal Move it, and validate with clairctl check-config before deploying
relation "scanner" does not exist (SQLSTATE 42P01) migrations not enabled on a fresh database Set migrations: true on every stanza that owns a database
Combo mode loops on dial unix /tmp/.s.PGSQL.5432 A required stanza is missing, so its connstring is empty and defaults to a local socket Combo mode needs indexer, matcher, and notifier blocks, even if unused
failed to validate config: matcher mode requires a remote Indexer address Distributed matcher without matcher.indexer_addr Set it; likewise indexer_addr and matcher_addr on the notifier
bad mode "scanner": unknown mode "scanner" Typo in CLAIR_MODE One of combo, indexer, matcher, notifier
Indexing fails with a fetch error on a private registry Clair, not the client, pulls layers; the presigned URL expired or Clair cannot reach blob storage Regenerate the manifest; open egress to the blob host, which is often not the registry API host
Indexing fails only in an air-gapped cluster indexer.airgap: true blocks non-private addresses, including foreign layers Mirror the layer internally, or unset airgap for that path
Indexer fails with no space left on device Layers stage under TMPDIR (/var/tmp), often the read-only container root or a tiny emptyDir Mount a volume and set TMPDIR; budget twice the largest uncompressed layer
429 from the indexer under load index_report_request_concurrency reached Raise it, set -1 for unlimited, or back off in the client
Postgres refusing connections in combo mode Three pools per process plus lock connections, multiplied by replicas Raise max_connections, or set pool_max_conns= in the connstring
Notifier silently does nothing No delivery mechanism configured Look for notifier disabled, reason: "no delivery mechanisms configured"; add a webhook block
Notifier configured but nothing arrives poll_interval/delivery_interval below 1 minute are silently replaced by the 6h/1h defaults Set realistic values; use NOTIFIER_TEST_MODE=1 to force synthetic notifications
All Alpine findings show Unknown severity Alpine's secdb carries no severity data Enable the clair.cvss updater set to attach NVD scores
Every image looks vulnerable after an upgrade Clair's internal scanners changed, so old IndexReports are stale Compare index_state ETags and re-index; Quay does this automatically
Updaters stop working behind a proxy Go's standard variables are honoured but must be set Set HTTPS_PROXY / SSL_CERT_DIR in the environment, not the config file
401 on every request including /openapi/v1 auth.psk is set and the token's iss is not in auth.psk.iss Add the issuer (clairctl, quay), or pass --issuer

Related Topics

The following topics complement this cheatsheet and would be valuable additions:

  1. Trivy - The CLI counterpart: severity gating, SBOM and VEX, IaC and secret scanning, and the CI ergonomics Clair deliberately lacks
  2. Grype and Syft - SBOM-first scanning and the detailed comparison table across all three scanners
  3. Container Registries - Quay and Harbor from the registry side: repository notifications, robot accounts, and the security tab Clair feeds
  4. PostgreSQL - Connection pooling, autovacuum, and sizing for the store that dominates Clair's operational cost
  5. Container Security - Runtime hardening, rootless isolation, and the supply-chain controls that scanning alone cannot provide
  6. Prometheus - Recording rules and alerting for the clair_http_* and claircore_* series, and the updater-freshness alert worth writing