Container Image Building
Choosing and driving an image builder: BuildKit/buildx, Buildah, Kaniko, Stacker, apko/melange, and ko, with rootless, multi-arch, and reproducibility trade-offs.
Container Image Building
Choosing an image builder and driving it: BuildKit/buildx, Buildah, Kaniko, Stacker, apko/melange, and ko.
Overview
An OCI image is a JSON config, a set of layer blobs, and a manifest that ties them together. Any tool that can emit those three things is an image builder, and the six here differ almost entirely in how they decide what goes into the layers and what privileges they need to do it.
This sheet is about the builders. Dockerfile technique — multi-stage builds, layer ordering, cache mounts, base image selection, .dockerignore, size and hardening work — is covered in depth in the Container Image Optimisation sheet and is not repeated here. Dockerfiles appear below only as far as they demonstrate a builder's behaviour.
A seventh model sits deliberately outside this comparison. Cloud Native Buildpacks inverts the question: you supply a source tree and no build definition at all, and a chain of buildpacks detects the language, builds it, and assembles the image against a base image the platform owns rather than you. That ownership split is the whole point — it is what makes pack rebase able to patch an OS CVE without rebuilding the application — and it makes buildpacks incomparable row-by-row with the tools here. It has its own Cloud Native Buildpacks sheet; the decision tree below shows where it displaces everything on this page.
flowchart LR
DF["Dockerfile"] --> BK["BuildKit + buildx<br/>daemon or standalone, root"]
DF --> BH["Buildah<br/>daemonless, rootless"]
DF --> KN["Kaniko (archived)<br/>daemonless, needs a container"]
SH["Shell script"] --> BH
YAML["stacker.yaml"] --> ST["Stacker<br/>daemonless, unprivileged"]
APY["apko.yaml + melange"] --> AP["apko<br/>daemonless, no root"]
GO["Go source"] --> KO["ko<br/>no build file at all"]
SRC["Any source tree"] --> CNB["Cloud Native Buildpacks<br/>lifecycle detects and builds"]
BK --> OCI["OCI image:<br/>config + layers + manifest"]
BH --> OCI
KN --> OCI
ST --> OCI
AP --> OCI
KO --> OCI
CNB --> OCI
The privilege question is the one that decides most real arguments:
flowchart TB
A["Do you need a Dockerfile?"] -->|No, Go binary| KO["ko: no daemon, no root, cross-compiles"]
A -->|"No, and no build file at all"| CNB["Cloud Native Buildpacks: pack build, platform owns the base image"]
A -->|"No, declarative"| B["Reproducibility the hard requirement?"]
A -->|Yes| C["Is a Docker daemon available?"]
B -->|Yes| AP["apko + melange: bitwise reproducible, APK-based"]
B -->|"No, want layer control"| ST["Stacker: YAML, unprivileged, OCI-layout out"]
C -->|Yes| BX["buildx: default driver, or docker-container for full features"]
C -->|"No, but rootless Linux"| BH["Buildah: rootless, daemonless, scripted or Dockerfile"]
C -->|"No, unprivileged K8s pod"| K8["buildkitd rootless pod + buildctl"]
KN["Kaniko"] -.->|"archived June 2025"| K8
Version baseline for this sheet
| Component | Version verified against |
|---|---|
| Docker Engine / CLI | 29.7.2 (containerd image store, overlayfs snapshotter) |
| buildx | v0.36.1 local; v0.37.0 current upstream |
| BuildKit | v0.32.2 local; v0.33.0 current upstream |
| Dockerfile frontend | docker/dockerfile:1 → 1.27.0 |
| Buildah | 1.39.3 on Debian 13 (trixie); 1.45.0 current upstream |
| Podman | 5.4.2 |
| Kaniko | v1.24.0 — final release; repo archived 2025-06-03 |
| Stacker | v1.2.0 binary (v1.2.1 tagged 2026-08-25) |
| apko / melange | v1.2.43 / v0.59.4 |
| ko | v0.19.1 |
| cosign | v3.1.3 |
Choosing a Builder
This is the table the rest of the sheet elaborates. It is opinionated on purpose.
| buildx / BuildKit | Buildah | Kaniko | Stacker | apko + melange | ko | |
|---|---|---|---|---|---|---|
| Input | Dockerfile (+ pluggable frontends) | Dockerfile or shell script driving buildah verbs |
Dockerfile | stacker.yaml |
apko.yaml (image) + melange YAML (packages) |
Go source tree, no build file |
| Daemon | Yes for the docker driver; docker-container/kubernetes/remote run buildkitd themselves |
None | None | None | None | None |
| Root | Daemon runs as root; rootless buildkitd supported |
No | No, but wants to be inside a container | No | No | No |
| Rootless story | Good but fiddly: moby/buildkit:*-rootless image, --oci-worker-no-process-sandbox, unconfined seccomp/AppArmor |
Best in class. Native user namespaces, buildah unshare, works out of the box given subuid/subgid |
Poor. Unpacks the base image into its own filesystem; needs an isolated container to be safe | Good. User namespaces, stacker unpriv-setup for subuid/subgid |
Excellent. apko needs nothing; melange's bubblewrap runner needs unprivileged userns | Excellent. Just a Go toolchain |
| Cache | Best in class: content-addressed DAG, inline/registry/local/gha/s3/azblob exporters, RUN --mount=type=cache |
--layers (off by default), --cache-from/--cache-to to a registry, --cache-ttl |
--cache + --cache-repo; stops reading cache after the first miss |
Content-hashed layers in .stacker/, incremental rebuild on input change |
APK resolution cache; apko lock pins the package set |
Go build cache; base image pulled once |
| Multi-arch | --platform with QEMU emulation, or native builder nodes appended to one builder; imagetools create after the fact |
--platform (needs host binfmt), --manifest for lists; no built-in emulation |
None. --custom-platform only relabels; stitch with an external tool |
Per-arch builds, no built-in manifest assembly | Native. archs: list, no emulation — it is just unpacking APKs |
Native. Go cross-compiles; no emulation ever |
| Reproducibility | Achievable: SOURCE_DATE_EPOCH + rewrite-timestamp=true. Not the default |
--timestamp; --source-date-epoch/--rewrite-timestamp from 1.41 |
--reproducible flag, patchy in practice |
Content-hashed layers, hash: pinning on imports, --require-hash |
Bitwise reproducible by design — verified below | Yes, with SOURCE_DATE_EPOCH — verified below |
| SBOM / provenance | SLSA provenance on by default (mode=min); --sbom=true uses the BuildKit Syft scanner |
--sbom preset scanners (1.35+), pluggable scanner image |
None | Embeds the whole stacker.yaml as an OCI annotation; bom: key |
Per-package SBOM inside each APK, composed into an image SBOM | SPDX by default (--sbom=spdx) |
| Maturity | The reference implementation. Actively developed, huge surface | Mature, stable, Red Hat-backed; Podman's build engine | Archived. Do not start new work on it | Real project, tiny community (~370 stars); latest tag shipped no binaries | Very active, Chainguard-backed; the model is unusual and the ecosystem is Wolfi-shaped | Active, CNCF Sandbox; deliberately narrow |
Short version. If you have a Dockerfile and a daemon, use buildx. If you need rootless and daemonless on Linux, use Buildah. If you need reproducibility as a contractual property, use apko/melange. If it is a Go binary, use ko and stop thinking about it. Do not adopt Kaniko in 2026. Stacker is a good tool with a very small blast radius of users — evaluate it only if its layer-control and unprivileged story is what you specifically need.
Where buildpacks displace all of this. Every tool in the table above assumes you want to own the base image and the build definition. If you do not — if you would rather hand over a source tree and have someone else's curated base image, language detection, and CVE patching applied for you — then Cloud Native Buildpacks is a different answer to the question, not a seventh column in it. The trade is control for operational leverage: you give up precise say over the final image and gain pack rebase, which swaps the OS layer under an unchanged application layer in seconds. See the Cloud Native Buildpacks sheet.
BuildKit and buildx
BuildKit is the build engine; docker buildx is the CLI in front of it. BuildKit has been the default since Docker 23, and docker build is now an alias for docker buildx build (along with docker builder build and docker image build). The classic builder has not been removed — DOCKER_BUILDKIT=0 docker build . still falls back to it on Engine 29.7.2, complete with Successfully built <id> output. Nothing in this sheet works there.
Builder Instances and Drivers
A builder is a named handle onto one or more BuildKit nodes. The driver decides where buildkitd lives.
| Driver | Where BuildKit runs | Notes |
|---|---|---|
docker |
Inside the Docker daemon | Default. Cannot be created manually; one per Docker context |
docker-container |
A container the driver manages | The workhorse. Full feature set, isolated cache |
cloud |
Docker Build Cloud | Managed, paid |
kubernetes |
Pods in a cluster | Deployment or StatefulSet; scales, supports native multi-arch nodes |
remote |
A buildkitd you run yourself |
tcp:// or unix://, TLS via driver-opts |
docker buildx ls
# NAME/NODE DRIVER/ENDPOINT STATUS BUILDKIT PLATFORMS
# desktop-linux* docker
# \_ desktop-linux \_ desktop-linux running v0.32.2 linux/amd64 (+2), linux/arm64, ...
docker buildx create --name ci --driver docker-container \
--driver-opt network=host --bootstrap --use
docker buildx inspect ci # BuildKit version, platforms, GC policy, worker labels
docker buildx rm ci # also removes its cache volume
# Kubernetes driver: BuildKit pods in a namespace
docker buildx create --name k8s --driver kubernetes \
--driver-opt namespace=buildkit,replicas=3,rootless=true \
--driver-opt requests.cpu=1,requests.memory=2Gi,limits.memory=4Gi,loadbalance=sticky
# Remote driver: an externally managed buildkitd
docker buildx create --name remote --driver remote tcp://buildkitd.internal:1234 \
--driver-opt cacert=/certs/ca.pem,cert=/certs/cert.pem,key=/certs/key.pem
Common docker-container driver-opts: image, network, cgroup-parent (default /docker/buildx), restart-policy (default unless-stopped), memory, cpu-quota, cpu-shares, cpuset-cpus, default-load (default false), env.<KEY>, provenance-add-gha (default true).
Naming note.
--buildkitd-configwas--configbefore buildx v0.13.0 (March 2024).--configstill works but is hidden. On versions < 0.13 use--config.
The docker Driver's Limits — and How the containerd Image Store Changed Them
The classic advice is that the default docker driver cannot do multi-platform, cache export, or attestations. That advice is now conditional on the image store, and the condition flipped under most people's feet.
# Which image store is the daemon using?
docker info -f '{{ .DriverStatus }}'
# [[driver-type io.containerd.snapshotter.v1]] <- containerd store
# [[Backing Filesystem extfs] ...] <- classic overlay2 store
With the containerd image store, the docker driver supports the inline, local, registry, and gha cache backends, multi-platform builds, --load of a multi-platform image, and attestations. With the classic overlay2 store none of that works, and the failures are not always loud.
The containerd store is the default for fresh installs of Docker Engine 29.0 (2025-11-10) and for Docker Desktop 4.34+. Upgrades keep overlay2 until you switch:
// /etc/docker/daemon.json
{
"features": { "containerd-snapshotter": true }
}
Switching does not migrate existing images — they stay in the old store and appear to vanish. That alone is why so many CI images are still on overlay2.
The buildx driver feature matrix in the docs still lists Multi-arch images as unsupported for the
dockerdriver with no containerd footnote. That row is stale; the containerd storage page and observed behaviour both contradict it. Verified locally on Engine 29.7.2 with the containerd store:--platform linux/amd64,linux/arm64builds, pushes a proper OCI index, and--loads.
For CI, use docker-container anyway. It gives you a predictable feature set regardless of what the runner's daemon is configured with, and an isolated cache you can size and garbage-collect.
Building in Anger
docker buildx build --builder ci -t registry.example.com/app:1.2.3 --push .
# Named stage, extra context, build args, labels, index annotations, metadata out
docker buildx build --builder ci \
--target runtime \
--build-context vendored=./third_party \
--build-arg VERSION=1.2.3 \
--label org.opencontainers.image.revision="$(git rev-parse HEAD)" \
--annotation "index:org.opencontainers.image.source=https://git.example.com/app" \
--metadata-file /tmp/meta.json \
-t registry.example.com/app:1.2.3 --push .
--check (buildx 0.15+, Dockerfile frontend 1.8+) lints without building and exits non-zero on any violation, where a normal build only warns — worth wiring into CI:
docker buildx build --check -f Dockerfile.lint .
# Check complete, 1 warning has been found!
#
# WARNING: FromAsCasing - https://docs.docker.com/go/dockerfile/rule/from-as-casing/
# 'as' and 'FROM' keywords' casing do not match
# Dockerfile.lint:1
Suppress rules in the Dockerfile rather than on the CLI, so the exception travels with the code. Both settings go on one directive, separated by a semicolon — a second # check= line is rejected with ERROR: only one check parser directive can be used:
# check=skip=JSONArgsRecommended,StageNameCasing;error=true
Output Types
--output (-o) decides what BuildKit does with the result. Seven exporters; the default is cacheonly except on the docker driver, which defaults to the docker exporter.
type= |
What you get |
|---|---|
image |
Image in the builder's store; push=true sends it onward |
registry |
Shorthand for image with push=true |
docker |
A Docker-format tarball, or loaded into the daemon |
oci |
An OCI image layout (tar or directory) |
local |
The filesystem unpacked to a directory — no image at all |
tar |
The filesystem as a tarball |
cacheonly |
Nothing but cache. Ideal for CI verification builds |
docker buildx build -t app:1 --push . # = --output=type=registry,unpack=false (0.31+)
docker buildx build -t app:1 --load . # = --output=type=docker
docker buildx build --target export -o type=local,dest=./dist . # artefact, no image
docker buildx build -o type=oci,dest=./app.oci.tar . # feed to skopeo/crane
tar tf app.oci.tar | head -3
# blobs/
# blobs/sha256/
# blobs/sha256/738128faa30f570583b0e57efd831e0e6a2a9aacf1be88c8f4c1ef8a5b7033cc
docker buildx build -o type=cacheonly . # verify without producing anything
local and tar split multi-platform output into per-platform subdirectories; platform-split=false merges them. That option is local/tar only — it does not exist on image or oci.
Cache Exporters and Importers
Six backends. mode=min (default) exports only the layers in the final image; mode=max exports every intermediate stage, which is what you want for multi-stage builds.
type= |
Use | Driver support |
|---|---|---|
inline |
Cache metadata embedded in the image itself | docker OK with containerd store |
registry |
Cache as a separate image, usually a :cache tag |
docker OK with containerd store |
local |
A directory on disk | Needs a non-default driver |
gha |
GitHub Actions cache | docker OK with containerd store |
s3 |
An S3 bucket | Needs a non-default driver |
azblob |
Azure Blob Storage | Needs a non-default driver |
# Registry cache, max mode — the default choice for CI
docker buildx build --builder ci \
--cache-from type=registry,ref=registry.example.com/app:cache \
--cache-to type=registry,ref=registry.example.com/app:cache,mode=max \
-t registry.example.com/app:1.2.3 --push .
# Local directory (self-hosted runners with a persistent volume)
docker buildx build --cache-from type=local,src=/var/cache/buildkit \
--cache-to type=local,dest=/var/cache/buildkit,mode=max .
# S3
docker buildx build --cache-to type=s3,region=eu-west-2,bucket=build-cache,name=app,mode=max \
--cache-from type=s3,region=eu-west-2,bucket=build-cache,name=app .
# GitHub Actions
docker buildx build --cache-from type=gha,scope=main --cache-to type=gha,scope=main,mode=max .
inline embeds cache metadata in the image, with no separation between artefact and cache, and does not scale to multi-stage builds because intermediate stages are not in the final image. Use registry with mode=max unless you genuinely have one stage.
GitHub Actions cache. GitHub retired the v1 cache service on 2025-04-15. buildx detects the v2 service from $ACTIONS_CACHE_SERVICE_V2 and switches automatically (from buildx v0.21.0); force it with version=2. url/token default to $ACTIONS_RESULTS_URL and $ACTIONS_RUNTIME_TOKEN, which docker/build-push-action sets for you but a bare docker buildx build step does not.
Multi-architecture Builds
# Emulation: one node, binfmt_misc handles foreign binaries
docker run --privileged --rm tonistiigi/binfmt --install all
docker buildx build --platform linux/amd64,linux/arm64,linux/arm/v7 -t app:1 --push .
# Native nodes: append a second machine to the same builder
docker buildx create --name multi --driver docker-container \
--platform linux/amd64 --node amd64 tcp://amd64-runner:2376
docker buildx create --append --name multi --driver docker-container \
--platform linux/arm64 --node arm64 tcp://arm64-runner:2376
docker buildx build --builder multi --platform linux/amd64,linux/arm64 -t app:1 --push .
Which of those to reach for, and the numbers behind the choice, are in Multi-architecture below.
imagetools and Manifest Lists
imagetools works on registries directly — no local image store involved.
docker buildx imagetools inspect registry.example.com/app:1
# Name: registry.example.com/app:1
# MediaType: application/vnd.oci.image.index.v1+json
# Digest: sha256:dbb89b78081e...
#
# Manifests:
# Name: ...@sha256:7c190ed5f921... Platform: linux/amd64
# Name: ...@sha256:096db7bb4ebd... Platform: linux/arm64
# Name: ...@sha256:9c7eab0e4ed5... Platform: unknown/unknown
# Annotations:
# vnd.docker.reference.type: attestation-manifest
docker buildx imagetools inspect --raw registry.example.com/app:1 # raw JSON
# Assemble a manifest list from per-arch tags built on separate machines
docker buildx imagetools create -t registry.example.com/app:1 \
--annotation "index:org.opencontainers.image.source=https://git.example.com/app" \
registry.example.com/app:1-amd64 registry.example.com/app:1-arm64
Those unknown/unknown entries are not corruption — they are the attestation manifests, addressed by the vnd.docker.reference.digest annotation pointing back at the runnable manifest they describe. Old tooling that iterates the index and tries to pull every entry chokes on them; that is the usual cause of "my deploy tool suddenly can't read my image".
bake: Multi-target Builds
bake is the build equivalent of a Makefile: a declarative set of targets, groups, and variables that one command builds concurrently.
# docker-bake.hcl
variable "REGISTRY" { default = "registry.example.com" }
variable "TAG" { default = "dev" }
group "default" { targets = ["api", "worker"] }
target "_common" {
context = "."
dockerfile = "Dockerfile"
platforms = ["linux/amd64", "linux/arm64"]
cache-from = ["type=registry,ref=${REGISTRY}/cache:build"]
cache-to = ["type=registry,ref=${REGISTRY}/cache:build,mode=max"]
}
target "api" { inherits = ["_common"], target = "api", tags = ["${REGISTRY}/api:${TAG}"] }
target "worker" { inherits = ["_common"], target = "worker", tags = ["${REGISTRY}/worker:${TAG}"] }
# matrix forks one target into variants; it requires a templated name
target "app" {
name = "app-${tgt}"
matrix = { tgt = ["debug", "release"] }
target = tgt
tags = ["app:${tgt}"]
}
docker buildx bake --print # resolve the full plan without building — do this first
docker buildx bake --list=targets # also --list=variables
# TARGET DESCRIPTION
# api
# default api, worker
# worker
docker buildx bake --builder ci --push # build the default group
docker buildx bake --set '*.platform=linux/amd64' --set 'api.tags=api:hotfix' api
docker buildx bake --var TAG=1.2.3 --push # --var is buildx 0.31+
docker buildx bake "https://github.com/org/repo.git#main" # remote definition
Files are searched in this order: compose.yaml, compose.yml, docker-compose.yml, docker-compose.yaml, docker-bake.json, docker-bake.hcl, docker-bake.override.json, docker-bake.override.hcl. Compose files are valid bake input directly — the cheapest way to get a multi-service build out of an existing compose.yaml.
Attestations: Provenance and SBOM
SLSA provenance at mode=min is added by default — has been since buildx v0.10. You do not opt in; you opt out.
docker buildx build -t app:1 --push . # min provenance, no SBOM
docker buildx build --provenance=mode=max --sbom=true -t app:1 --push .
docker buildx build --attest type=sbom,generator=<scanner-image> -t app:1 --push .
docker buildx build --provenance=false --sbom=false -t app:1 --push .
export BUILDX_NO_DEFAULT_ATTESTATIONS=1 # same, environment-wide
docker buildx imagetools inspect app:1 --format '{{ json .Provenance }}'
docker buildx imagetools inspect app:1 --format '{{ json .SBOM.SPDX }}'
# spdxVersion SPDX-2.3 | name sbom
# creators ['Organization: Anchore, Inc', 'Tool: syft-v1.51.0', 'Tool: buildkit-v0.32.2']
docker buildx imagetools inspect app:1 --format '{{ json (index .SBOM "linux/amd64").SPDX }}'
To scan the build context and named stages as well as the final image, add real ARG instructions to the Dockerfile — these cannot be set from the environment:
ARG BUILDKIT_SBOM_SCAN_CONTEXT=true
ARG BUILDKIT_SBOM_SCAN_STAGE=build,runtime
Footgun.
--provenance=mode=maxrecords every build argument verbatim in the pushed attestation. Verified: a build run with--build-arg API_TOKEN=s3cr3t-do-not-shipproduced provenance containing"build-arg:API_TOKEN": "s3cr3t-do-not-ship"underbuildDefinition.externalParameters.request.args. The defaultmode=mindoes not. See Build Secrets below.
Frontends and the # syntax= Directive
BuildKit does not natively understand Dockerfiles. It runs a frontend — a container image — that compiles the input into BuildKit's LLB graph. That is why Dockerfile features can ship independently of your Docker version.
# syntax=docker/dockerfile:1
Channels: docker/dockerfile:1 tracks the latest stable v1 release (currently 1.27.0) and takes minor and patch updates; docker/dockerfile:1-labs adds experimental instructions. Pin a full version (:1.27.0) plus digest if you need bit-stability.
# Same effect without touching the Dockerfile
docker buildx build --build-arg BUILDKIT_SYNTAX=docker/dockerfile:1 .
# A completely custom frontend
# syntax=registry.example.com/frontends/mylang:2@sha256:abc...
If you use RUN --mount=type=cache, RUN --mount=type=secret, heredocs, or COPY --parents, you need the directive. Without it you get whatever frontend is baked into the daemon, which on older hosts silently rejects the syntax.
buildkitd Standalone in Kubernetes
For a build service not tied to a Docker daemon, run buildkitd yourself and drive it with buildctl or a remote buildx builder. Upstream ships example manifests in examples/kubernetes/: pod.rootless.yaml, deployment+service.rootless.yaml, statefulset.rootless.yaml, plus .privileged. and .userns. variants of each, and create-certs.sh for TLS.
apiVersion: v1
kind: Pod
metadata:
name: buildkitd
spec:
containers:
- name: buildkitd
image: moby/buildkit:v0.33.0-rootless
args: [--oci-worker-no-process-sandbox]
securityContext:
seccompProfile: { type: Unconfined }
appArmorProfile: { type: Unconfined }
runAsUser: 1000
runAsGroup: 1000
buildctl --addr kube-pod://buildkitd build \
--frontend dockerfile.v0 --local context=. --local dockerfile=. \
--output type=image,name=registry.example.com/app:1,push=true
docker buildx create --name k8s-remote --driver remote kube-pod://buildkitd
The rootless variant needs unconfined seccomp and AppArmor because it uses nested user namespaces; hardened clusters that forbid unconfined profiles cannot run it. Note the manifests now use the securityContext.seccompProfile / appArmorProfile fields — the old container.apparmor.security.beta.kubernetes.io/... annotation form is stale. There is no official Helm chart or operator.
Buildah
Buildah is what Podman uses to build images — podman build calls Buildah's Go API. As a CLI it does two distinct things, and the second is the reason to reach for it.
Dockerfile-compatible Builds
buildah build -t app:1 . # `bud` is a retained alias
buildah build --layers -t app:1 . # keep intermediate layers (off by default)
buildah build --format docker -t app:1 . # Docker v2s2 manifest instead of OCI
buildah build --platform linux/arm64 -t app:1 .
buildah build --jobs 4 -t app:1 . # parallel stages
buildah build --secret id=npmrc,src=$HOME/.npmrc -t app:1 .
buildah build --ssh default -t app:1 .
buildah build --cache-to registry.example.com/app:cache \
--cache-from registry.example.com/app:cache --cache-ttl 24h -t app:1 .
Defaults, verified on 1.39.3 (Debian 13, rootless):
| Flag | Default | Override |
|---|---|---|
--format |
oci |
BUILDAH_FORMAT=docker |
--isolation |
rootless for unprivileged users, oci for root |
BUILDAH_ISOLATION |
--layers |
false |
BUILDAH_LAYERS=true |
--jobs |
1 |
— |
--squash on buildah build squashes all layers including the base image's into one — stronger than podman build --squash, which only squashes the layers the build itself added. The flag name is the same in both tools and the semantics are not. There is no --squash-all in Buildah.
The Scripted, Dockerfile-free Build
This is what makes Buildah different. Instead of a declarative file you get shell verbs against a working container: no parser, no frontend, no DSL — just the filesystem and buildah config.
#!/usr/bin/env bash
set -euo pipefail
# Harvest a static binary from a published image
src=$(buildah from docker.io/library/busybox:1.37-musl)
srcmnt=$(buildah mount "$src")
# A scratch container is an empty rootfs — no base layer at all
ctr=$(buildah from scratch)
mnt=$(buildah mount "$ctr")
mkdir -p "$mnt/bin" "$mnt/etc"
cp "$srcmnt/bin/busybox" "$mnt/bin/busybox"
ln -sf busybox "$mnt/bin/sh"
printf 'app:x:1000:1000::/:/bin/sh\n' > "$mnt/etc/passwd"
printf 'app:x:1000:\n' > "$mnt/etc/group"
buildah umount "$src" >/dev/null && buildah rm "$src" >/dev/null
buildah config \
--entrypoint '["/bin/sh"]' \
--cmd '["-c","echo hello from scratch"]' \
--user 1000:1000 --workingdir / --env APP_ENV=prod \
--label org.opencontainers.image.title=scripted-demo \
"$ctr"
buildah umount "$ctr" >/dev/null
buildah commit --format oci --rm "$ctr" scripted-demo:1
Run it through buildah unshare so the mount happens inside your own user namespace:
buildah unshare -- bash ./build.sh
# 3bcbc8c0c6af73a88bb50ff558eb40faf22c98e96385253f43782812072e5012
buildah images --format '{{.Name}}:{{.Tag}} {{.Size}}'
# localhost/scripted-demo:1 1.22 MB
buildah inspect --type image --format '{{len .OCIv1.RootFS.DiffIDs}}' scripted-demo:1
# 1 <- one layer, one history entry
podman inspect scripted-demo:1 --format '{{.ManifestType}}'
# application/vnd.oci.image.manifest.v1+json
podman run --rm scripted-demo:1
# hello from scratch
Why bother: exact control over layer count and contents, any host tool available to populate the rootfs (a package manager, rsync, a Python script), and no Dockerfile grammar between you and the filesystem. The classic use is assembling minimal images from artefacts a separate pipeline already produced.
buildah config takes JSON array form for --entrypoint and --cmd; a bare string is split on whitespace, which is rarely what you want. Other flags: --env, --label, --annotation, --port, --user, --workingdir (one word, no hyphen), --volume, --author, --created-by, --arch, --os, --shell, --stop-signal, --healthcheck (plus --healthcheck-interval, --healthcheck-retries, --healthcheck-timeout, --healthcheck-start-period), --history-comment, --onbuild, --unsetlabel, and --unsetannotation (1.41+; unknown flag on 1.39.3).
Rootless Mechanics
Buildah's rootless story is the best of the six, and it hinges on user namespaces.
grep "^$(id -un)" /etc/subuid /etc/subgid # ranges must exist, or nothing works
# /etc/subuid:mike:165536:65536
# /etc/subgid:mike:165536:65536
podman info --format '{{.Store.GraphDriverName}} {{.Store.GraphStatus}}'
# overlay map[Backing Filesystem:extfs Native Overlay Diff:true ...]
ctr=$(buildah from scratch)
buildah mount "$ctr"
# Error: cannot mount using driver overlay in rootless mode.
# You need to run it in a `buildah unshare` session
buildah build handles the namespace for you; anything touching the container's filesystem directly does not. That error is the single most common rootless Buildah failure, and the fix is in the message — buildah unshare drops you into a user namespace where your UID maps to root and the subuid range maps everything else, so mount, chown, and direct writes behave.
Rootless storage lives in ~/.local/share/containers/storage (root: /var/lib/containers/storage). On kernels ≥ 5.11 you get native rootless overlayfs (Native Overlay Diff: true above); older kernels fall back to fuse-overlayfs, and if that is missing you silently drop to the vfs driver, which copies the whole filesystem per layer and is glacial.
Output Format, Podman, and Skopeo
Buildah defaults to OCI manifests; Docker's tooling historically defaults to Docker v2s2, and some older registries and admission controllers still choke on OCI media types. buildah build --format docker (or BUILDAH_FORMAT=docker) switches.
Multi-arch is manual — build per platform, then assemble. Buildah has no built-in emulation, so without host binfmt handlers a cross-arch RUN fails immediately:
buildah manifest create app:1
buildah build --platform linux/amd64 --manifest app:1 -t app:1-amd64 .
buildah build --platform linux/arm64 --manifest app:1 -t app:1-arm64 .
buildah manifest push --all app:1 docker://registry.example.com/app:1
buildah build --platform linux/arm64 -t app:1 . # no binfmt registered
# STEP 2/2: RUN echo hello > /hello.txt
# exec container process `/bin/sh`: Exec format error
# Error: building at STEP "RUN ...": while running runtime: exit status 1
Podman and Buildah share image storage but not container state ("you can not see Podman containers from within Buildah or vice versa"). Skopeo is the third of the trio, for copying and inspecting images between registries and stores without a daemon. See the Podman and Container Registries sheets.
Kaniko
Status: Archived
GoogleContainerTools/kaniko was archived on 2025-06-03 and is read-only. The README opens with a banner stating the project is archived and no longer developed or maintained. The final release is v1.24.0 (2025-05-23). There is no named successor, no new registry location, and ~760 issues frozen in place. It was never an officially supported Google product.
You will still find it recommended in blog posts, Helm charts, and internal CI templates written between 2019 and 2024, because for several years it was the only credible answer to "build an image in an unprivileged Kubernetes pod". It no longer is. Migrate to rootless buildkitd (see above) or Buildah in a pod.
It is documented here because you will inherit it, and because understanding its mechanism explains its constraints.
How It Actually Works
Kaniko has no daemon and does not use overlayfs. It extracts the base image into its own container's root filesystem, then executes each Dockerfile instruction in that filesystem, taking a userspace snapshot after each one and diffing to produce a layer.
INFO Unpacking rootfs as cmd RUN echo hello > /hello.txt requires it.
INFO Initializing snapshotter ...
INFO Taking snapshot of full filesystem...
INFO Running: [/bin/sh -c echo hello > /hello.txt]
INFO Taking snapshot of full filesystem...
That is the whole design, and every consequence follows from it:
- It must run inside a disposable container. It writes to
/in whatever it is running in. Outside a container it guards itself and refuses:kaniko should only be run inside of a container, run with the --force flag if you are sure you want to continue.--forceis documented as an escape hatch for gVisor, where containerisation cannot be detected — not as a general blessing. Never use it on a machine whose filesystem you care about. - Run only the official image. Running the executor binary in some other image "is not supported due to implementation details" — it cannot chroot or bind-mount, so it will overwrite whatever is already there.
- Full-filesystem snapshots are expensive.
--snapshot-modetrades correctness for speed:full(default, hashes file contents),redo(mtime + size + permissions),time(mtime only, fastest, misses same-second changes). - No BuildKit features. No
RUN --mount=type=cache, no--mount=type=secret, no heredocs, no# syntax=frontends — there is no BuildKit and no frontend. Note this is an inference from the archived feature set rather than an explicit statement in the README.
Running It
apiVersion: v1
kind: Pod
metadata:
name: kaniko-build
spec:
restartPolicy: Never
containers:
- name: kaniko
image: gcr.io/kaniko-project/executor:v1.24.0
args:
- --context=git://github.com/org/repo.git#refs/heads/main
- --dockerfile=Dockerfile
- --destination=registry.example.com/app:1.2.3
- --cache=true
- --cache-repo=registry.example.com/app/cache
- --snapshot-mode=redo
volumeMounts:
- { name: docker-config, mountPath: /kaniko/.docker }
volumes:
- name: docker-config
secret:
secretName: regcred
items: [{ key: .dockerconfigjson, path: config.json }]
Credentials go in a Docker config JSON at /kaniko/.docker/config.json. Contexts: dir://, tar:// (including tar://stdin), git://<url>#<ref>#<commit>, gs://, s3://, and Azure Blob over https://.
# Local trial: no push, tarball out
podman run --rm -v "$PWD/ctx":/workspace -v "$PWD":/out \
gcr.io/kaniko-project/executor:v1.24.0 \
--context dir:///workspace --dockerfile /workspace/Dockerfile \
--destination example.invalid/demo:1 --no-push --tar-path /out/image.tar
Other flags: --cache-ttl (default two weeks), --cache-dir (default /cache), --cache-copy-layers, --cache-run-layers (default true), --single-snapshot, --reproducible, --digest-file, --target, --build-arg, --ignore-path, --push-retry, --registry-mirror, --skip-tls-verify. A companion gcr.io/kaniko-project/warmer image pre-populates the base-image cache; the :debug tag adds a busybox shell.
Cache reads stop after the first miss — every layer after that rebuilds locally, even if a later one was cached. Multi-arch does not exist: --custom-platform (hyphenated; not --customPlatform) only changes the declared platform and, per the README, "is not virtualization and cannot help to build an architecture not natively supported by the build host". Stitch per-arch builds with imagetools create or manifest-tool.
Stacker
Stacker builds OCI images from a YAML file, unprivileged, using LXC-sealed build containers. Apache-2.0, Cisco-originated, and genuinely maintained — v1.2.1 landed 2026-08-25 — but small: ~370 stars, ~113 open issues, and the published docs site is still versioned at v1.0.0 while releases are at v1.2.1. The v1.2.1 tag shipped no release binaries (v1.2.0 did), so plan on pinning an older tag or building from source, which needs Go, lxc-devel, libcap, libacl, gpgme, and mksquashfs.
# stacker.yaml
build:
build_only: true # produces artefacts, not an output image
from:
type: docker
url: docker://docker.io/library/alpine:3.22
run: |
mkdir -p /out
echo "built by stacker" > /out/hello.txt
demo:
from:
type: docker
url: docker://docker.io/library/busybox:1.37-musl
imports:
- stacker://build/out/hello.txt # pull an artefact from another layer
run: |
cp /stacker/imports/hello.txt /hello.txt
entrypoint: /bin/sh
cmd: ["-c", "cat /hello.txt"]
environment:
APP_ENV: prod
labels:
org.opencontainers.image.title: stacker-demo
stacker check # are the kernel features present?
stacker build
# preparing image build...
# + mkdir -p /out
# preparing image demo...
# + cp /stacker/imports/hello.txt /hello.txt
# filesystem demo built successfully
ls # oci/ roots/ .stacker/ stacker.yaml
stacker build --layer-type squashfs # tar (default), squashfs, or erofs
stacker build --substitute VERSION=1.2.3 # ${{VERSION}} in the YAML
stacker build --require-hash # refuse imports without a pinned sha256
stacker recursive-build -d ./images # build every stacker.yaml under a tree
stacker publish --url docker://registry.example.com --tag 1.2.3
stacker unpriv-setup # one-time subuid/subgid configuration
Layer keys: from (type: one of docker, oci, tar, built, scratch), imports (plural — import is the deprecated legacy form), overlay_dirs, run, cmd, entrypoint, full_command, environment (not env), build_env, build_env_passthrough, volumes, labels, generate_labels, annotations, working_dir, runtime_user, build_only, binds, os, arch, bom.
Output is a real OCI layout in oci/, unpacked rootfs in roots/, cache in .stacker/. Unprivileged builds need /etc/subuid and /etc/subgid ranges (stacker unpriv-setup writes them) and kernel ≥ 5.11 for unprivileged overlay mounts. Verified on Debian 13 (kernel 6.12) as an ordinary user with only pre-existing subuid/subgid ranges — no setup step needed.
The provenance story is the interesting part: layers are content-hashed and rebuilt only when inputs change, imports can be pinned with a hash: sha256 (--require-hash enforces it fleet-wide), and the entire stacker.yaml is embedded in the image as an OCI annotation:
Annotations:
io.stackeroci.stacker.stacker_version: v1.2.0
io.stackeroci.stacker.stacker_yaml: |
build:
build_only: true
...
The recipe travels with the artefact, which is genuinely nice. There is no published guarantee of bit-identical layer digests, so do not promise one. Its niche is base images and OS-like layers where you want explicit layer control and no daemon; documented adopters are thin on the ground beyond Cisco.
apko and melange
This is the most different model in the set, and it is worth understanding even if you do not adopt it.
There is no RUN. An apko image is composed from a list of APK packages, not built by executing commands. Everything bespoke must first be packaged, and melange is what packages it. Removing arbitrary execution from image assembly is what buys the two properties apko claims: bitwise reproducibility, and complete SBOM coverage.
flowchart LR
SRC["Your source"] --> ML["melange build"]
ML --> APK["signed .apk + APKINDEX.tar.gz"]
UP["Upstream Wolfi repo"] --> AK["apko build"]
APK --> AK
AK --> IMG["OCI image + SPDX SBOM"]
apko: Composing the Image
# apko.yaml
contents:
repositories: [https://packages.wolfi.dev/os]
keyring: [https://packages.wolfi.dev/os/wolfi-signing.rsa.pub]
packages:
- wolfi-baselayout
- ca-certificates-bundle
- busybox
accounts:
groups: [{ groupname: app, gid: 65532 }]
users: [{ username: app, uid: 65532, gid: 65532 }]
run-as: 65532
entrypoint:
command: /bin/sh -l
environment:
APP_ENV: prod
archs: [x86_64, aarch64]
# apko build <config> <tag> <output> -- exactly three arguments
podman run --rm -v "$PWD":/work -w /work cgr.dev/chainguard/apko \
build apko.yaml demo:latest demo.tar
ls
# apko.yaml demo.tar sbom-aarch64.spdx.json sbom-index.spdx.json sbom-x86_64.spdx.json
apko publish apko.yaml registry.example.com/demo:latest # build straight to a registry
The schema's mixed casing matters. Hyphenated: shell-fragment, stop-signal, work-dir, run-as, vcs-url. Underscored: build_repositories, runtime_repositories, runtime_keyring. Multi-layer output is opt-in via layering: {strategy: origin, budget: 10}; the default is one layer. There is no top-level options key any more, and no --rewrite-tags flag.
Reproducibility — verified. Two independent builds of the same config, minutes apart:
sha256sum demo.tar demo2.tar
# 26a3326bde49c53ba3014693b32a7e8de7b7f2712b8978946df8c38ab92da48c demo.tar
# 26a3326bde49c53ba3014693b32a7e8de7b7f2712b8978946df8c38ab92da48c demo2.tar
Byte-identical, per-arch layer digests included. SOURCE_DATE_EPOCH is honoured and always overrides --build-date; when it is unset the default epoch is the build time of the newest installed APK, not the current time. What breaks reproducibility is the package set moving underneath you, so pin it — apko lock apko.yaml writes apko.yaml.lock.json, replayed with apko build --lockfile. (lock and resolve are hidden from --help pending feedback; resolve is deprecated in favour of lock.)
SBOMs are SPDX JSON written beside the output as sbom-<arch>.spdx.json plus sbom-index.spdx.json. They are not attached as OCI attestations — that is a separate cosign step. apko composes them from per-package SBOMs it finds inside each APK at /var/lib/db/sbom/, which is why coverage is complete rather than inferred by scanning. apko needs no root and no runner: where it cannot chown or create device nodes directly, it records the intent and writes correct ownership into the layer tar stream.
melange: Building the Packages
# hello.yaml
package:
name: hello-demo
version: 1.0.0
epoch: 0
description: A tiny demo package built by melange
copyright:
- license: Apache-2.0
environment:
contents:
repositories: [https://packages.wolfi.dev/os]
keyring: [https://packages.wolfi.dev/os/wolfi-signing.rsa.pub]
packages: [busybox, wolfi-baselayout]
pipeline:
- name: Install the script
runs: |
mkdir -p "${{targets.destdir}}/usr/bin"
install -m0755 hello "${{targets.destdir}}/usr/bin/hello"
melange keygen # melange.rsa + melange.rsa.pub, 4096-bit
melange build hello.yaml --arch x86_64 --signing-key melange.rsa --runner bubblewrap
# INFO wrote packages/x86_64/hello-demo-1.0.0-r0.apk
# INFO generating apk index from packages in packages/x86_64
# INFO signing index packages/x86_64/APKINDEX.tar.gz with key melange.rsa
Built-in pipelines cover most of the ground — fetch, git-checkout, patch, strip, autoconf/configure, autoconf/make, autoconf/make-install, cmake/configure, cmake/build, go/build, go/install, split/dev, split/manpages, split/debug — invoked as - uses: autoconf/make. Runners are bubblewrap, docker, and qemu, defaulting by platform. --out-dir defaults to ./packages/, --cache-dir to ./melange-cache/, --generate-index to true.
Feeding melange output back into apko uses a labelled local repository — a bare filesystem path, not a file:// URL:
contents:
repositories:
- https://packages.wolfi.dev/os
- "@local /work/packages"
keyring:
- https://packages.wolfi.dev/os/wolfi-signing.rsa.pub
- /work/melange.rsa.pub
packages:
- wolfi-baselayout
- busybox
- hello-demo@local # pin this package to the @local repo
Verified end to end: melange build → apko build → podman run printed the packaged script's output, and the resulting image SBOM listed 43 packages including hello-demo.
What It Costs You
"Distroless by construction" is real — the image holds exactly the packages you named and their dependencies, with no shell, no apk, and no package database at runtime unless you add busybox and apk-tools yourself. Excellent for attack surface and SBOM fidelity.
The price is that your whole dependency graph has to exist as APKs. For anything Wolfi already ships that is free; for your own application it means writing and maintaining melange definitions, which is a packaging discipline most teams do not have. Dockerfile → apko is not a translation, it is a change of model. Budget accordingly.
ko
ko builds container images from Go source with no Dockerfile, no daemon, and no privileges. Deliberately narrow and, within that scope, unbeatable.
export KO_DOCKER_REPO=registry.example.com/org
ko build ./cmd/app # build and push
ko build ./cmd/app --platform=linux/amd64,linux/arm64
ko build ./cmd/app --platform=all # every platform the base image has
ko build ./cmd/app --push=false --oci-layout-path=./oci
ko build ./cmd/app --local # side-load into the local daemon
ko build ./cmd/app --bare # image name == KO_DOCKER_REPO exactly
ko build ./cmd/app --sbom=spdx # default; only other value is "none"
ko apply -f config/ # resolve ko:// refs, then kubectl apply
ko resolve -f config/ > release.yaml
In a manifest, image: ko://example.com/org/cmd/app is replaced by the pushed digest at ko apply/ko resolve time. KO_DOCKER_REPO=ko.local side-loads into the Docker daemon; kind.local loads into kind nodes. Static assets live in <importpath>/kodata/ and are found at runtime via $KO_DATA_PATH; the binary is always installed as /ko-app/<name>.
# .ko.yaml
defaultBaseImage: cgr.dev/chainguard/static
baseImageOverrides:
example.com/org/cmd/needs-libc: cgr.dev/chainguard/glibc-dynamic
defaultPlatforms: [linux/amd64, linux/arm64]
builds:
- id: app
dir: ./cmd/app
ldflags: ["-s -w", "-X main.version={{.Env.VERSION}}"]
env: [CGO_ENABLED=0]
The default base image is cgr.dev/chainguard/static (not distroless, as older material says) — verified from a live build log:
Using base cgr.dev/chainguard/static:latest@sha256:f51c2493951313c3ad... for .../cmd/hello
Building example.invalid/hello/cmd/hello for linux/amd64
Building example.invalid/hello/cmd/hello for linux/arm64
Both architectures were built on an x86_64 host with no emulation and no QEMU — Go cross-compiles, so multi-arch costs a second go build and nothing else. That is ko's single biggest practical advantage over every Dockerfile-based builder.
Reproducibility — verified. Two independent builds with the same SOURCE_DATE_EPOCH produced the identical index digest:
SOURCE_DATE_EPOCH=1700000000 ko build ./cmd/hello --push=false --oci-layout-path=./oci2 --bare
# ./oci2@sha256:e56d0a6ed8fc9bce8269f1f84cd45c4da18f6d214b0957fbd5ac9ccfc33a8365
SOURCE_DATE_EPOCH=1700000000 ko build ./cmd/hello --push=false --oci-layout-path=./oci3 --bare
# ./oci3@sha256:e56d0a6ed8fc9bce8269f1f84cd45c4da18f6d214b0957fbd5ac9ccfc33a8365
By default ko embeds no timestamps at all (images show 1970), which is why they reproduce. SOURCE_DATE_EPOCH sets the image creation time; KO_DATA_DATE_EPOCH sets modtimes for kodata assets.
Hard limits. Go only. CGO_ENABLED=0 by default, so cgo needs a custom base and explicit configuration. No OS packages — ca-certificates, ICU data, or a shell is a base-image change, not a build step. ko publish and ko deps no longer exist, and --sbom accepts only spdx and none (anything else silently becomes SPDX). ko is a CNCF Sandbox project.
Rootless and CI
What each tool actually needs on a shared runner:
| Tool | Requirement | The trap |
|---|---|---|
buildx, docker driver |
Access to the Docker socket | Socket access is root on the host. A "non-privileged" job with /var/run/docker.sock mounted can start a privileged container |
buildx, docker-container |
Same socket | Same trap. The builder container is created by the daemon |
buildx, rootless buildkitd |
Unprivileged user namespaces; unconfined seccomp and AppArmor | Some hardened clusters forbid unconfined profiles outright |
| Buildah | subuid/subgid ranges; newuidmap/newgidmap setuid helpers |
Missing ranges give a confusing lchown ... invalid argument from the storage driver |
| Kaniko | An isolated pod it can safely trash | Running it anywhere it can reach a real filesystem |
| Stacker | subuid/subgid; kernel ≥ 5.11 for unprivileged overlay | stacker unpriv-setup writes the ranges, and that step itself needs privilege |
| apko | Nothing | — |
| melange | Unprivileged user namespaces for bwrap |
Nesting inside a rootless container fails: bwrap: Creating new namespace failed: Operation not permitted |
| ko | A Go toolchain | — |
The privilege-escalation trap is worth stating plainly: mounting the Docker socket into a CI job grants root on the runner. A job that can call docker run --privileged -v /:/host owns the machine and every other tenant's secrets on it. If your runners are shared, the honest options are rootless BuildKit, Buildah in a rootless container, or a remote builder the job cannot reach except through the build API.
# Rootless Buildah inside a container: the container itself needs userns support
podman run --rm --device /dev/fuse \
--security-opt seccomp=unconfined --security-opt apparmor=unconfined \
-v "$PWD":/src -w /src quay.io/buildah/stable \
buildah build -t app:1 .
--device /dev/fuse is needed when the nested storage driver falls back to fuse-overlayfs. Without it you land on vfs and builds become minutes-per-layer.
Multi-architecture
The performance reality first. Measured on Docker Engine 29.7.2 / BuildKit v0.32.2 on Apple Silicon — same shell-loop workload, --no-cache, base image pre-warmed:
linux/arm64 (native) 0.4s
linux/amd64 (Rosetta) 0.5s ~1.2x
linux/arm/v7 (QEMU) 4.0s ~10x
linux/s390x (QEMU) 4.9s ~12x
The lesson is not "emulation is slow"; it is that QEMU is slow and Rosetta is not QEMU. On an Apple Silicon workstation amd64 emulation is nearly free, because Docker Desktop uses Rosetta 2. On a Linux CI runner with qemu-user-static you pay the QEMU multiplier, and a compile-heavy build that takes four minutes natively takes tens of minutes. That is the "why is my ARM build taking an hour" question, answered — and why you cannot benchmark emulation on a Mac and extrapolate.
The multiplier itself is not portable either. The table above is an Apple Silicon measurement. The same workload shape on a Linux/amd64 box (1 vCPU, Debian 13, tonistiigi/binfmt) measured native 1.6s, arm64 8.1s, arm/v7 8.0s, s390x 9.3s — roughly 5–6x, not 10–12x. Take the order of magnitude, not the number: QEMU costs you multiples, and which multiple depends on the host, the guest architecture and how compute-bound the step is.
Three approaches, in decreasing order of how much you should like them:
1. Native nodes. One builder, several machines, each advertising its own platforms. Fastest and least surprising; costs you runners.
2. Cross-compilation in the Dockerfile. Build natively on $BUILDPLATFORM and copy the artefact into a $TARGETPLATFORM stage, so only the final stage is emulated and it usually runs nothing. This is what ko does implicitly, and why ko is fast.
FROM --platform=$BUILDPLATFORM golang:1.26 AS build
ARG TARGETOS TARGETARCH
WORKDIR /src
COPY . .
RUN CGO_ENABLED=0 GOOS=$TARGETOS GOARCH=$TARGETARCH go build -o /out/app ./cmd/app
FROM gcr.io/distroless/static-debian12
COPY --from=build /out/app /app
ENTRYPOINT ["/app"]
3. Emulation. Register binfmt handlers and build. Correct, universal, and slow.
Assembly after the fact. If per-arch builds run as separate jobs, stitch them at the end; --metadata-file gives you the digests to feed in. Buildah's equivalent is buildah manifest create / add / push --all. Kaniko has no equivalent and needs an external tool.
docker buildx imagetools create -t registry.example.com/app:1 \
registry.example.com/app@sha256:aaa... registry.example.com/app@sha256:bbb...
Reproducibility
Two things get conflated here. Stable means the digest does not change when nothing changed. Reproducible means someone else, on another machine, gets your exact digest from your exact inputs. Most pipelines want the first and claim the second.
Plain buildx builds are neither, because every layer carries a fresh timestamp. Fix it with SOURCE_DATE_EPOCH and rewrite-timestamp:
docker buildx build --no-cache -t app:a --push . # sha256:df331dfe244340...
docker buildx build --no-cache -t app:b --push . # sha256:c07489586118fb... <- differs
export SOURCE_DATE_EPOCH=1700000000
docker buildx build --no-cache --provenance=false --sbom=false \
-o type=registry,name=registry.example.com/app:c,rewrite-timestamp=true .
# sha256:<digest A>
docker buildx build --no-cache --provenance=false --sbom=false \
-o type=registry,name=registry.example.com/app:d,rewrite-timestamp=true .
# sha256:<digest A> <- identical, and stable across hosts for the same inputs
The split of responsibilities matters:
SOURCE_DATE_EPOCH(BuildKit 0.11+, auto-forwarded from the client environment by buildx 0.10+) rewrites the image configcreated, each history entry'screated, and theorg.opencontainers.image.createdannotation. It does not touch file mtimes inside the layers.rewrite-timestamp=true(BuildKit 0.13+, on theimage,registry,oci, anddockerexporters) rewrites those file timestamps to the same epoch.- Provenance attestations embed build timing and environment, so the index digest moves even when the image manifest does not. Turn them off, or accept it.
SOURCE_DATE_EPOCH=contextderives the value from the build context — a git commit time, or an HTTPLast-Modified.- BuildKit also exposes a
compatibility-versionexporter option pinning digest-affecting assembly behaviour (10= v0.13/v0.14,20= v0.15–v0.31,30= current), for matching a digest computed by an older BuildKit.
What genuinely reproduces versus what merely looks stable:
| Reproduces across machines? | |
|---|---|
| apko | Yes, bitwise — same config plus same package set, verified above |
| ko | Yes with SOURCE_DATE_EPOCH, verified above; Go builds are deterministic and -trimpath is always on |
| Stacker | Content-hashed layers plus hash:-pinned imports; no published bit-identical guarantee |
buildx + SOURCE_DATE_EPOCH + rewrite-timestamp |
Stable, and reproducible if the Dockerfile is |
Buildah --timestamp |
Same caveats; --source-date-epoch and --rewrite-timestamp need 1.41+ (absent from 1.39.3) |
Kaniko --reproducible |
Exists; unmaintained since 2025 |
The uncomfortable truth: the builder is rarely why a build fails to reproduce — unpinned package managers are. apt-get install without pinned versions, pip install without a lockfile, or any curl | sh breaks it regardless of the builder. Pin the base image by digest and dependencies by lockfile, and only then blame the tooling.
Build Secrets
Never use ARG or ENV for a secret. Verified — a build arg lands in the image history, readable by anyone who can pull the image:
docker buildx build --build-arg API_TOKEN=s3cr3t-do-not-ship -t app:1 --push .
docker buildx imagetools inspect app:1 --format '{{ json .Image.History }}'
# { "created_by": "ARG API_TOKEN=s3cr3t-do-not-ship", "empty_layer": true },
# { "created_by": "RUN |1 API_TOKEN=s3cr3t-do-not-ship /bin/sh -c echo ... # buildkit" }
And with --provenance=mode=max it also lands in the pushed attestation:
docker buildx imagetools inspect app:1 --format '{{ json .Provenance }}' | python3 -c \
'import json,sys; print(json.load(sys.stdin)["SLSA"]["buildDefinition"]["externalParameters"]["request"]["args"])'
# {'build-arg:API_TOKEN': 's3cr3t-do-not-ship', 'no-cache': ''}
The default mode=min does not carry args — which is exactly why people turn mode=max on for supply-chain compliance and hand out their tokens in the process.
Use secret mounts instead. The secret is bind-mounted for the lifetime of one RUN and never enters a layer, the history, or the provenance:
# syntax=docker/dockerfile:1
FROM alpine:3.22
RUN --mount=type=secret,id=api_token,required=true \
curl -H "Authorization: Bearer $(cat /run/secrets/api_token)" -O https://internal/artifact
RUN --mount=type=ssh git clone git@git.example.com:org/private.git /src
docker buildx build --secret id=api_token,src=./token.txt . # from a file
API_TOKEN=... docker buildx build --secret id=api_token,env=API_TOKEN . # from the environment
docker buildx build --ssh default . # agent forwarding
docker buildx build --ssh mykey=$HOME/.ssh/id_ed25519 .
Mount options: id, target (default /run/secrets/<id>), env (mount as an environment variable instead of a file — Dockerfile frontend 1.10+), required, mode, uid, gid. The env= form on --secret has been available since buildx v0.6.
Buildah takes the same flags (buildah build --secret id=npmrc,src=$HOME/.npmrc --ssh default .); Kaniko supports neither. The SecretsUsedInArgOrEnv build check catches the obvious cases — docker buildx build --check will flag them. For where the secrets come from in the first place, see the SOPS and HashiCorp Vault sheets.
Signing and Attestations
Build-time attestations prove how an image was built; signatures prove who published it. You want both, and they are separate artefacts in the registry.
cosign generate-key-pair # cosign.key + cosign.pub
# Always sign by DIGEST. A tag is mutable, and a signature over a tag proves nothing.
DIGEST=$(docker buildx imagetools inspect registry.example.com/app:1 \
--format '{{.Manifest.Digest}}')
cosign sign --key cosign.key registry.example.com/app@$DIGEST
cosign verify --key cosign.pub registry.example.com/app@$DIGEST
# The signature lands as an extra tag alongside the image
curl -s registry.example.com/v2/app/tags/list
# {"name":"app","tags":["1","sha256-fd26cffe341fe85ff138704f419049637b644368422b40c420f60a2db0016d56"]}
# Attach an externally generated SBOM — apko's, for instance, which is not attached by default
cosign attest --key cosign.key --type spdxjson \
--predicate sbom-x86_64.spdx.json registry.example.com/app@$DIGEST
cosign verify-attestation --key cosign.pub --type spdxjson registry.example.com/app@$DIGEST
# Keyless via OIDC — the better default in CI: no key material, identity bound to the workflow
cosign sign --yes registry.example.com/app@$DIGEST
cosign verify \
--certificate-identity-regexp 'https://github.com/org/repo/.github/workflows/.*' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
registry.example.com/app@$DIGEST
cosign 3.x change.
--tlog-upload=falseis deprecated and now conflicts with the default signing config: "--tlog-upload=false is not supported with --signing-config". To sign without a transparency log — an air-gapped or internal registry — pass a signing config with the Rekor services removed, built withcosign signing-config create. On cosign 2.x,--tlog-upload=falsestill works.
For scanning the result see the Trivy and Grype and Syft sheets; for registry-side policy and admission control, Container Registries.
CI Integration
GitHub Actions with buildx
name: build
on:
push:
tags: ['v*']
permissions:
contents: read
packages: write
id-token: write # required for keyless cosign
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: docker/setup-qemu-action@v4 # only if you actually emulate
- uses: docker/setup-buildx-action@v4
- uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- id: meta
uses: docker/metadata-action@v6
with:
images: ghcr.io/${{ github.repository }}
tags: |
type=semver,pattern={{version}}
type=sha,format=long
- id: build
uses: docker/build-push-action@v7
with:
context: .
platforms: linux/amd64,linux/arm64
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=gha
cache-to: type=gha,mode=max
provenance: mode=max
sbom: true
secrets: |
npmrc=${{ secrets.NPMRC }}
- uses: sigstore/cosign-installer@v4
- run: cosign sign --yes ghcr.io/${{ github.repository }}@${{ steps.build.outputs.digest }}
Points that bite people:
docker/build-push-actionexports the Actions cache environment for you; a baredocker buildx buildstep does not, sotype=ghasilently no-ops unless you addcrazy-max/ghaction-github-runtimeor setACTIONS_*yourself.- Sign
steps.build.outputs.digest. Do not re-resolve the tag. platforms:with QEMU onubuntu-latestis where the 10x tax lands. For compile-heavy builds, split into an amd64 job and anubuntu-24.04-armjob and assemble withimagetools create.- For multi-target repos,
docker/bake-action@v7runs yourdocker-bake.hcldirectly and keeps the build definition in the repo rather than the workflow.
See the GitHub Actions sheet for the surrounding workflow patterns.
Kubernetes-native with rootless BuildKit
apiVersion: batch/v1
kind: Job
metadata:
name: build-app
spec:
template:
spec:
restartPolicy: Never
containers:
- name: buildkit
image: moby/buildkit:v0.33.0-rootless
command: [buildctl-daemonless.sh]
args:
- build
- --frontend=dockerfile.v0
- --local=context=/src
- --local=dockerfile=/src
- --opt=build-arg:VERSION=1.2.3
- --output=type=image,name=registry.example.com/app:1.2.3,push=true
- --export-cache=type=registry,ref=registry.example.com/app:cache,mode=max
- --import-cache=type=registry,ref=registry.example.com/app:cache
env:
- { name: BUILDKITD_FLAGS, value: --oci-worker-no-process-sandbox }
- { name: DOCKER_CONFIG, value: /home/user/.docker }
securityContext:
seccompProfile: { type: Unconfined }
appArmorProfile: { type: Unconfined }
runAsUser: 1000
runAsGroup: 1000
volumeMounts:
- { name: src, mountPath: /src }
- { name: regcred, mountPath: /home/user/.docker }
volumes:
- { name: src, emptyDir: {} }
- name: regcred
secret:
secretName: regcred
items: [{ key: .dockerconfigjson, path: config.json }]
buildctl-daemonless.sh starts a buildkitd for the life of one build, which suits a Job. For a persistent build service, run deployment+service.rootless.yaml from moby/buildkit/examples/kubernetes/ and point a remote or kubernetes buildx builder at it.
Quick Reference
# --- buildx -------------------------------------------------------------
docker buildx ls # builders, nodes, platforms
docker buildx create --name ci --driver docker-container --bootstrap --use
docker buildx inspect ci # platforms, GC policy, worker labels
docker buildx rm ci # CI teardown — do not skip this
docker buildx build -t app:1 --push . # build and push
docker buildx build -t app:1 --load . # into the local daemon
docker buildx build -o type=local,dest=./out . # filesystem only, no image
docker buildx build -o type=oci,dest=./app.tar . # OCI layout
docker buildx build -o type=cacheonly . # verify only
docker buildx build --check . # lint the Dockerfile
docker buildx build --platform linux/amd64,linux/arm64 -t app:1 --push .
docker buildx build --cache-from type=registry,ref=app:cache \
--cache-to type=registry,ref=app:cache,mode=max .
docker buildx build --secret id=tok,env=TOKEN --ssh default .
docker buildx build --provenance=mode=max --sbom=true -t app:1 --push .
docker buildx imagetools inspect app:1
docker buildx imagetools inspect app:1 --format '{{ json .Provenance }}'
docker buildx imagetools create -t app:1 app:1-amd64 app:1-arm64
docker buildx bake --print # resolve without building
docker buildx bake --list=targets
docker buildx du && docker buildx prune --filter until=48h
# --- Buildah ------------------------------------------------------------
buildah build -t app:1 . # Dockerfile-compatible
buildah build --layers --format docker -t app:1 .
buildah unshare -- bash ./build.sh # user namespace for mount/chown
buildah from scratch # empty rootfs working container
buildah mount|umount|copy|add|run|config <ctr> ...
buildah commit --format oci --rm <ctr> app:1
buildah manifest create app:1 && buildah manifest push --all app:1 docker://reg/app:1
# --- the rest -----------------------------------------------------------
/kaniko/executor --context dir:///workspace --dockerfile Dockerfile \
--destination reg/app:1 --cache=true --cache-repo=reg/app/cache
stacker check && stacker build && stacker publish --url docker://reg --tag 1.2.3
melange keygen && melange build pkg.yaml --arch x86_64 --signing-key melange.rsa
apko build apko.yaml app:1 app.tar # config, tag, output — three args
KO_DOCKER_REPO=reg/org ko build ./cmd/app --platform=linux/amd64,linux/arm64
| Variable | Effect |
|---|---|
SOURCE_DATE_EPOCH |
Forwarded by buildx as a build arg; honoured by apko, ko, Buildah 1.41+ |
BUILDX_NO_DEFAULT_ATTESTATIONS |
Suppress the default provenance attestation |
BUILDKIT_SYNTAX |
Override the Dockerfile frontend without editing the file |
BUILDAH_FORMAT |
docker or oci (default) |
BUILDAH_ISOLATION |
oci, rootless (default when unprivileged), chroot |
BUILDAH_LAYERS |
true to keep intermediate layers |
KO_DOCKER_REPO |
ko's target repo; ko.local / kind.local side-load instead of pushing |
KO_DATA_DATE_EPOCH |
Modtime for ko's kodata static assets |
Common Issues and Solutions
| Issue | Cause | Fix |
|---|---|---|
| Multi-arch build takes an hour on CI | QEMU emulation on compute-bound steps — measured ~10x on Apple Silicon, ~5x on a Linux/amd64 runner; multiples either way | Native builder nodes, or cross-compile on $BUILDPLATFORM and copy the artefact into the target stage |
| Same build is fast locally, glacial in CI | Apple Silicon uses Rosetta for amd64 (~1.2x); Linux runners use QEMU (5–12x depending on host and guest arch) | Do not benchmark emulation on a Mac and extrapolate |
--cache-to type=registry silently does nothing |
Default docker driver on the classic overlay2 image store |
Use docker-container, or enable the containerd image store |
--load rejects a multi-platform build |
Classic image store cannot hold an index | Enable the containerd image store, or --load one platform at a time |
| Images "disappeared" after enabling containerd | The two stores are separate; nothing is migrated | Re-pull, or rebuild. Check with docker info -f '{{ .DriverStatus }}' |
type=gha cache no-ops in a bare docker buildx build step |
ACTIONS_RUNTIME_TOKEN / ACTIONS_RESULTS_URL are not exported to plain steps |
Use docker/build-push-action, or crazy-max/ghaction-github-runtime |
Deploy tool chokes on unknown/unknown manifests |
Default provenance attestations sit in the index | Update the tool, or --provenance=false / BUILDX_NO_DEFAULT_ATTESTATIONS=1 |
Secret visible in docker history |
ARG/ENV values persist in the image config |
RUN --mount=type=secret plus --secret id=...,src= or ,env= |
| Secret visible in the pushed attestation | --provenance=mode=max records every build arg verbatim |
Same fix. mode=min does not record args, but is not a substitute for secret mounts |
RUN --mount=type=cache is rejected |
No # syntax= directive, so an old bundled frontend is used |
Add # syntax=docker/dockerfile:1 as the first line |
Buildah: cannot mount using driver overlay in rootless mode |
buildah mount needs a user namespace you own |
Wrap the script in buildah unshare |
Buildah: lchown ... invalid argument, or storage silently on vfs |
Missing /etc/subuid and /etc/subgid ranges, or no fuse-overlayfs |
usermod --add-subuids 100000-165535 --add-subgids 100000-165535 "$USER", install fuse-overlayfs, then podman system migrate |
Buildah cross-arch build: Exec format error |
No binfmt handlers registered on the host | Register QEMU handlers, or build on a native host and use buildah manifest |
Buildah --source-date-epoch not recognised |
Added in 1.41.0 | Use --timestamp <epoch> on versions < 1.41 |
buildah build --squash removes more than expected |
Buildah squashes base-image layers too; podman build --squash does not |
Use --layers plus explicit staging if you want the weaker behaviour |
| Kaniko destroyed a build machine's filesystem | It unpacks the base image into its own /; run outside a container with --force and that is the host |
Only ever run the official image in a disposable pod. Never --force outside gVisor |
| Kaniko cache hit rate collapses mid-build | Cache reads stop after the first miss | Reorder so volatile steps come last, or accept it — the project is archived, migrate |
melange: bwrap: Creating new namespace failed |
Nested user namespaces blocked inside a rootless container | Run melange on the host, or in a container that permits nested userns |
apko image has no shell and docker exec fails |
"Distroless by construction" — nothing is installed that you did not name | Add busybox to contents.packages for a debug variant, and keep it out of production |
| Two identical builds produce different digests | Fresh timestamps in every layer | SOURCE_DATE_EPOCH and rewrite-timestamp=true, and turn provenance off or accept a moving index digest |
| Reproducible builder, irreproducible image | Unpinned apt-get/pip/npm in the Dockerfile |
Pin the base by digest and dependencies by lockfile. The builder is not the problem |
| CI job with the Docker socket "isn't privileged" | Socket access is root-equivalent on the host | Rootless BuildKit or Buildah, or a remote builder reachable only via the build API |
cosign 3.x: --tlog-upload=false is not supported with --signing-config |
cosign 3 deprecated the flag in favour of signing configs | Supply a signing config with no Rekor services, via cosign signing-config create |
Related Topics
The following topics complement this cheatsheet and would be valuable additions:
- Container Image Optimisation - The Dockerfile half of the story: multi-stage builds, layer ordering, cache mounts, base image selection, and size/security work
- Cloud Native Buildpacks - The no-Dockerfile model: the CNB lifecycle,
pack, builders, and rebasing an OS layer without rebuilding the application - Podman - Rootless containers and the engine that uses Buildah as its build library;
podman build, storage, and user-namespace mechanics - Container Registries - Where the output goes: manifest lists, OCI media types, garbage collection, and registry-side authentication and policy
- GitHub Actions - Workflow structure, OIDC, caching, and matrix strategy around the build steps shown here
- Grype and Syft - SBOM generation and vulnerability scanning against the images and attestations these builders produce