Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Container Image Building

Choosing and driving an image builder: BuildKit/buildx, Buildah, Kaniko, Stacker, apko/melange, and ko, with rootless, multi-arch, and reproducibility trade-offs.

Container Image Building

Choosing an image builder and driving it: BuildKit/buildx, Buildah, Kaniko, Stacker, apko/melange, and ko.

Overview

An OCI image is a JSON config, a set of layer blobs, and a manifest that ties them together. Any tool that can emit those three things is an image builder, and the six here differ almost entirely in how they decide what goes into the layers and what privileges they need to do it.

This sheet is about the builders. Dockerfile technique — multi-stage builds, layer ordering, cache mounts, base image selection, .dockerignore, size and hardening work — is covered in depth in the Container Image Optimisation sheet and is not repeated here. Dockerfiles appear below only as far as they demonstrate a builder's behaviour.

A seventh model sits deliberately outside this comparison. Cloud Native Buildpacks inverts the question: you supply a source tree and no build definition at all, and a chain of buildpacks detects the language, builds it, and assembles the image against a base image the platform owns rather than you. That ownership split is the whole point — it is what makes pack rebase able to patch an OS CVE without rebuilding the application — and it makes buildpacks incomparable row-by-row with the tools here. It has its own Cloud Native Buildpacks sheet; the decision tree below shows where it displaces everything on this page.

DockerfileBuildKit + buildxdaemon orstandalone, rootBuildahdaemonless, rootlessKaniko (archived)daemonless, needs acontainerShell scriptstacker.yamlStackerdaemonless,unprivilegedapko.yaml + melangeapkodaemonless, no rootGo sourcekono build file at allAny source treeCloud NativeBuildpackslifecycle detectsand buildsOCI image:config + layers +manifestDockerfileBuildKit + buildxdaemon orstandalone, rootBuildahdaemonless, rootlessKaniko (archived)daemonless, needs acontainerShell scriptstacker.yamlStackerdaemonless,unprivilegedapko.yaml + melangeapkodaemonless, no rootGo sourcekono build file at allAny source treeCloud NativeBuildpackslifecycle detectsand buildsOCI image:config + layers +manifest

The privilege question is the one that decides most real arguments:

No, Go binaryNo, and no buildfile at allNo, declarativeYesYesNo, want layercontrolYesNo, but rootlessLinuxNo, unprivileged K8spodarchived June 2025Do you need aDockerfile?ko: no daemon, noroot, cross-compilesCloud NativeBuildpacks: packbuild, platform ownsthe base imageReproducibility thehard requirement?Is a Docker daemonavailable?apko + melange:bitwisereproducible,APK-basedStacker: YAML,unprivileged,OCI-layout outbuildx: defaultdriver, ordocker-container forfull featuresBuildah: rootless,daemonless, scriptedor Dockerfilebuildkitd rootlesspod + buildctlKanikoNo, Go binaryNo, and no buildfile at allNo, declarativeYesYesNo, want layercontrolYesNo, but rootlessLinuxNo, unprivileged K8spodarchived June 2025Do you need aDockerfile?ko: no daemon, noroot, cross-compilesCloud NativeBuildpacks: packbuild, platform ownsthe base imageReproducibility thehard requirement?Is a Docker daemonavailable?apko + melange:bitwisereproducible,APK-basedStacker: YAML,unprivileged,OCI-layout outbuildx: defaultdriver, ordocker-container forfull featuresBuildah: rootless,daemonless, scriptedor Dockerfilebuildkitd rootlesspod + buildctlKaniko

Version baseline for this sheet

Component Version verified against
Docker Engine / CLI 29.7.2 (containerd image store, overlayfs snapshotter)
buildx v0.36.1 local; v0.37.0 current upstream
BuildKit v0.32.2 local; v0.33.0 current upstream
Dockerfile frontend docker/dockerfile:1 → 1.27.0
Buildah 1.39.3 on Debian 13 (trixie); 1.45.0 current upstream
Podman 5.4.2
Kaniko v1.24.0 — final release; repo archived 2025-06-03
Stacker v1.2.0 binary (v1.2.1 tagged 2026-08-25)
apko / melange v1.2.43 / v0.59.4
ko v0.19.1
cosign v3.1.3

Choosing a Builder

This is the table the rest of the sheet elaborates. It is opinionated on purpose.

buildx / BuildKit Buildah Kaniko Stacker apko + melange ko
Input Dockerfile (+ pluggable frontends) Dockerfile or shell script driving buildah verbs Dockerfile stacker.yaml apko.yaml (image) + melange YAML (packages) Go source tree, no build file
Daemon Yes for the docker driver; docker-container/kubernetes/remote run buildkitd themselves None None None None None
Root Daemon runs as root; rootless buildkitd supported No No, but wants to be inside a container No No No
Rootless story Good but fiddly: moby/buildkit:*-rootless image, --oci-worker-no-process-sandbox, unconfined seccomp/AppArmor Best in class. Native user namespaces, buildah unshare, works out of the box given subuid/subgid Poor. Unpacks the base image into its own filesystem; needs an isolated container to be safe Good. User namespaces, stacker unpriv-setup for subuid/subgid Excellent. apko needs nothing; melange's bubblewrap runner needs unprivileged userns Excellent. Just a Go toolchain
Cache Best in class: content-addressed DAG, inline/registry/local/gha/s3/azblob exporters, RUN --mount=type=cache --layers (off by default), --cache-from/--cache-to to a registry, --cache-ttl --cache + --cache-repo; stops reading cache after the first miss Content-hashed layers in .stacker/, incremental rebuild on input change APK resolution cache; apko lock pins the package set Go build cache; base image pulled once
Multi-arch --platform with QEMU emulation, or native builder nodes appended to one builder; imagetools create after the fact --platform (needs host binfmt), --manifest for lists; no built-in emulation None. --custom-platform only relabels; stitch with an external tool Per-arch builds, no built-in manifest assembly Native. archs: list, no emulation — it is just unpacking APKs Native. Go cross-compiles; no emulation ever
Reproducibility Achievable: SOURCE_DATE_EPOCH + rewrite-timestamp=true. Not the default --timestamp; --source-date-epoch/--rewrite-timestamp from 1.41 --reproducible flag, patchy in practice Content-hashed layers, hash: pinning on imports, --require-hash Bitwise reproducible by design — verified below Yes, with SOURCE_DATE_EPOCH — verified below
SBOM / provenance SLSA provenance on by default (mode=min); --sbom=true uses the BuildKit Syft scanner --sbom preset scanners (1.35+), pluggable scanner image None Embeds the whole stacker.yaml as an OCI annotation; bom: key Per-package SBOM inside each APK, composed into an image SBOM SPDX by default (--sbom=spdx)
Maturity The reference implementation. Actively developed, huge surface Mature, stable, Red Hat-backed; Podman's build engine Archived. Do not start new work on it Real project, tiny community (~370 stars); latest tag shipped no binaries Very active, Chainguard-backed; the model is unusual and the ecosystem is Wolfi-shaped Active, CNCF Sandbox; deliberately narrow

Short version. If you have a Dockerfile and a daemon, use buildx. If you need rootless and daemonless on Linux, use Buildah. If you need reproducibility as a contractual property, use apko/melange. If it is a Go binary, use ko and stop thinking about it. Do not adopt Kaniko in 2026. Stacker is a good tool with a very small blast radius of users — evaluate it only if its layer-control and unprivileged story is what you specifically need.

Where buildpacks displace all of this. Every tool in the table above assumes you want to own the base image and the build definition. If you do not — if you would rather hand over a source tree and have someone else's curated base image, language detection, and CVE patching applied for you — then Cloud Native Buildpacks is a different answer to the question, not a seventh column in it. The trade is control for operational leverage: you give up precise say over the final image and gain pack rebase, which swaps the OS layer under an unchanged application layer in seconds. See the Cloud Native Buildpacks sheet.


BuildKit and buildx

BuildKit is the build engine; docker buildx is the CLI in front of it. BuildKit has been the default since Docker 23, and docker build is now an alias for docker buildx build (along with docker builder build and docker image build). The classic builder has not been removed — DOCKER_BUILDKIT=0 docker build . still falls back to it on Engine 29.7.2, complete with Successfully built <id> output. Nothing in this sheet works there.

Builder Instances and Drivers

A builder is a named handle onto one or more BuildKit nodes. The driver decides where buildkitd lives.

Driver Where BuildKit runs Notes
docker Inside the Docker daemon Default. Cannot be created manually; one per Docker context
docker-container A container the driver manages The workhorse. Full feature set, isolated cache
cloud Docker Build Cloud Managed, paid
kubernetes Pods in a cluster Deployment or StatefulSet; scales, supports native multi-arch nodes
remote A buildkitd you run yourself tcp:// or unix://, TLS via driver-opts
docker buildx ls
# NAME/NODE           DRIVER/ENDPOINT     STATUS    BUILDKIT   PLATFORMS
# desktop-linux*      docker
#  \_ desktop-linux    \_ desktop-linux   running   v0.32.2    linux/amd64 (+2), linux/arm64, ...

docker buildx create --name ci --driver docker-container \
  --driver-opt network=host --bootstrap --use
docker buildx inspect ci          # BuildKit version, platforms, GC policy, worker labels
docker buildx rm ci               # also removes its cache volume

# Kubernetes driver: BuildKit pods in a namespace
docker buildx create --name k8s --driver kubernetes \
  --driver-opt namespace=buildkit,replicas=3,rootless=true \
  --driver-opt requests.cpu=1,requests.memory=2Gi,limits.memory=4Gi,loadbalance=sticky

# Remote driver: an externally managed buildkitd
docker buildx create --name remote --driver remote tcp://buildkitd.internal:1234 \
  --driver-opt cacert=/certs/ca.pem,cert=/certs/cert.pem,key=/certs/key.pem

Common docker-container driver-opts: image, network, cgroup-parent (default /docker/buildx), restart-policy (default unless-stopped), memory, cpu-quota, cpu-shares, cpuset-cpus, default-load (default false), env.<KEY>, provenance-add-gha (default true).

Naming note. --buildkitd-config was --config before buildx v0.13.0 (March 2024). --config still works but is hidden. On versions < 0.13 use --config.

The docker Driver's Limits — and How the containerd Image Store Changed Them

The classic advice is that the default docker driver cannot do multi-platform, cache export, or attestations. That advice is now conditional on the image store, and the condition flipped under most people's feet.

# Which image store is the daemon using?
docker info -f '{{ .DriverStatus }}'
# [[driver-type io.containerd.snapshotter.v1]]   <- containerd store
# [[Backing Filesystem extfs] ...]               <- classic overlay2 store

With the containerd image store, the docker driver supports the inline, local, registry, and gha cache backends, multi-platform builds, --load of a multi-platform image, and attestations. With the classic overlay2 store none of that works, and the failures are not always loud.

The containerd store is the default for fresh installs of Docker Engine 29.0 (2025-11-10) and for Docker Desktop 4.34+. Upgrades keep overlay2 until you switch:

// /etc/docker/daemon.json
{
  "features": { "containerd-snapshotter": true }
}

Switching does not migrate existing images — they stay in the old store and appear to vanish. That alone is why so many CI images are still on overlay2.

The buildx driver feature matrix in the docs still lists Multi-arch images as unsupported for the docker driver with no containerd footnote. That row is stale; the containerd storage page and observed behaviour both contradict it. Verified locally on Engine 29.7.2 with the containerd store: --platform linux/amd64,linux/arm64 builds, pushes a proper OCI index, and --loads.

For CI, use docker-container anyway. It gives you a predictable feature set regardless of what the runner's daemon is configured with, and an isolated cache you can size and garbage-collect.

Building in Anger

docker buildx build --builder ci -t registry.example.com/app:1.2.3 --push .

# Named stage, extra context, build args, labels, index annotations, metadata out
docker buildx build --builder ci \
  --target runtime \
  --build-context vendored=./third_party \
  --build-arg VERSION=1.2.3 \
  --label org.opencontainers.image.revision="$(git rev-parse HEAD)" \
  --annotation "index:org.opencontainers.image.source=https://git.example.com/app" \
  --metadata-file /tmp/meta.json \
  -t registry.example.com/app:1.2.3 --push .

--check (buildx 0.15+, Dockerfile frontend 1.8+) lints without building and exits non-zero on any violation, where a normal build only warns — worth wiring into CI:

docker buildx build --check -f Dockerfile.lint .
# Check complete, 1 warning has been found!
#
# WARNING: FromAsCasing - https://docs.docker.com/go/dockerfile/rule/from-as-casing/
# 'as' and 'FROM' keywords' casing do not match
# Dockerfile.lint:1

Suppress rules in the Dockerfile rather than on the CLI, so the exception travels with the code. Both settings go on one directive, separated by a semicolon — a second # check= line is rejected with ERROR: only one check parser directive can be used:

# check=skip=JSONArgsRecommended,StageNameCasing;error=true

Output Types

--output (-o) decides what BuildKit does with the result. Seven exporters; the default is cacheonly except on the docker driver, which defaults to the docker exporter.

type= What you get
image Image in the builder's store; push=true sends it onward
registry Shorthand for image with push=true
docker A Docker-format tarball, or loaded into the daemon
oci An OCI image layout (tar or directory)
local The filesystem unpacked to a directory — no image at all
tar The filesystem as a tarball
cacheonly Nothing but cache. Ideal for CI verification builds
docker buildx build -t app:1 --push .             # = --output=type=registry,unpack=false (0.31+)
docker buildx build -t app:1 --load .             # = --output=type=docker
docker buildx build --target export -o type=local,dest=./dist .   # artefact, no image
docker buildx build -o type=oci,dest=./app.oci.tar .              # feed to skopeo/crane
tar tf app.oci.tar | head -3
# blobs/
# blobs/sha256/
# blobs/sha256/738128faa30f570583b0e57efd831e0e6a2a9aacf1be88c8f4c1ef8a5b7033cc
docker buildx build -o type=cacheonly .           # verify without producing anything

local and tar split multi-platform output into per-platform subdirectories; platform-split=false merges them. That option is local/tar only — it does not exist on image or oci.

Cache Exporters and Importers

Six backends. mode=min (default) exports only the layers in the final image; mode=max exports every intermediate stage, which is what you want for multi-stage builds.

type= Use Driver support
inline Cache metadata embedded in the image itself docker OK with containerd store
registry Cache as a separate image, usually a :cache tag docker OK with containerd store
local A directory on disk Needs a non-default driver
gha GitHub Actions cache docker OK with containerd store
s3 An S3 bucket Needs a non-default driver
azblob Azure Blob Storage Needs a non-default driver
# Registry cache, max mode — the default choice for CI
docker buildx build --builder ci \
  --cache-from type=registry,ref=registry.example.com/app:cache \
  --cache-to   type=registry,ref=registry.example.com/app:cache,mode=max \
  -t registry.example.com/app:1.2.3 --push .

# Local directory (self-hosted runners with a persistent volume)
docker buildx build --cache-from type=local,src=/var/cache/buildkit \
                    --cache-to   type=local,dest=/var/cache/buildkit,mode=max .

# S3
docker buildx build --cache-to type=s3,region=eu-west-2,bucket=build-cache,name=app,mode=max \
                    --cache-from type=s3,region=eu-west-2,bucket=build-cache,name=app .

# GitHub Actions
docker buildx build --cache-from type=gha,scope=main --cache-to type=gha,scope=main,mode=max .

inline embeds cache metadata in the image, with no separation between artefact and cache, and does not scale to multi-stage builds because intermediate stages are not in the final image. Use registry with mode=max unless you genuinely have one stage.

GitHub Actions cache. GitHub retired the v1 cache service on 2025-04-15. buildx detects the v2 service from $ACTIONS_CACHE_SERVICE_V2 and switches automatically (from buildx v0.21.0); force it with version=2. url/token default to $ACTIONS_RESULTS_URL and $ACTIONS_RUNTIME_TOKEN, which docker/build-push-action sets for you but a bare docker buildx build step does not.

Multi-architecture Builds

# Emulation: one node, binfmt_misc handles foreign binaries
docker run --privileged --rm tonistiigi/binfmt --install all
docker buildx build --platform linux/amd64,linux/arm64,linux/arm/v7 -t app:1 --push .

# Native nodes: append a second machine to the same builder
docker buildx create --name multi --driver docker-container \
  --platform linux/amd64 --node amd64 tcp://amd64-runner:2376
docker buildx create --append --name multi --driver docker-container \
  --platform linux/arm64 --node arm64 tcp://arm64-runner:2376
docker buildx build --builder multi --platform linux/amd64,linux/arm64 -t app:1 --push .

Which of those to reach for, and the numbers behind the choice, are in Multi-architecture below.

imagetools and Manifest Lists

imagetools works on registries directly — no local image store involved.

docker buildx imagetools inspect registry.example.com/app:1
# Name:      registry.example.com/app:1
# MediaType: application/vnd.oci.image.index.v1+json
# Digest:    sha256:dbb89b78081e...
#
# Manifests:
#   Name:      ...@sha256:7c190ed5f921...  Platform: linux/amd64
#   Name:      ...@sha256:096db7bb4ebd...  Platform: linux/arm64
#   Name:      ...@sha256:9c7eab0e4ed5...  Platform: unknown/unknown
#   Annotations:
#     vnd.docker.reference.type:   attestation-manifest

docker buildx imagetools inspect --raw registry.example.com/app:1   # raw JSON

# Assemble a manifest list from per-arch tags built on separate machines
docker buildx imagetools create -t registry.example.com/app:1 \
  --annotation "index:org.opencontainers.image.source=https://git.example.com/app" \
  registry.example.com/app:1-amd64 registry.example.com/app:1-arm64

Those unknown/unknown entries are not corruption — they are the attestation manifests, addressed by the vnd.docker.reference.digest annotation pointing back at the runnable manifest they describe. Old tooling that iterates the index and tries to pull every entry chokes on them; that is the usual cause of "my deploy tool suddenly can't read my image".

bake: Multi-target Builds

bake is the build equivalent of a Makefile: a declarative set of targets, groups, and variables that one command builds concurrently.

# docker-bake.hcl
variable "REGISTRY" { default = "registry.example.com" }
variable "TAG"      { default = "dev" }

group "default" { targets = ["api", "worker"] }

target "_common" {
  context    = "."
  dockerfile = "Dockerfile"
  platforms  = ["linux/amd64", "linux/arm64"]
  cache-from = ["type=registry,ref=${REGISTRY}/cache:build"]
  cache-to   = ["type=registry,ref=${REGISTRY}/cache:build,mode=max"]
}

target "api"    { inherits = ["_common"], target = "api",    tags = ["${REGISTRY}/api:${TAG}"] }
target "worker" { inherits = ["_common"], target = "worker", tags = ["${REGISTRY}/worker:${TAG}"] }

# matrix forks one target into variants; it requires a templated name
target "app" {
  name   = "app-${tgt}"
  matrix = { tgt = ["debug", "release"] }
  target = tgt
  tags   = ["app:${tgt}"]
}
docker buildx bake --print              # resolve the full plan without building — do this first
docker buildx bake --list=targets       # also --list=variables
# TARGET   DESCRIPTION
# api
# default  api, worker
# worker

docker buildx bake --builder ci --push  # build the default group
docker buildx bake --set '*.platform=linux/amd64' --set 'api.tags=api:hotfix' api
docker buildx bake --var TAG=1.2.3 --push                  # --var is buildx 0.31+
docker buildx bake "https://github.com/org/repo.git#main"  # remote definition

Files are searched in this order: compose.yaml, compose.yml, docker-compose.yml, docker-compose.yaml, docker-bake.json, docker-bake.hcl, docker-bake.override.json, docker-bake.override.hcl. Compose files are valid bake input directly — the cheapest way to get a multi-service build out of an existing compose.yaml.

Attestations: Provenance and SBOM

SLSA provenance at mode=min is added by default — has been since buildx v0.10. You do not opt in; you opt out.

docker buildx build -t app:1 --push .                              # min provenance, no SBOM
docker buildx build --provenance=mode=max --sbom=true -t app:1 --push .
docker buildx build --attest type=sbom,generator=<scanner-image> -t app:1 --push .
docker buildx build --provenance=false --sbom=false -t app:1 --push .
export BUILDX_NO_DEFAULT_ATTESTATIONS=1                            # same, environment-wide

docker buildx imagetools inspect app:1 --format '{{ json .Provenance }}'
docker buildx imagetools inspect app:1 --format '{{ json .SBOM.SPDX }}'
# spdxVersion SPDX-2.3 | name sbom
# creators ['Organization: Anchore, Inc', 'Tool: syft-v1.51.0', 'Tool: buildkit-v0.32.2']
docker buildx imagetools inspect app:1 --format '{{ json (index .SBOM "linux/amd64").SPDX }}'

To scan the build context and named stages as well as the final image, add real ARG instructions to the Dockerfile — these cannot be set from the environment:

ARG BUILDKIT_SBOM_SCAN_CONTEXT=true
ARG BUILDKIT_SBOM_SCAN_STAGE=build,runtime

Footgun. --provenance=mode=max records every build argument verbatim in the pushed attestation. Verified: a build run with --build-arg API_TOKEN=s3cr3t-do-not-ship produced provenance containing "build-arg:API_TOKEN": "s3cr3t-do-not-ship" under buildDefinition.externalParameters.request.args. The default mode=min does not. See Build Secrets below.

Frontends and the # syntax= Directive

BuildKit does not natively understand Dockerfiles. It runs a frontend — a container image — that compiles the input into BuildKit's LLB graph. That is why Dockerfile features can ship independently of your Docker version.

# syntax=docker/dockerfile:1

Channels: docker/dockerfile:1 tracks the latest stable v1 release (currently 1.27.0) and takes minor and patch updates; docker/dockerfile:1-labs adds experimental instructions. Pin a full version (:1.27.0) plus digest if you need bit-stability.

# Same effect without touching the Dockerfile
docker buildx build --build-arg BUILDKIT_SYNTAX=docker/dockerfile:1 .

# A completely custom frontend
# syntax=registry.example.com/frontends/mylang:2@sha256:abc...

If you use RUN --mount=type=cache, RUN --mount=type=secret, heredocs, or COPY --parents, you need the directive. Without it you get whatever frontend is baked into the daemon, which on older hosts silently rejects the syntax.

buildkitd Standalone in Kubernetes

For a build service not tied to a Docker daemon, run buildkitd yourself and drive it with buildctl or a remote buildx builder. Upstream ships example manifests in examples/kubernetes/: pod.rootless.yaml, deployment+service.rootless.yaml, statefulset.rootless.yaml, plus .privileged. and .userns. variants of each, and create-certs.sh for TLS.

apiVersion: v1
kind: Pod
metadata:
  name: buildkitd
spec:
  containers:
    - name: buildkitd
      image: moby/buildkit:v0.33.0-rootless
      args: [--oci-worker-no-process-sandbox]
      securityContext:
        seccompProfile:  { type: Unconfined }
        appArmorProfile: { type: Unconfined }
        runAsUser: 1000
        runAsGroup: 1000
buildctl --addr kube-pod://buildkitd build \
  --frontend dockerfile.v0 --local context=. --local dockerfile=. \
  --output type=image,name=registry.example.com/app:1,push=true

docker buildx create --name k8s-remote --driver remote kube-pod://buildkitd

The rootless variant needs unconfined seccomp and AppArmor because it uses nested user namespaces; hardened clusters that forbid unconfined profiles cannot run it. Note the manifests now use the securityContext.seccompProfile / appArmorProfile fields — the old container.apparmor.security.beta.kubernetes.io/... annotation form is stale. There is no official Helm chart or operator.


Buildah

Buildah is what Podman uses to build images — podman build calls Buildah's Go API. As a CLI it does two distinct things, and the second is the reason to reach for it.

Dockerfile-compatible Builds

buildah build -t app:1 .                       # `bud` is a retained alias
buildah build --layers -t app:1 .              # keep intermediate layers (off by default)
buildah build --format docker -t app:1 .       # Docker v2s2 manifest instead of OCI
buildah build --platform linux/arm64 -t app:1 .
buildah build --jobs 4 -t app:1 .              # parallel stages
buildah build --secret id=npmrc,src=$HOME/.npmrc -t app:1 .
buildah build --ssh default -t app:1 .
buildah build --cache-to registry.example.com/app:cache \
              --cache-from registry.example.com/app:cache --cache-ttl 24h -t app:1 .

Defaults, verified on 1.39.3 (Debian 13, rootless):

Flag Default Override
--format oci BUILDAH_FORMAT=docker
--isolation rootless for unprivileged users, oci for root BUILDAH_ISOLATION
--layers false BUILDAH_LAYERS=true
--jobs 1 —

--squash on buildah build squashes all layers including the base image's into one — stronger than podman build --squash, which only squashes the layers the build itself added. The flag name is the same in both tools and the semantics are not. There is no --squash-all in Buildah.

The Scripted, Dockerfile-free Build

This is what makes Buildah different. Instead of a declarative file you get shell verbs against a working container: no parser, no frontend, no DSL — just the filesystem and buildah config.

#!/usr/bin/env bash
set -euo pipefail

# Harvest a static binary from a published image
src=$(buildah from docker.io/library/busybox:1.37-musl)
srcmnt=$(buildah mount "$src")

# A scratch container is an empty rootfs — no base layer at all
ctr=$(buildah from scratch)
mnt=$(buildah mount "$ctr")

mkdir -p "$mnt/bin" "$mnt/etc"
cp "$srcmnt/bin/busybox" "$mnt/bin/busybox"
ln -sf busybox "$mnt/bin/sh"
printf 'app:x:1000:1000::/:/bin/sh\n' > "$mnt/etc/passwd"
printf 'app:x:1000:\n'                > "$mnt/etc/group"

buildah umount "$src" >/dev/null && buildah rm "$src" >/dev/null

buildah config \
  --entrypoint '["/bin/sh"]' \
  --cmd        '["-c","echo hello from scratch"]' \
  --user 1000:1000 --workingdir / --env APP_ENV=prod \
  --label org.opencontainers.image.title=scripted-demo \
  "$ctr"

buildah umount "$ctr" >/dev/null
buildah commit --format oci --rm "$ctr" scripted-demo:1

Run it through buildah unshare so the mount happens inside your own user namespace:

buildah unshare -- bash ./build.sh
# 3bcbc8c0c6af73a88bb50ff558eb40faf22c98e96385253f43782812072e5012

buildah images --format '{{.Name}}:{{.Tag}}  {{.Size}}'
# localhost/scripted-demo:1   1.22 MB
buildah inspect --type image --format '{{len .OCIv1.RootFS.DiffIDs}}' scripted-demo:1
# 1                                       <- one layer, one history entry
podman inspect scripted-demo:1 --format '{{.ManifestType}}'
# application/vnd.oci.image.manifest.v1+json
podman run --rm scripted-demo:1
# hello from scratch

Why bother: exact control over layer count and contents, any host tool available to populate the rootfs (a package manager, rsync, a Python script), and no Dockerfile grammar between you and the filesystem. The classic use is assembling minimal images from artefacts a separate pipeline already produced.

buildah config takes JSON array form for --entrypoint and --cmd; a bare string is split on whitespace, which is rarely what you want. Other flags: --env, --label, --annotation, --port, --user, --workingdir (one word, no hyphen), --volume, --author, --created-by, --arch, --os, --shell, --stop-signal, --healthcheck (plus --healthcheck-interval, --healthcheck-retries, --healthcheck-timeout, --healthcheck-start-period), --history-comment, --onbuild, --unsetlabel, and --unsetannotation (1.41+; unknown flag on 1.39.3).

Rootless Mechanics

Buildah's rootless story is the best of the six, and it hinges on user namespaces.

grep "^$(id -un)" /etc/subuid /etc/subgid     # ranges must exist, or nothing works
# /etc/subuid:mike:165536:65536
# /etc/subgid:mike:165536:65536
podman info --format '{{.Store.GraphDriverName}} {{.Store.GraphStatus}}'
# overlay map[Backing Filesystem:extfs Native Overlay Diff:true ...]

ctr=$(buildah from scratch)
buildah mount "$ctr"
# Error: cannot mount using driver overlay in rootless mode.
# You need to run it in a `buildah unshare` session

buildah build handles the namespace for you; anything touching the container's filesystem directly does not. That error is the single most common rootless Buildah failure, and the fix is in the message — buildah unshare drops you into a user namespace where your UID maps to root and the subuid range maps everything else, so mount, chown, and direct writes behave.

Rootless storage lives in ~/.local/share/containers/storage (root: /var/lib/containers/storage). On kernels ≥ 5.11 you get native rootless overlayfs (Native Overlay Diff: true above); older kernels fall back to fuse-overlayfs, and if that is missing you silently drop to the vfs driver, which copies the whole filesystem per layer and is glacial.

Output Format, Podman, and Skopeo

Buildah defaults to OCI manifests; Docker's tooling historically defaults to Docker v2s2, and some older registries and admission controllers still choke on OCI media types. buildah build --format docker (or BUILDAH_FORMAT=docker) switches.

Multi-arch is manual — build per platform, then assemble. Buildah has no built-in emulation, so without host binfmt handlers a cross-arch RUN fails immediately:

buildah manifest create app:1
buildah build --platform linux/amd64 --manifest app:1 -t app:1-amd64 .
buildah build --platform linux/arm64 --manifest app:1 -t app:1-arm64 .
buildah manifest push --all app:1 docker://registry.example.com/app:1

buildah build --platform linux/arm64 -t app:1 .          # no binfmt registered
# STEP 2/2: RUN echo hello > /hello.txt
# exec container process `/bin/sh`: Exec format error
# Error: building at STEP "RUN ...": while running runtime: exit status 1

Podman and Buildah share image storage but not container state ("you can not see Podman containers from within Buildah or vice versa"). Skopeo is the third of the trio, for copying and inspecting images between registries and stores without a daemon. See the Podman and Container Registries sheets.


Kaniko

Status: Archived

GoogleContainerTools/kaniko was archived on 2025-06-03 and is read-only. The README opens with a banner stating the project is archived and no longer developed or maintained. The final release is v1.24.0 (2025-05-23). There is no named successor, no new registry location, and ~760 issues frozen in place. It was never an officially supported Google product.

You will still find it recommended in blog posts, Helm charts, and internal CI templates written between 2019 and 2024, because for several years it was the only credible answer to "build an image in an unprivileged Kubernetes pod". It no longer is. Migrate to rootless buildkitd (see above) or Buildah in a pod.

It is documented here because you will inherit it, and because understanding its mechanism explains its constraints.

How It Actually Works

Kaniko has no daemon and does not use overlayfs. It extracts the base image into its own container's root filesystem, then executes each Dockerfile instruction in that filesystem, taking a userspace snapshot after each one and diffing to produce a layer.

INFO Unpacking rootfs as cmd RUN echo hello > /hello.txt requires it.
INFO Initializing snapshotter ...
INFO Taking snapshot of full filesystem...
INFO Running: [/bin/sh -c echo hello > /hello.txt]
INFO Taking snapshot of full filesystem...

That is the whole design, and every consequence follows from it:

  • It must run inside a disposable container. It writes to / in whatever it is running in. Outside a container it guards itself and refuses: kaniko should only be run inside of a container, run with the --force flag if you are sure you want to continue. --force is documented as an escape hatch for gVisor, where containerisation cannot be detected — not as a general blessing. Never use it on a machine whose filesystem you care about.
  • Run only the official image. Running the executor binary in some other image "is not supported due to implementation details" — it cannot chroot or bind-mount, so it will overwrite whatever is already there.
  • Full-filesystem snapshots are expensive. --snapshot-mode trades correctness for speed: full (default, hashes file contents), redo (mtime + size + permissions), time (mtime only, fastest, misses same-second changes).
  • No BuildKit features. No RUN --mount=type=cache, no --mount=type=secret, no heredocs, no # syntax= frontends — there is no BuildKit and no frontend. Note this is an inference from the archived feature set rather than an explicit statement in the README.

Running It

apiVersion: v1
kind: Pod
metadata:
  name: kaniko-build
spec:
  restartPolicy: Never
  containers:
    - name: kaniko
      image: gcr.io/kaniko-project/executor:v1.24.0
      args:
        - --context=git://github.com/org/repo.git#refs/heads/main
        - --dockerfile=Dockerfile
        - --destination=registry.example.com/app:1.2.3
        - --cache=true
        - --cache-repo=registry.example.com/app/cache
        - --snapshot-mode=redo
      volumeMounts:
        - { name: docker-config, mountPath: /kaniko/.docker }
  volumes:
    - name: docker-config
      secret:
        secretName: regcred
        items: [{ key: .dockerconfigjson, path: config.json }]

Credentials go in a Docker config JSON at /kaniko/.docker/config.json. Contexts: dir://, tar:// (including tar://stdin), git://<url>#<ref>#<commit>, gs://, s3://, and Azure Blob over https://.

# Local trial: no push, tarball out
podman run --rm -v "$PWD/ctx":/workspace -v "$PWD":/out \
  gcr.io/kaniko-project/executor:v1.24.0 \
  --context dir:///workspace --dockerfile /workspace/Dockerfile \
  --destination example.invalid/demo:1 --no-push --tar-path /out/image.tar

Other flags: --cache-ttl (default two weeks), --cache-dir (default /cache), --cache-copy-layers, --cache-run-layers (default true), --single-snapshot, --reproducible, --digest-file, --target, --build-arg, --ignore-path, --push-retry, --registry-mirror, --skip-tls-verify. A companion gcr.io/kaniko-project/warmer image pre-populates the base-image cache; the :debug tag adds a busybox shell.

Cache reads stop after the first miss — every layer after that rebuilds locally, even if a later one was cached. Multi-arch does not exist: --custom-platform (hyphenated; not --customPlatform) only changes the declared platform and, per the README, "is not virtualization and cannot help to build an architecture not natively supported by the build host". Stitch per-arch builds with imagetools create or manifest-tool.


Stacker

Stacker builds OCI images from a YAML file, unprivileged, using LXC-sealed build containers. Apache-2.0, Cisco-originated, and genuinely maintained — v1.2.1 landed 2026-08-25 — but small: ~370 stars, ~113 open issues, and the published docs site is still versioned at v1.0.0 while releases are at v1.2.1. The v1.2.1 tag shipped no release binaries (v1.2.0 did), so plan on pinning an older tag or building from source, which needs Go, lxc-devel, libcap, libacl, gpgme, and mksquashfs.

# stacker.yaml
build:
  build_only: true                       # produces artefacts, not an output image
  from:
    type: docker
    url: docker://docker.io/library/alpine:3.22
  run: |
    mkdir -p /out
    echo "built by stacker" > /out/hello.txt

demo:
  from:
    type: docker
    url: docker://docker.io/library/busybox:1.37-musl
  imports:
    - stacker://build/out/hello.txt      # pull an artefact from another layer
  run: |
    cp /stacker/imports/hello.txt /hello.txt
  entrypoint: /bin/sh
  cmd: ["-c", "cat /hello.txt"]
  environment:
    APP_ENV: prod
  labels:
    org.opencontainers.image.title: stacker-demo
stacker check                                # are the kernel features present?
stacker build
# preparing image build...
# + mkdir -p /out
# preparing image demo...
# + cp /stacker/imports/hello.txt /hello.txt
# filesystem demo built successfully
ls                                           # oci/  roots/  .stacker/  stacker.yaml

stacker build --layer-type squashfs          # tar (default), squashfs, or erofs
stacker build --substitute VERSION=1.2.3     # ${{VERSION}} in the YAML
stacker build --require-hash                 # refuse imports without a pinned sha256
stacker recursive-build -d ./images          # build every stacker.yaml under a tree
stacker publish --url docker://registry.example.com --tag 1.2.3
stacker unpriv-setup                         # one-time subuid/subgid configuration

Layer keys: from (type: one of docker, oci, tar, built, scratch), imports (plural — import is the deprecated legacy form), overlay_dirs, run, cmd, entrypoint, full_command, environment (not env), build_env, build_env_passthrough, volumes, labels, generate_labels, annotations, working_dir, runtime_user, build_only, binds, os, arch, bom.

Output is a real OCI layout in oci/, unpacked rootfs in roots/, cache in .stacker/. Unprivileged builds need /etc/subuid and /etc/subgid ranges (stacker unpriv-setup writes them) and kernel ≥ 5.11 for unprivileged overlay mounts. Verified on Debian 13 (kernel 6.12) as an ordinary user with only pre-existing subuid/subgid ranges — no setup step needed.

The provenance story is the interesting part: layers are content-hashed and rebuilt only when inputs change, imports can be pinned with a hash: sha256 (--require-hash enforces it fleet-wide), and the entire stacker.yaml is embedded in the image as an OCI annotation:

Annotations:
  io.stackeroci.stacker.stacker_version: v1.2.0
  io.stackeroci.stacker.stacker_yaml: |
    build:
      build_only: true
      ...

The recipe travels with the artefact, which is genuinely nice. There is no published guarantee of bit-identical layer digests, so do not promise one. Its niche is base images and OS-like layers where you want explicit layer control and no daemon; documented adopters are thin on the ground beyond Cisco.


apko and melange

This is the most different model in the set, and it is worth understanding even if you do not adopt it.

There is no RUN. An apko image is composed from a list of APK packages, not built by executing commands. Everything bespoke must first be packaged, and melange is what packages it. Removing arbitrary execution from image assembly is what buys the two properties apko claims: bitwise reproducibility, and complete SBOM coverage.

Your sourcemelange buildsigned .apk +APKINDEX.tar.gzUpstream Wolfi repoapko buildOCI image + SPDXSBOMYour sourcemelange buildsigned .apk +APKINDEX.tar.gzUpstream Wolfi repoapko buildOCI image + SPDXSBOM

apko: Composing the Image

# apko.yaml
contents:
  repositories: [https://packages.wolfi.dev/os]
  keyring:      [https://packages.wolfi.dev/os/wolfi-signing.rsa.pub]
  packages:
    - wolfi-baselayout
    - ca-certificates-bundle
    - busybox

accounts:
  groups: [{ groupname: app, gid: 65532 }]
  users:  [{ username: app, uid: 65532, gid: 65532 }]
  run-as: 65532

entrypoint:
  command: /bin/sh -l

environment:
  APP_ENV: prod

archs: [x86_64, aarch64]
# apko build <config> <tag> <output>   -- exactly three arguments
podman run --rm -v "$PWD":/work -w /work cgr.dev/chainguard/apko \
  build apko.yaml demo:latest demo.tar
ls
# apko.yaml  demo.tar  sbom-aarch64.spdx.json  sbom-index.spdx.json  sbom-x86_64.spdx.json

apko publish apko.yaml registry.example.com/demo:latest   # build straight to a registry

The schema's mixed casing matters. Hyphenated: shell-fragment, stop-signal, work-dir, run-as, vcs-url. Underscored: build_repositories, runtime_repositories, runtime_keyring. Multi-layer output is opt-in via layering: {strategy: origin, budget: 10}; the default is one layer. There is no top-level options key any more, and no --rewrite-tags flag.

Reproducibility — verified. Two independent builds of the same config, minutes apart:

sha256sum demo.tar demo2.tar
# 26a3326bde49c53ba3014693b32a7e8de7b7f2712b8978946df8c38ab92da48c  demo.tar
# 26a3326bde49c53ba3014693b32a7e8de7b7f2712b8978946df8c38ab92da48c  demo2.tar

Byte-identical, per-arch layer digests included. SOURCE_DATE_EPOCH is honoured and always overrides --build-date; when it is unset the default epoch is the build time of the newest installed APK, not the current time. What breaks reproducibility is the package set moving underneath you, so pin it — apko lock apko.yaml writes apko.yaml.lock.json, replayed with apko build --lockfile. (lock and resolve are hidden from --help pending feedback; resolve is deprecated in favour of lock.)

SBOMs are SPDX JSON written beside the output as sbom-<arch>.spdx.json plus sbom-index.spdx.json. They are not attached as OCI attestations — that is a separate cosign step. apko composes them from per-package SBOMs it finds inside each APK at /var/lib/db/sbom/, which is why coverage is complete rather than inferred by scanning. apko needs no root and no runner: where it cannot chown or create device nodes directly, it records the intent and writes correct ownership into the layer tar stream.

melange: Building the Packages

# hello.yaml
package:
  name: hello-demo
  version: 1.0.0
  epoch: 0
  description: A tiny demo package built by melange
  copyright:
    - license: Apache-2.0

environment:
  contents:
    repositories: [https://packages.wolfi.dev/os]
    keyring:      [https://packages.wolfi.dev/os/wolfi-signing.rsa.pub]
    packages:     [busybox, wolfi-baselayout]

pipeline:
  - name: Install the script
    runs: |
      mkdir -p "${{targets.destdir}}/usr/bin"
      install -m0755 hello "${{targets.destdir}}/usr/bin/hello"
melange keygen                       # melange.rsa + melange.rsa.pub, 4096-bit
melange build hello.yaml --arch x86_64 --signing-key melange.rsa --runner bubblewrap
# INFO wrote packages/x86_64/hello-demo-1.0.0-r0.apk
# INFO generating apk index from packages in packages/x86_64
# INFO signing index packages/x86_64/APKINDEX.tar.gz with key melange.rsa

Built-in pipelines cover most of the ground — fetch, git-checkout, patch, strip, autoconf/configure, autoconf/make, autoconf/make-install, cmake/configure, cmake/build, go/build, go/install, split/dev, split/manpages, split/debug — invoked as - uses: autoconf/make. Runners are bubblewrap, docker, and qemu, defaulting by platform. --out-dir defaults to ./packages/, --cache-dir to ./melange-cache/, --generate-index to true.

Feeding melange output back into apko uses a labelled local repository — a bare filesystem path, not a file:// URL:

contents:
  repositories:
    - https://packages.wolfi.dev/os
    - "@local /work/packages"
  keyring:
    - https://packages.wolfi.dev/os/wolfi-signing.rsa.pub
    - /work/melange.rsa.pub
  packages:
    - wolfi-baselayout
    - busybox
    - hello-demo@local          # pin this package to the @local repo

Verified end to end: melange build → apko build → podman run printed the packaged script's output, and the resulting image SBOM listed 43 packages including hello-demo.

What It Costs You

"Distroless by construction" is real — the image holds exactly the packages you named and their dependencies, with no shell, no apk, and no package database at runtime unless you add busybox and apk-tools yourself. Excellent for attack surface and SBOM fidelity.

The price is that your whole dependency graph has to exist as APKs. For anything Wolfi already ships that is free; for your own application it means writing and maintaining melange definitions, which is a packaging discipline most teams do not have. Dockerfile → apko is not a translation, it is a change of model. Budget accordingly.


ko

ko builds container images from Go source with no Dockerfile, no daemon, and no privileges. Deliberately narrow and, within that scope, unbeatable.

export KO_DOCKER_REPO=registry.example.com/org

ko build ./cmd/app                                  # build and push
ko build ./cmd/app --platform=linux/amd64,linux/arm64
ko build ./cmd/app --platform=all                   # every platform the base image has
ko build ./cmd/app --push=false --oci-layout-path=./oci
ko build ./cmd/app --local                          # side-load into the local daemon
ko build ./cmd/app --bare                           # image name == KO_DOCKER_REPO exactly
ko build ./cmd/app --sbom=spdx                      # default; only other value is "none"

ko apply -f config/                                 # resolve ko:// refs, then kubectl apply
ko resolve -f config/ > release.yaml

In a manifest, image: ko://example.com/org/cmd/app is replaced by the pushed digest at ko apply/ko resolve time. KO_DOCKER_REPO=ko.local side-loads into the Docker daemon; kind.local loads into kind nodes. Static assets live in <importpath>/kodata/ and are found at runtime via $KO_DATA_PATH; the binary is always installed as /ko-app/<name>.

# .ko.yaml
defaultBaseImage: cgr.dev/chainguard/static
baseImageOverrides:
  example.com/org/cmd/needs-libc: cgr.dev/chainguard/glibc-dynamic
defaultPlatforms: [linux/amd64, linux/arm64]
builds:
  - id: app
    dir: ./cmd/app
    ldflags: ["-s -w", "-X main.version={{.Env.VERSION}}"]
    env: [CGO_ENABLED=0]

The default base image is cgr.dev/chainguard/static (not distroless, as older material says) — verified from a live build log:

Using base cgr.dev/chainguard/static:latest@sha256:f51c2493951313c3ad... for .../cmd/hello
Building example.invalid/hello/cmd/hello for linux/amd64
Building example.invalid/hello/cmd/hello for linux/arm64

Both architectures were built on an x86_64 host with no emulation and no QEMU — Go cross-compiles, so multi-arch costs a second go build and nothing else. That is ko's single biggest practical advantage over every Dockerfile-based builder.

Reproducibility — verified. Two independent builds with the same SOURCE_DATE_EPOCH produced the identical index digest:

SOURCE_DATE_EPOCH=1700000000 ko build ./cmd/hello --push=false --oci-layout-path=./oci2 --bare
# ./oci2@sha256:e56d0a6ed8fc9bce8269f1f84cd45c4da18f6d214b0957fbd5ac9ccfc33a8365
SOURCE_DATE_EPOCH=1700000000 ko build ./cmd/hello --push=false --oci-layout-path=./oci3 --bare
# ./oci3@sha256:e56d0a6ed8fc9bce8269f1f84cd45c4da18f6d214b0957fbd5ac9ccfc33a8365

By default ko embeds no timestamps at all (images show 1970), which is why they reproduce. SOURCE_DATE_EPOCH sets the image creation time; KO_DATA_DATE_EPOCH sets modtimes for kodata assets.

Hard limits. Go only. CGO_ENABLED=0 by default, so cgo needs a custom base and explicit configuration. No OS packages — ca-certificates, ICU data, or a shell is a base-image change, not a build step. ko publish and ko deps no longer exist, and --sbom accepts only spdx and none (anything else silently becomes SPDX). ko is a CNCF Sandbox project.


Rootless and CI

What each tool actually needs on a shared runner:

Tool Requirement The trap
buildx, docker driver Access to the Docker socket Socket access is root on the host. A "non-privileged" job with /var/run/docker.sock mounted can start a privileged container
buildx, docker-container Same socket Same trap. The builder container is created by the daemon
buildx, rootless buildkitd Unprivileged user namespaces; unconfined seccomp and AppArmor Some hardened clusters forbid unconfined profiles outright
Buildah subuid/subgid ranges; newuidmap/newgidmap setuid helpers Missing ranges give a confusing lchown ... invalid argument from the storage driver
Kaniko An isolated pod it can safely trash Running it anywhere it can reach a real filesystem
Stacker subuid/subgid; kernel ≥ 5.11 for unprivileged overlay stacker unpriv-setup writes the ranges, and that step itself needs privilege
apko Nothing —
melange Unprivileged user namespaces for bwrap Nesting inside a rootless container fails: bwrap: Creating new namespace failed: Operation not permitted
ko A Go toolchain —

The privilege-escalation trap is worth stating plainly: mounting the Docker socket into a CI job grants root on the runner. A job that can call docker run --privileged -v /:/host owns the machine and every other tenant's secrets on it. If your runners are shared, the honest options are rootless BuildKit, Buildah in a rootless container, or a remote builder the job cannot reach except through the build API.

# Rootless Buildah inside a container: the container itself needs userns support
podman run --rm --device /dev/fuse \
  --security-opt seccomp=unconfined --security-opt apparmor=unconfined \
  -v "$PWD":/src -w /src quay.io/buildah/stable \
  buildah build -t app:1 .

--device /dev/fuse is needed when the nested storage driver falls back to fuse-overlayfs. Without it you land on vfs and builds become minutes-per-layer.


Multi-architecture

The performance reality first. Measured on Docker Engine 29.7.2 / BuildKit v0.32.2 on Apple Silicon — same shell-loop workload, --no-cache, base image pre-warmed:

linux/arm64    (native)     0.4s
linux/amd64    (Rosetta)    0.5s    ~1.2x
linux/arm/v7   (QEMU)       4.0s     ~10x
linux/s390x    (QEMU)       4.9s     ~12x

The lesson is not "emulation is slow"; it is that QEMU is slow and Rosetta is not QEMU. On an Apple Silicon workstation amd64 emulation is nearly free, because Docker Desktop uses Rosetta 2. On a Linux CI runner with qemu-user-static you pay the QEMU multiplier, and a compile-heavy build that takes four minutes natively takes tens of minutes. That is the "why is my ARM build taking an hour" question, answered — and why you cannot benchmark emulation on a Mac and extrapolate.

The multiplier itself is not portable either. The table above is an Apple Silicon measurement. The same workload shape on a Linux/amd64 box (1 vCPU, Debian 13, tonistiigi/binfmt) measured native 1.6s, arm64 8.1s, arm/v7 8.0s, s390x 9.3s — roughly 5–6x, not 10–12x. Take the order of magnitude, not the number: QEMU costs you multiples, and which multiple depends on the host, the guest architecture and how compute-bound the step is.

Three approaches, in decreasing order of how much you should like them:

1. Native nodes. One builder, several machines, each advertising its own platforms. Fastest and least surprising; costs you runners.

2. Cross-compilation in the Dockerfile. Build natively on $BUILDPLATFORM and copy the artefact into a $TARGETPLATFORM stage, so only the final stage is emulated and it usually runs nothing. This is what ko does implicitly, and why ko is fast.

FROM --platform=$BUILDPLATFORM golang:1.26 AS build
ARG TARGETOS TARGETARCH
WORKDIR /src
COPY . .
RUN CGO_ENABLED=0 GOOS=$TARGETOS GOARCH=$TARGETARCH go build -o /out/app ./cmd/app

FROM gcr.io/distroless/static-debian12
COPY --from=build /out/app /app
ENTRYPOINT ["/app"]

3. Emulation. Register binfmt handlers and build. Correct, universal, and slow.

Assembly after the fact. If per-arch builds run as separate jobs, stitch them at the end; --metadata-file gives you the digests to feed in. Buildah's equivalent is buildah manifest create / add / push --all. Kaniko has no equivalent and needs an external tool.

docker buildx imagetools create -t registry.example.com/app:1 \
  registry.example.com/app@sha256:aaa... registry.example.com/app@sha256:bbb...

Reproducibility

Two things get conflated here. Stable means the digest does not change when nothing changed. Reproducible means someone else, on another machine, gets your exact digest from your exact inputs. Most pipelines want the first and claim the second.

Plain buildx builds are neither, because every layer carries a fresh timestamp. Fix it with SOURCE_DATE_EPOCH and rewrite-timestamp:

docker buildx build --no-cache -t app:a --push .   # sha256:df331dfe244340...
docker buildx build --no-cache -t app:b --push .   # sha256:c07489586118fb...  <- differs

export SOURCE_DATE_EPOCH=1700000000
docker buildx build --no-cache --provenance=false --sbom=false \
  -o type=registry,name=registry.example.com/app:c,rewrite-timestamp=true .
# sha256:<digest A>
docker buildx build --no-cache --provenance=false --sbom=false \
  -o type=registry,name=registry.example.com/app:d,rewrite-timestamp=true .
# sha256:<digest A>   <- identical, and stable across hosts for the same inputs

The split of responsibilities matters:

  • SOURCE_DATE_EPOCH (BuildKit 0.11+, auto-forwarded from the client environment by buildx 0.10+) rewrites the image config created, each history entry's created, and the org.opencontainers.image.created annotation. It does not touch file mtimes inside the layers.
  • rewrite-timestamp=true (BuildKit 0.13+, on the image, registry, oci, and docker exporters) rewrites those file timestamps to the same epoch.
  • Provenance attestations embed build timing and environment, so the index digest moves even when the image manifest does not. Turn them off, or accept it.
  • SOURCE_DATE_EPOCH=context derives the value from the build context — a git commit time, or an HTTP Last-Modified.
  • BuildKit also exposes a compatibility-version exporter option pinning digest-affecting assembly behaviour (10 = v0.13/v0.14, 20 = v0.15–v0.31, 30 = current), for matching a digest computed by an older BuildKit.

What genuinely reproduces versus what merely looks stable:

Reproduces across machines?
apko Yes, bitwise — same config plus same package set, verified above
ko Yes with SOURCE_DATE_EPOCH, verified above; Go builds are deterministic and -trimpath is always on
Stacker Content-hashed layers plus hash:-pinned imports; no published bit-identical guarantee
buildx + SOURCE_DATE_EPOCH + rewrite-timestamp Stable, and reproducible if the Dockerfile is
Buildah --timestamp Same caveats; --source-date-epoch and --rewrite-timestamp need 1.41+ (absent from 1.39.3)
Kaniko --reproducible Exists; unmaintained since 2025

The uncomfortable truth: the builder is rarely why a build fails to reproduce — unpinned package managers are. apt-get install without pinned versions, pip install without a lockfile, or any curl | sh breaks it regardless of the builder. Pin the base image by digest and dependencies by lockfile, and only then blame the tooling.


Build Secrets

Never use ARG or ENV for a secret. Verified — a build arg lands in the image history, readable by anyone who can pull the image:

docker buildx build --build-arg API_TOKEN=s3cr3t-do-not-ship -t app:1 --push .
docker buildx imagetools inspect app:1 --format '{{ json .Image.History }}'
# { "created_by": "ARG API_TOKEN=s3cr3t-do-not-ship", "empty_layer": true },
# { "created_by": "RUN |1 API_TOKEN=s3cr3t-do-not-ship /bin/sh -c echo ... # buildkit" }

And with --provenance=mode=max it also lands in the pushed attestation:

docker buildx imagetools inspect app:1 --format '{{ json .Provenance }}' | python3 -c \
  'import json,sys; print(json.load(sys.stdin)["SLSA"]["buildDefinition"]["externalParameters"]["request"]["args"])'
# {'build-arg:API_TOKEN': 's3cr3t-do-not-ship', 'no-cache': ''}

The default mode=min does not carry args — which is exactly why people turn mode=max on for supply-chain compliance and hand out their tokens in the process.

Use secret mounts instead. The secret is bind-mounted for the lifetime of one RUN and never enters a layer, the history, or the provenance:

# syntax=docker/dockerfile:1
FROM alpine:3.22
RUN --mount=type=secret,id=api_token,required=true \
    curl -H "Authorization: Bearer $(cat /run/secrets/api_token)" -O https://internal/artifact

RUN --mount=type=ssh git clone git@git.example.com:org/private.git /src
docker buildx build --secret id=api_token,src=./token.txt .              # from a file
API_TOKEN=... docker buildx build --secret id=api_token,env=API_TOKEN .  # from the environment
docker buildx build --ssh default .                                      # agent forwarding
docker buildx build --ssh mykey=$HOME/.ssh/id_ed25519 .

Mount options: id, target (default /run/secrets/<id>), env (mount as an environment variable instead of a file — Dockerfile frontend 1.10+), required, mode, uid, gid. The env= form on --secret has been available since buildx v0.6.

Buildah takes the same flags (buildah build --secret id=npmrc,src=$HOME/.npmrc --ssh default .); Kaniko supports neither. The SecretsUsedInArgOrEnv build check catches the obvious cases — docker buildx build --check will flag them. For where the secrets come from in the first place, see the SOPS and HashiCorp Vault sheets.


Signing and Attestations

Build-time attestations prove how an image was built; signatures prove who published it. You want both, and they are separate artefacts in the registry.

cosign generate-key-pair                        # cosign.key + cosign.pub

# Always sign by DIGEST. A tag is mutable, and a signature over a tag proves nothing.
DIGEST=$(docker buildx imagetools inspect registry.example.com/app:1 \
           --format '{{.Manifest.Digest}}')
cosign sign   --key cosign.key registry.example.com/app@$DIGEST
cosign verify --key cosign.pub registry.example.com/app@$DIGEST

# The signature lands as an extra tag alongside the image
curl -s registry.example.com/v2/app/tags/list
# {"name":"app","tags":["1","sha256-fd26cffe341fe85ff138704f419049637b644368422b40c420f60a2db0016d56"]}

# Attach an externally generated SBOM — apko's, for instance, which is not attached by default
cosign attest --key cosign.key --type spdxjson \
  --predicate sbom-x86_64.spdx.json registry.example.com/app@$DIGEST
cosign verify-attestation --key cosign.pub --type spdxjson registry.example.com/app@$DIGEST

# Keyless via OIDC — the better default in CI: no key material, identity bound to the workflow
cosign sign --yes registry.example.com/app@$DIGEST
cosign verify \
  --certificate-identity-regexp 'https://github.com/org/repo/.github/workflows/.*' \
  --certificate-oidc-issuer https://token.actions.githubusercontent.com \
  registry.example.com/app@$DIGEST

cosign 3.x change. --tlog-upload=false is deprecated and now conflicts with the default signing config: "--tlog-upload=false is not supported with --signing-config". To sign without a transparency log — an air-gapped or internal registry — pass a signing config with the Rekor services removed, built with cosign signing-config create. On cosign 2.x, --tlog-upload=false still works.

For scanning the result see the Trivy and Grype and Syft sheets; for registry-side policy and admission control, Container Registries.


CI Integration

GitHub Actions with buildx

name: build
on:
  push:
    tags: ['v*']

permissions:
  contents: read
  packages: write
  id-token: write            # required for keyless cosign

jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v6
      - uses: docker/setup-qemu-action@v4      # only if you actually emulate
      - uses: docker/setup-buildx-action@v4
      - uses: docker/login-action@v4
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}

      - id: meta
        uses: docker/metadata-action@v6
        with:
          images: ghcr.io/${{ github.repository }}
          tags: |
            type=semver,pattern={{version}}
            type=sha,format=long

      - id: build
        uses: docker/build-push-action@v7
        with:
          context: .
          platforms: linux/amd64,linux/arm64
          push: true
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
          cache-from: type=gha
          cache-to: type=gha,mode=max
          provenance: mode=max
          sbom: true
          secrets: |
            npmrc=${{ secrets.NPMRC }}

      - uses: sigstore/cosign-installer@v4
      - run: cosign sign --yes ghcr.io/${{ github.repository }}@${{ steps.build.outputs.digest }}

Points that bite people:

  • docker/build-push-action exports the Actions cache environment for you; a bare docker buildx build step does not, so type=gha silently no-ops unless you add crazy-max/ghaction-github-runtime or set ACTIONS_* yourself.
  • Sign steps.build.outputs.digest. Do not re-resolve the tag.
  • platforms: with QEMU on ubuntu-latest is where the 10x tax lands. For compile-heavy builds, split into an amd64 job and an ubuntu-24.04-arm job and assemble with imagetools create.
  • For multi-target repos, docker/bake-action@v7 runs your docker-bake.hcl directly and keeps the build definition in the repo rather than the workflow.

See the GitHub Actions sheet for the surrounding workflow patterns.

Kubernetes-native with rootless BuildKit

apiVersion: batch/v1
kind: Job
metadata:
  name: build-app
spec:
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: buildkit
          image: moby/buildkit:v0.33.0-rootless
          command: [buildctl-daemonless.sh]
          args:
            - build
            - --frontend=dockerfile.v0
            - --local=context=/src
            - --local=dockerfile=/src
            - --opt=build-arg:VERSION=1.2.3
            - --output=type=image,name=registry.example.com/app:1.2.3,push=true
            - --export-cache=type=registry,ref=registry.example.com/app:cache,mode=max
            - --import-cache=type=registry,ref=registry.example.com/app:cache
          env:
            - { name: BUILDKITD_FLAGS, value: --oci-worker-no-process-sandbox }
            - { name: DOCKER_CONFIG,   value: /home/user/.docker }
          securityContext:
            seccompProfile: { type: Unconfined }
            appArmorProfile: { type: Unconfined }
            runAsUser: 1000
            runAsGroup: 1000
          volumeMounts:
            - { name: src,     mountPath: /src }
            - { name: regcred, mountPath: /home/user/.docker }
      volumes:
        - { name: src, emptyDir: {} }
        - name: regcred
          secret:
            secretName: regcred
            items: [{ key: .dockerconfigjson, path: config.json }]

buildctl-daemonless.sh starts a buildkitd for the life of one build, which suits a Job. For a persistent build service, run deployment+service.rootless.yaml from moby/buildkit/examples/kubernetes/ and point a remote or kubernetes buildx builder at it.


Quick Reference

# --- buildx -------------------------------------------------------------
docker buildx ls                                  # builders, nodes, platforms
docker buildx create --name ci --driver docker-container --bootstrap --use
docker buildx inspect ci                          # platforms, GC policy, worker labels
docker buildx rm ci                               # CI teardown — do not skip this

docker buildx build -t app:1 --push .             # build and push
docker buildx build -t app:1 --load .             # into the local daemon
docker buildx build -o type=local,dest=./out .    # filesystem only, no image
docker buildx build -o type=oci,dest=./app.tar .  # OCI layout
docker buildx build -o type=cacheonly .           # verify only
docker buildx build --check .                     # lint the Dockerfile

docker buildx build --platform linux/amd64,linux/arm64 -t app:1 --push .
docker buildx build --cache-from type=registry,ref=app:cache \
                    --cache-to   type=registry,ref=app:cache,mode=max .
docker buildx build --secret id=tok,env=TOKEN --ssh default .
docker buildx build --provenance=mode=max --sbom=true -t app:1 --push .

docker buildx imagetools inspect app:1
docker buildx imagetools inspect app:1 --format '{{ json .Provenance }}'
docker buildx imagetools create -t app:1 app:1-amd64 app:1-arm64

docker buildx bake --print                        # resolve without building
docker buildx bake --list=targets
docker buildx du && docker buildx prune --filter until=48h

# --- Buildah ------------------------------------------------------------
buildah build -t app:1 .                          # Dockerfile-compatible
buildah build --layers --format docker -t app:1 .
buildah unshare -- bash ./build.sh                # user namespace for mount/chown
buildah from scratch                              # empty rootfs working container
buildah mount|umount|copy|add|run|config <ctr> ...
buildah commit --format oci --rm <ctr> app:1
buildah manifest create app:1 && buildah manifest push --all app:1 docker://reg/app:1

# --- the rest -----------------------------------------------------------
/kaniko/executor --context dir:///workspace --dockerfile Dockerfile \
                 --destination reg/app:1 --cache=true --cache-repo=reg/app/cache
stacker check && stacker build && stacker publish --url docker://reg --tag 1.2.3
melange keygen && melange build pkg.yaml --arch x86_64 --signing-key melange.rsa
apko build apko.yaml app:1 app.tar                # config, tag, output — three args
KO_DOCKER_REPO=reg/org ko build ./cmd/app --platform=linux/amd64,linux/arm64
Variable Effect
SOURCE_DATE_EPOCH Forwarded by buildx as a build arg; honoured by apko, ko, Buildah 1.41+
BUILDX_NO_DEFAULT_ATTESTATIONS Suppress the default provenance attestation
BUILDKIT_SYNTAX Override the Dockerfile frontend without editing the file
BUILDAH_FORMAT docker or oci (default)
BUILDAH_ISOLATION oci, rootless (default when unprivileged), chroot
BUILDAH_LAYERS true to keep intermediate layers
KO_DOCKER_REPO ko's target repo; ko.local / kind.local side-load instead of pushing
KO_DATA_DATE_EPOCH Modtime for ko's kodata static assets

Common Issues and Solutions

Issue Cause Fix
Multi-arch build takes an hour on CI QEMU emulation on compute-bound steps — measured ~10x on Apple Silicon, ~5x on a Linux/amd64 runner; multiples either way Native builder nodes, or cross-compile on $BUILDPLATFORM and copy the artefact into the target stage
Same build is fast locally, glacial in CI Apple Silicon uses Rosetta for amd64 (~1.2x); Linux runners use QEMU (5–12x depending on host and guest arch) Do not benchmark emulation on a Mac and extrapolate
--cache-to type=registry silently does nothing Default docker driver on the classic overlay2 image store Use docker-container, or enable the containerd image store
--load rejects a multi-platform build Classic image store cannot hold an index Enable the containerd image store, or --load one platform at a time
Images "disappeared" after enabling containerd The two stores are separate; nothing is migrated Re-pull, or rebuild. Check with docker info -f '{{ .DriverStatus }}'
type=gha cache no-ops in a bare docker buildx build step ACTIONS_RUNTIME_TOKEN / ACTIONS_RESULTS_URL are not exported to plain steps Use docker/build-push-action, or crazy-max/ghaction-github-runtime
Deploy tool chokes on unknown/unknown manifests Default provenance attestations sit in the index Update the tool, or --provenance=false / BUILDX_NO_DEFAULT_ATTESTATIONS=1
Secret visible in docker history ARG/ENV values persist in the image config RUN --mount=type=secret plus --secret id=...,src= or ,env=
Secret visible in the pushed attestation --provenance=mode=max records every build arg verbatim Same fix. mode=min does not record args, but is not a substitute for secret mounts
RUN --mount=type=cache is rejected No # syntax= directive, so an old bundled frontend is used Add # syntax=docker/dockerfile:1 as the first line
Buildah: cannot mount using driver overlay in rootless mode buildah mount needs a user namespace you own Wrap the script in buildah unshare
Buildah: lchown ... invalid argument, or storage silently on vfs Missing /etc/subuid and /etc/subgid ranges, or no fuse-overlayfs usermod --add-subuids 100000-165535 --add-subgids 100000-165535 "$USER", install fuse-overlayfs, then podman system migrate
Buildah cross-arch build: Exec format error No binfmt handlers registered on the host Register QEMU handlers, or build on a native host and use buildah manifest
Buildah --source-date-epoch not recognised Added in 1.41.0 Use --timestamp <epoch> on versions < 1.41
buildah build --squash removes more than expected Buildah squashes base-image layers too; podman build --squash does not Use --layers plus explicit staging if you want the weaker behaviour
Kaniko destroyed a build machine's filesystem It unpacks the base image into its own /; run outside a container with --force and that is the host Only ever run the official image in a disposable pod. Never --force outside gVisor
Kaniko cache hit rate collapses mid-build Cache reads stop after the first miss Reorder so volatile steps come last, or accept it — the project is archived, migrate
melange: bwrap: Creating new namespace failed Nested user namespaces blocked inside a rootless container Run melange on the host, or in a container that permits nested userns
apko image has no shell and docker exec fails "Distroless by construction" — nothing is installed that you did not name Add busybox to contents.packages for a debug variant, and keep it out of production
Two identical builds produce different digests Fresh timestamps in every layer SOURCE_DATE_EPOCH and rewrite-timestamp=true, and turn provenance off or accept a moving index digest
Reproducible builder, irreproducible image Unpinned apt-get/pip/npm in the Dockerfile Pin the base by digest and dependencies by lockfile. The builder is not the problem
CI job with the Docker socket "isn't privileged" Socket access is root-equivalent on the host Rootless BuildKit or Buildah, or a remote builder reachable only via the build API
cosign 3.x: --tlog-upload=false is not supported with --signing-config cosign 3 deprecated the flag in favour of signing configs Supply a signing config with no Rekor services, via cosign signing-config create

Related Topics

The following topics complement this cheatsheet and would be valuable additions:

  1. Container Image Optimisation - The Dockerfile half of the story: multi-stage builds, layer ordering, cache mounts, base image selection, and size/security work
  2. Cloud Native Buildpacks - The no-Dockerfile model: the CNB lifecycle, pack, builders, and rebasing an OS layer without rebuilding the application
  3. Podman - Rootless containers and the engine that uses Buildah as its build library; podman build, storage, and user-namespace mechanics
  4. Container Registries - Where the output goes: manifest lists, OCI media types, garbage collection, and registry-side authentication and policy
  5. GitHub Actions - Workflow structure, OIDC, caching, and matrix strategy around the build steps shown here
  6. Grype and Syft - SBOM generation and vulnerability scanning against the images and attestations these builders produce