ADR-1306: Replace nvidia/cuda base images with digest-pinned Ubuntu and version-locked apt installation¶
- Status: Accepted
- Date: 2026-09-24
- Deciders: Lusoris
- Tags:
docker,cuda,build,ci,dependencies,security
Context¶
Following ADR-1231 and ADR-1285, container base images and the coordinated CUDA release pin were centralized in build-config.env. However, the CUDA build and runtime containers (Dockerfile, docker/Dockerfile.production-gpu, docker/Dockerfile.node, docker/dev/ubuntu-26.04-cuda.Dockerfile) still built directly FROM nvidia/cuda:<version>-devel-ubuntu26.04 and FROM nvidia/cuda:<version>-runtime-ubuntu26.04.
This created a severe dependency bottleneck:
- Upstream OCI publication lag: NVIDIA publishes deb packages to its apt repository on day 1 of a release, but its OCI container images on Docker Hub frequently lag by weeks or skip point releases altogether. For example, CUDA 13.4.2 was available immediately in the official
ubuntu2604apt repository (cuda-toolkit-13-4v13.4.2-1,cuda-nvcc-13-4v13.4.92-1) and in the official redistributable JSON manifest (redistrib_13.4.2.json), but NVIDIA published nonvidia/cuda:13.4.2-*container images. - Coupling inversion: ADR-1300 freed CI runners from the
Jimver/cuda-toolkitaction by installing directly from NVIDIA's apt repository. Yet container images remained tied tonvidia/cudabase image publication, meaning CI could test on newer CUDA releases that Dockerfiles could not yet build against. - Redundant image sites: The 6
imagesites acrossbuild-config.envand the Dockerfile ARG mirrors represented 60% of the active sites inscripts/ci/check-cuda-pin-lockstep.py, requiring digest coordination for vendor images whose base OS was already Ubuntu 26.04.
Decision¶
We replace all nvidia/cuda base images with digest-pinned Ubuntu 26.04 (ubuntu:26.04@sha256:...) matching DEV_BASE / DEV_UBUNTU, installing the version-locked CUDA toolkit explicitly from NVIDIA's apt repository via a shared, hardened script.
-
Unified Base Image: In
build-config.env,CUDA_BUILDERandCUDA_RUNTIMEare set to the identical digest-pinned Ubuntu 26.04 base asDEV_BASE:ubuntu:26.04@sha256:da6fc2be547864451aa253836dd926da33623312df4a9a243e35dc877c378a78.scripts/ci/check-base-image-single-source.shrequires both keys to equalDEV_BASEbyte-for-byte, including the digest. The narrowdocker/dev/ubuntu-26.04-cuda.Dockerfilecompatibility image joins the generated base-image inventory; the unrelated Alpine, Arch, and Fedora compatibility images remain independent. -
Hardened Shared Installer (
scripts/ci/install-cuda-toolkit.sh): The existing CI installer script is generalized and hardened to support: - Root container environments where
sudois absent, detectingid -uand omittingsudowhen executing as root. - Non-root CI/host environments where
sudois used for privilege escalation. - Distinct installation modes:
--mode=builder: installs exactcuda-nvcc-${series},cuda-cudart-dev-${series}, andcuda-cudart-${series}versions (compiler + dev headers + runtime).--mode=runtime: installs the exactcuda-cudart-${series}version (minimal runtime shared libraries).--mode=full: installs the exactcuda-toolkit-${series}meta-package plus the exact compiler and cudart components needed by the all-backend dev container.
- Every apt operand uses
package=version, and every installed package is checked afterward withdpkg-query; resolving a different candidate is a hard failure. CUDA_APT_LOCK_RELEASE,CUDA_APT_TOOLKIT_VERSION,CUDA_APT_NVCC_VERSION, andCUDA_APT_CUDART_VERSIONrecord the live NVIDIA package mapping inbuild-config.env. The lock deliberately goes stale when Renovate changesCUDA_VERSION, forcing manual metadata review before the bump can pass.- Automated prerequisite detection (
curl,ca-certificates). -
Idempotent
/usr/local/cudasymlink creation pointing to/usr/local/cuda-${dotted}. -
Container Adoption:
Dockerfile: buildsFROM ${CUDA_BUILDER}, executes/tmp/install-cuda-toolkit.sh --mode=builder /tmp.docker/Dockerfile.production-gpu:builder-cuda13executes--mode=builder;final-cuda13buildsFROM ${CUDA_RUNTIME}and executes--mode=runtime.docker/Dockerfile.node:cuda-runtime-libsbuildsFROM ${CUDA_RUNTIME}and executes--mode=runtime, providing/usr/local/cuda/lib64/libcudart.so*to the node stage.docker/dev/ubuntu-26.04-cuda.Dockerfile: buildsFROM ${CUDA_BUILDER}and executes--mode=builder.-
dev/Containerfile: executes the same shared installer in--mode=full; it no longer carries a second keyring/bootstrap or a floatingcuda-toolkit-${series}apt command. -
Coordinated Pin Reduction:
scripts/ci/check-cuda-pin-lockstep.pyretires theimageshape fromSITE_SHAPES. The coordinated pin is reduced from 10 sites to 7 sites across 2 files: build-config.env:CUDA_VERSION="13.4.2"build-config.env:CUDA_APT_PACKAGE="cuda-toolkit-13-4"build-config.env:CUDA_APT_LOCK_RELEASE="13.4.2"build-config.env: exact toolkit, nvcc, and cudart Debian versions-
docker/Dockerfile.production-gpu:"VMAFX production CUDA 13.4.2 runtime"The residual regexRESIDUAL_REsweeps for any unrecognized CUDA release literal and catches any re-introducednvidia/cuda:*image tags. -
Renovate Release Ownership: The old
nvidia/cudaDocker datasource andCUDA release (coordinated pin)group are removed rather than retained as a hidden OCI dependency. Acustom.nvidia-cuda-redistHTML datasource reads NVIDIA's official redist index, and the CUDA regex manager accepts onlyredistrib_X.Y.Z.jsonlinks throughextractVersionTemplate. The index exposes no release timestamps, so one datasource-scoped package rule setsminimumReleaseAgeBehaviourtotimestamp-optional; CUDA updates remain manual-review and never automerge. Renovate owns onlyCUDA_VERSION; the gate derives the apt series and runtime label, while the exact metadata lock remains human-reviewed authority.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
Wait for upstream nvidia/cuda:13.4.2 OCI images | Zero Dockerfile script modifications | Point releases are frequently skipped or delayed by weeks; stalls security updates | Unacceptable release latency; blocked #1525 |
| Build and host private base images in an external registry | Avoids apt-get in downstream Dockerfiles | Requires extra build pipelines, secrets, registry hosting, and provenance signing | High infrastructure overhead for a standard apt package set |
| Base on digest-pinned Ubuntu + shared in-tree install script | Day-1 availability of NVIDIA releases; hermetic single source; unifies CI and container recipes | Build stages run apt-get during container build | Chosen. BuildKit cache mounts minimize rebuild times; aligns with ADR-1300 |
Consequences¶
- Positive:
- Bumping CUDA versions is completely unblocked from third-party OCI image publication.
- CUDA builders and runtimes share the exact same Ubuntu 26.04 base OS, glibc, and package ecosystem as the rest of the repository.
- CUDA apt installs are exact package locks rather than floating
13-4series selections, and installed versions are independently verified. - Six fragile image tag and digest mirrors are eliminated from
check-cuda-pin-lockstep.py. - Renovate discovers CUDA from the same official redist publication channel that exists before vendor OCI images, so automated discovery no longer reintroduces the original bottleneck.
- Minimal runtime container size:
final-cuda13installs onlycuda-cudart-13-4without unnecessary build tools. - Negative:
- Container build stages without BuildKit cache mounts download apt metadata from NVIDIA on first build.
- Each CUDA release bump requires checking NVIDIA's live redist manifest and Ubuntu package index to refresh component versions; those versions cannot be derived safely from the marketing release.
- Neutral / follow-ups:
- Amends ADR-1231 and ADR-1285.
- Windows CI continues using
install-cuda-toolkit.ps1. - ROCm and oneAPI retain their respective vendor configurations.