Skip to content

Rebase-sensitive invariants

Cross-package invariants that any upstream-sync or rebase agent must preserve. Referenced from the canonical AGENTS.md harness. Per-subtree detail lives in the AGENTS.md under each subtree; this page is the index. When a rebase touches a cited translation unit, read that subtree harness before resolving conflicts. A subtree AGENTS.md with an AGENTS.d/ next to it is a generated index: read the pages its table names for the paths you touch, and add an invariant as a page there (agents index and topic pages).

Cross-package invariants that any upstream-sync / rebase agent must preserve. Per-subtree details (the load-bearing reasons + load-bearing mechanics) live in the relevant AGENTS.md under that subtree; this list is the index. When a rebase touches the cited TUs, walk the linked AGENTS.md before resolving conflicts.

  • Documentation entry points: keep README.md concise and link to the topic guides for changing build requirements, backend coverage and model defaults. docs/index.md and docs/backends/index.md should link to backend guides rather than repeat kernel counts or maturity summaries. Keep the repository-root build instructions in docs/getting-started/index.md and include Meson's core/ source directory when showing a configure command.

  • Meson test secret environment sanitization (ADR-1333): scripts/ci/run_meson_test.py deletes sensitive GitHub credential keys before Meson starts and records its raw parent environment in testlog.txt. Every supported Make, workflow, preflight, bisection, setup-guidance, and Zed entry point must remain on that wrapper. core/meson.build retains a default test setup using environment().unset() for (GITHUB_PERSONAL_ACCESS_TOKEN, GITHUB_TOKEN, GH_TOKEN, GH_ENTERPRISE_TOKEN, GITHUB_ENTERPRISE_TOKEN, GITHUB_PAT, GH_PAT, GITHUB_AUTH_TOKEN, GITHUB_API_TOKEN, HOMEBREW_GITHUB_API_TOKEN, ACTIONS_ID_TOKEN_REQUEST_TOKEN, ACTIONS_RUNTIME_TOKEN) at the child and JSON-log layer. The regression contract rejects raw supported-entry-point bypasses, alternate setups, and explicit forbidden-name reintroduction. Preserve the runner, callers, setup, and core/test/test_meson_secret_env_sanitization.py together. Raw external Meson/Ninja test commands are outside this bounded guarantee.

  • Zed project settings are project-scoped: .zed/settings.json is parsed as Zed's ProjectSettingsContent, so it must not regain agent, agent_servers, provider/model pins, or permission policy. Preserve the current docker exec -i vmaf-dev-mcp vmafx-mcp context-server entry, the three Standards: tasks, and the contract in scripts/ci/tests/test_zed_project_config.py. The scoped mechanics live in .zed/AGENTS.md.

  • GPU long-tail terminus reached — every registered feature extractor has at least one GPU twin (lpips remains ORT-delegated per ADR-0022). Cross-backend tolerances live in scripts/ci/cross_backend_parity_gate.py. Governing ADRs: ADR-0182 (batch 1: psnr / ciede / moment), ADR-0188 (batch 2: ssim / ms_ssim / psnr_hvs), ADR-0192 (batch 3: motion_v2 / float-twins / ssimulacra2 / cambi; float_ansnr removed in commit 70ed8b3ce3 / PR #38). See core/src/feature/AGENTS.md.

  • Vulkan backend removed (ADR-0726) — the Vulkan backend, its libvmaf_vulkan.h surface, the core/src/vulkan/ tree, the Volk-symbol-hiding machinery, and all *_vulkan GLSL kernels (ssim / ms_ssim / motion_v2 / cambi / psnr chroma) no longer exist in the tree. No rebase invariant survives. Treat any lingering Vulkan reference as stale.
  • MCP embedded scaffold (T5-2a, ADR-0209): ADR-0209. Public header libvmaf_mcp.h, audit-first -ENOSYS stubs in core/src/mcp/mcp.c, enable_mcp + 3 transport sub-flags. T5-2b (cJSON + mongoose + transport bodies) is open. See core/AGENTS.md §Rebase-sensitive invariants.
  • HIP scaffold (T7-10, ADR-0212 placeholder, PR #200) — audit-first AMD HIP backend scaffold. Public libvmaf_hip.h, 19 registered feature extractors + 3 unregistered legacy stubs, enable_hip meson option default false.
  • SVE2 SIMD ports (T7-38, ADR-0213 placeholder, PR #201) — SSIMULACRA 2 PTLR + IIR-blur SVE2 ports developed against qemu-aarch64-static. Same bit-exact contract as the existing NEON ports.
  • GPU-parity CI gate (T6-8, ADR-0214): ADR-0214. Single source of truth for cross-backend tolerances: scripts/ci/cross_backend_parity_gate.py. Adding a new GPU twin requires (1) FEATURE_METRICS entry, (2) FEATURE_TOLERANCE entry if it relaxes places=4, (3) row in docs/development/cross-backend-gate.md. Declaring a twin bit-identical adds one file scripts/ci/exact_twins.d/<feature>.<backend> (ADR-1428) and edits no shared line; on a conflict in the generated docs/development/cross-backend-exact-twins.md take master's side and run make docs-fragments-write. See core/AGENTS.md.
  • psnr_hvs_cuda returns the CPU's scores bit for bit (ADR-1397): psnr_hvs_score.cu stores the 64 terms calc_psnrhvs() sums per block, in the CPU's arithmetic (double masking table and threshold, integer coefficient difference, fatbin built with --fmad=false), and core/src/feature/psnr_hvs_score.c adds them into one running float in the CPU's order. A change to calc_psnrhvs() or extract() in third_party/xiph/psnr_hvs.c changes the kernel and that file in the same PR. core/test/test_psnr_hvs_twin_exact_sum_contract.py and test_psnr_hvs_score guard it without a device, test_cuda_psnr_hvs_parity on one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md.
  • psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit (ADR-1401): both store the 64 terms calc_psnrhvs() sums per block and call core/src/feature/psnr_hvs_score.c, as the CUDA twin does. The HIP kernel (psnr_hvs_score.hip) takes the masking table and the threshold in double and is built with -ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt (hip_cu_extra_flags). The SYCL kernel has no fp64: its masking table is a compile-time constant and its threshold comes from sqrt_prod_rn() in core/src/feature/sycl/sycl_exact_fp.h (integer product and integer square root); it must stay free of scratch memory. A change to calc_psnrhvs() or extract() in third_party/xiph/psnr_hvs.c changes both kernels in the same PR. test_psnr_hvs_twin_exact_sum_contract.py guards all three twins without a device; test_sycl_psnr_hvs_parity, test_hip_psnr_hvs_parity and test_sycl_fp_arith_contract on one. See core/src/feature/sycl/AGENTS.md and core/src/feature/hip/AGENTS.md.
  • FastDVDnet temporal pre-filter (T6-7, ADR-0215 placeholder, PR #203) — 5-frame window pre-filter feeding ssim/ms_ssim.
  • psnr chroma GPU twins (T3-15(b), PR #204) — psnr_cb / psnr_cr device kernels alongside the existing psnr_y from ADR-0182. (The original Vulkan implementation was removed with the backend in ADR-0726.)
  • MobileSal saliency extractor (T6-2a, ADR-0218 placeholder, PR #208) — first half of T6-2 (encoder-side ROI bundle). Saliency-weighted VMAF, sidecar emit for tools/vmaf-roi.
  • TransNet V2 shot-boundary extractor (T6-3a, PR #210) — ~1M params; feeds tools/vmaf-perShot CRF predictor.
  • SYCL SpEED device-resident pipeline (ADR-1358): every SpEED kernel lives in core/src/feature/sycl/speed_sycl_pipeline.cpp and reproduces speed.c operation for operation; like every SYCL feature TU the SpEED TUs build with contraction off (sycl_strict_fp_args, ADR-1367), divide and take square roots through div_rn() / sqrt_rn(), and never wait on the queue mid-frame. core/test/test_sycl_kernel_source_contract.py guards the layout; scripts/dev/speed_gpu_parity.py --backend sycl re-checks bit parity. See core/src/feature/sycl/AGENTS.md.
  • ssimulacra2_hip returns the CPU's score bit for bit (ADR-1445): ssimulacra2_device.hip evaluates the six per-pixel terms with the CPU's fp64 expressions (ss2h_terms(), no fp32 pairs) and forms their sums with core/src/feature/ordered_sum.h in the four kernels of the CUDA twin (ADR-1433): 1024-pixel chunks in raster order, lanes composed in lane order, a checked walk, term-by-term fallback in pixel order. A change to ssim_map() / edge_diff_map() in ssimulacra2.c changes ss2h_terms() in the same PR. core/test/test_hip_ssimulacra2_exact_contract.py and core/test/test_ordered_sum.c guard it without a device, test_hip_ssimulacra2_parity (==) on one. See core/src/feature/hip/AGENTS.md.
  • SYCL ssimulacra2 / float_ms_ssim single wait (ADR-1363): ssimulacra2_sycl.cpp runs the whole frame on the device and reads one block of per-scale sums in collect(). The exact-fp helpers live in core/src/feature/sycl/sycl_exact_fp.h and need contraction off, which every SYCL feature TU has (ADR-1367). integer_ms_ssim_sycl.cpp enqueues every scale in submit() into its own partials span and waits once. core/test/test_sycl_kernel_source_contract.py guards all of it.
  • ssimulacra2_sycl returns the CPU's score bit for bit (ADR-1446): a SYCL kernel has no fp64 type, so the six per-sample terms are the CPU's doubles computed in 64-bit integers (core/src/feature/sycl/sycl_ssimulacra2_math.h), and their sums come from core/src/feature/ordered_sum.h through its _bits forms (core/src/feature/sycl/sycl_ordered_sum.h): 512-pixel chunks in raster order, lanes composed in lane order, a checked walk, term-by-term fallback in pixel order. The fp32 pair sums are advice for the walk's plan and never a result. ordered_sum.h is shared with the CUDA and HIP twins and keeps every double inside #ifndef VMAF_ORDSUM_NO_FP64. A change to ssim_map() / edge_diff_map() in ssimulacra2.c changes the math header and reference_terms() of core/test/test_sycl_ssimulacra2_math.c in the same PR. core/test/test_sycl_ssimulacra2_exact_contract.py guards it without a device; test_sycl_ssimulacra2_math, test_sycl_ordered_sum, test_sycl_ssimulacra2_parity (==) and scripts/dev/speed_gpu_parity.py --backend sycl --feature ssimulacra2 on one. See core/src/feature/sycl/AGENTS.md.
  • CUDA CAMBI and SpEED device-resident (ADR-1379, ADR-1380): cambi_cuda, speed_chroma_cuda and speed_temporal_cuda read back one result block and wait once per frame, in collect(); a sync must not bring back the host c-values, host pooling, host SpEED linear algebra or a mid-frame cuStreamSynchronize. The host constants come from cambi.c (vmaf_cambi_* helpers in cambi_internal.h) and speed_internal_gpu_configure(), shared with the SYCL twins; speed/speed_score.cu keeps its __f*_rn intrinsics and --fmad=false (every CUDA fatbin's, ADR-1403). core/test/test_cuda_device_resident_contract.py guards the design. See core/src/feature/cuda/AGENTS.md.
  • CUDA speed_chroma gate cell (ADR-1430): the twin's log2 stays correctly rounded; the only difference from a glibc CPU is that library's log2f (13 of 789 measured values, 1.4e-6 at most, none with a correctly rounded log2f preloaded). The cell's bound is LIBM_TWINS["speed_chroma"] in scripts/ci/cross_backend_calibration.py (5e-6, sized for scores below 16), and core/test/test_cuda_speed_chroma_parity.c keeps its 960x960 textured fixture: a smaller or ramp fixture has a singular covariance and never reaches the scoring path. The HIP twin is listed at the same bound (ADR-1452: 13 of 990 values on a gfx1036); the fixture and the comparison of both tests are core/test/speed_chroma_twin_parity.h.
  • No C or C++ translation unit is built with FP contraction (ADR-1461): core/src/meson.build declares vmaf_strict_fp_args as a project argument for C and C++ directly after the VMAF strict FP compiler-argument policy block, above the first build target. Keep both there on a rebase (Meson refuses add_project_arguments() after a target), and never give a target vmaf_fp_model_args alone or any flag that turns contraction back on. core/test/test_strict_fp_compiler_args.py reads the compile database of the build it runs in; make test-netflix-golden-arm64 runs the golden gate on an aarch64 cross build, where a clang build and a GCC build used to differ. See core/AGENTS.md.
  • SYCL strict FP line on every feature TU (ADR-1367): core/src/meson.build defines sycl_strict_fp_args once, between the BEGIN/END VMAF SYCL strict FP policy markers: icpx gets -fp-model=precise -ffp-contract=off -foffload-fp32-prec-div -foffload-fp32-prec-sqrt in that order (precise implies contraction on, so contraction-off must follow it), AdaptiveCpp -ffp-contract=off. Every feature TU takes it through sycl_feature_tail_args; no TU gets a private FP list. sycl_link_args also carries sycl_fp32_prec_args to every link the icpx driver runs, because the SPIR-V JIT image is generated there; dropping it leaves -Dsycl_icpx_aot_targets= builds with approximate / and sqrt. The MSVC build's explicit device link (ADR-1364) generates every image and takes sycl_strict_fp_args whole. core/test/test_strict_fp_compiler_args.py executes the policy and test_sycl_fp_arith_contract checks the device arithmetic.
  • HIP CAMBI and SpEED device-resident pipelines (ADR-1378, ADR-1384): no host stage of cambi.c / speed.c and no mid-frame wait; one staged upload, one readback, the wait in collect(). Per-work-item math lives in integer_cambi/cambi_hip_device.h and speed/speed_hip_device.h, which the host replay tests compile; the SpEED kernel TU keeps -ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt. The init-time helpers are cambi.c's (cambi_internal.h) and speed_internal_gpu_configure(), shared with SYCL. core/test/test_hip_device_resident_contract.py guards the layout. See core/src/feature/hip/AGENTS.md.
  • CUDA ssimulacra2 single readback (ADR-1391): ssimulacra2_cuda.c enqueues the whole frame in submit() on the picture stream and reads one block of per-scale sums in collect(); no host compute or host wait mid-frame. The device TUs ssimulacra2_device and ssimulacra2_blur build with --fmad=false (as every CUDA fatbin does, ADR-1403), and ssimulacra2_device.cu compiles the shared feature/ssimulacra2_math.h, ssimulacra2_score.h and ssimulacra2_eotf_lut.h into device code through the VMAF_SS2_FUNC / VMAF_SS2_EOTF_LUT_STORAGE hooks, so an upstream change to those helpers must stay valid CUDA device code. The per-pixel SSIM / edge terms are fp64, and their sums are the sums of the CPU's loops (ADR-1433): chunks of 1024 pixels in raster order, integer increments of the running sum's binade composed in pixel order (feature/ordered_sum.h, compiled into device code through the VMAF_ORDSUM_* hooks and tested on the host by core/test/test_ordered_sum.c), one walk per sum with a term-by-term fallback. The terms must stay non-negative or NaN, and a change to ssim_map() / edge_diff_map() in ssimulacra2.c changes ss2c_terms() in the same PR. core/test/test_cuda_ssimulacra2_parity.c (==), core/test/test_cuda_ssimulacra2_exact_contract.py and scripts/dev/speed_gpu_parity.py --backend cuda --feature ssimulacra2 re-check parity. See core/src/feature/cuda/AGENTS.md.
  • SYCL kernels use no scratch memory (ADR-1395): on an Arc A-series GPU under the Linux xe driver, kernels with a private array in memory or spilled registers return wrong values. test_sycl_kernel_scratch fails on a scratch kernel missing from core/src/sycl/scratch_ratchet.txt, whose extractors must match kScratchExtractors in core/src/sycl/scratch_check.cpp; the list only shrinks. integer_vif_sycl's SIMD-32 kernels keep VmafSyclKernelShape<32, 256>. See core/src/sycl/AGENTS.md and core/src/feature/sycl/AGENTS.md.
  • float_ms_ssim_cuda per-scale sums are the CPU's, in the CPU's order (ADR-1465): ms_ssim_vert_lcs in core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu stores every window's l, c and s at its raster position and integer_ms_ssim_cuda.c::ms_ssim_scale_sums() adds the three planes of a scale in index order, as iqa/ssim_tools.c::iqa_ssim() adds them. A sync must not bring back a device reduction of the terms: on the frame of core/test/float_ms_ssim_order_frame.h per-block sums return the neighbouring float for float_ms_ssim_c_scale1. That header is shared with the HIP and SYCL twin tests and its bytes are fixed. Preserve the kernel, the host loop, core/test/test_cuda_float_ms_ssim_order.c and core/test/test_cuda_float_ms_ssim_exact_contract.py together.
  • CUDA device FP policy (ADR-1403): every CUDA fatbin takes cuda_device_strict_fp_args (--fmad=false under nvcc, -ffp-contract=off under clang CUDA), defined once between the VMAF CUDA device strict FP policy markers in core/src/meson.build; cuda_cu_extra_flags carries no floating-point flag. A kernel whose reference fuses writes __fmaf_rn(). float_ms_ssim_cuda reproduces ms_ssim_decimate.c, iqa_convolve() and ssim_accumulate_default_scalar() operation for operation and is bit-identical to the CPU. core/test/test_strict_fp_compiler_args.py, core/test/test_cuda_kernel_source_contract.py and core/test/test_cuda_float_ms_ssim_parity.c guard it. See core/src/cuda/AGENTS.md and core/src/feature/cuda/AGENTS.md.
  • float_motion_cuda adds its SAD in the CPU's order (ADR-1409): float_motion.c::compute_motion_simd() keeps one fp32 running sum per row and one over the rows. The twin's float_motion_row_sad kernel runs one thread per row with a plain left-to-right loop, and the host finishes through core/src/feature/float_motion_sad.h; the scores are the CPU's bit for bit and the parity gate compares them with tolerance 0 (EXACT_TWINS). A change to the CPU's SAD order, or to convolution_f32_c_s()'s tap order, changes the kernel and the helper in the same PR. core/test/test_cuda_float_motion_parity.c, core/test/test_float_motion_sad.c and core/test/test_cuda_kernel_source_contract.py guard it.
  • float_motion_sycl adds its SAD in the CPU's order (ADR-1411): the same contract as the CUDA twin above. fm_row_sad() in core/src/feature/sycl/float_motion_sycl.cpp is one plain left-to-right loop per work-item, launched over sycl::range<1>(height) at sub-group size 8, and collect() finishes through core/src/feature/float_motion_sad.h. No group, sub-group or atomic reduction may return to the TU, and the blur needs the SYCL strict FP line (ADR-1367). core/test/test_sycl_float_motion_parity.c (==) and core/test/test_sycl_kernel_source_contract.py guard it; the row kernel must stay free of scratch memory (test_sycl_kernel_scratch, ADR-1395).
  • float_ms_ssim_sycl is the CPU's arithmetic (ADR-1414): the decimate spells each tap sycl::fma() as ms_ssim_decimate.c fuses it; the window sums and the l / c / s terms come from core/src/feature/sycl/sycl_ssim_terms.h, shared with float_ssim_sycl (the window sums as exact fp32 pairs, l and c as the CPU's doubles in 64-bit integers); every window's l, c and s of every scale is stored unreduced and the host adds them in iqa_ssim()'s raster order (ADR-1466; no reduction may return to the twin); the host rounds each per-scale mean to fp32 and combines as ms_ssim.c does. A change to ms_ssim_decimate.c, iqa/convolve.c, iqa/ssim_tools.c, iqa/ssim_accumulate_lane.h or ms_ssim.c changes the header or the twin in the same PR. core/test/test_sycl_ms_ssim_parity.c (== on 18 outputs of 3 frames, and two order pairs: seeded noise and core/test/float_ms_ssim_order_frame.h, a file shared byte for byte with the CUDA and HIP tests) and core/test/test_sycl_kernel_source_contract.py guard it. See core/src/feature/sycl/AGENTS.md.
  • float_vif_sycl returns the CPU's scores bit for bit (ADR-1422): the same contract as the CUDA twin, without an fp64 type. The host takes each scale's Gaussian from vif_get_filter() and hands it to the kernels by value. core/src/feature/sycl/sycl_float_vif_math.h is vif_pixel_statistic_s() and log2f_approx() operation for operation; its one_plus_ratio() evaluates the reference's two fp64 expressions as exact fp32 pairs and replays the fp64 operations in integers next to a rounding boundary. vif_row_sums() adds the terms of a row in one work-item and sum_vif_rows() adds the rows on the host, both in fp32. A change to vif_get_filter(), to VIF_OPT_FAST_LOG2 / log2f_approx(), to vif_pixel_statistic_s() or to vif_statistic_s() in vif_tools.c changes that header in the same PR. core/test/test_sycl_float_vif_math.c (host and device), core/test/test_sycl_float_vif_exact_contract.py and core/test/test_sycl_float_vif_parity.c guard it; every kernel must stay free of scratch memory (test_sycl_kernel_scratch, ADR-1395). See core/src/feature/sycl/AGENTS.md.
  • vif_sycl returns the CPU's scores bit for bit (ADR-1432): core/src/feature/sycl/sycl_integer_vif_math.h returns the two integers integer_vif.c::vif_accumulate_pixel() truncates from its fp64 gain (sigma2_sq - g * sigma12 and g * g * sigma1_sq), from one integer division and, for a sample within the fp64 chain's rounding error of an integer, from the reference's fp64 operations replayed in 64-bit integers (core/src/feature/sycl/sycl_soft_double.h, shared with float_vif_sycl). The host tail rounds each scale's sums to float as vif_store_residuals() does. A change to those lines of integer_vif.c (the same lines are in x86/vif_avx2.c, x86/vif_avx512.c and arm64/vif_neon.c) changes the header in the same PR. The kernels stay free of fp64, of sycl::mul_hi() on 64-bit operands (wrong values on an Arc A380) and of scratch memory. core/test/test_sycl_integer_vif_math.c, core/test/test_sycl_vif_exact_gain_contract.py and core/test/test_sycl_vif_parity.c guard it. See core/src/feature/sycl/AGENTS.md.
  • integer_ssim_sycl returns the CPU's score bit for bit (ADR-1443): core/src/feature/sycl/sycl_integer_ssim_math.h runs the fp64 operations of integer_ssim.c::ssim_reduce_row_range()'s per-pixel term, one for one and in the reference's order, on values held in 64-bit integers (core/src/feature/sycl/sycl_soft_signed.h, on sycl_soft_double.h). The kernel stores the bit pattern of every term unreduced and the host adds the plane in calc_ssim()'s raster order. A change to that expression in integer_ssim.c changes the header in the same PR. The twin stays free of float in its term, of a device reduction of the terms, and of scratch memory (SIMD-16 with the 256-entry register file). core/test/test_sycl_integer_ssim_math.c (host and device), core/test/test_sycl_ssim_exact_contract.py and core/test/test_sycl_ssim_parity.c guard it. See core/src/feature/sycl/AGENTS.md.
  • float_adm_sycl returns the CPU's scores bit for bit (ADR-1434): core/src/feature/sycl/sycl_float_adm_math.h is the same arithmetic as the CUDA twin's device header, without an fp64 type: the three expressions adm_tools.c evaluates in double (the enhancement gain, the 1/30 product and the centre tap's 1/15 product) are exact fp32 pairs, and a result next to an fp32 rounding boundary replays the fp64 operations in 64-bit integers (sycl_soft_double.h). The decouple's quotient is fp32 n / d, as the reference's DIVS() is since ADR-1442; it must not become a product with a reciprocal. The header also holds what one work-item of the decouple, term and row-sum kernels does; float_adm_sycl.cpp only launches them. A row is added by one work-item and the rows by the host, both in fp32. The weights, the region, the pooling and the floor are the reference's own. A change to adm_decouple_s(), adm_csf_s(), adm_cm_thresh3x3_s(), adm_csf_den_scale_s() or adm_cm_s() changes this header and the CUDA one in the same PR. core/test/test_sycl_float_adm_math.c and core/test/test_sycl_float_adm_exact_contract.py guard it, test_sycl_float_adm_parity on a device; the twin is declared exact by scripts/ci/exact_twins.d/float_adm.sycl. See core/src/feature/sycl/AGENTS.md.
  • SYCL fp64-less device contract (T7-17, ADR-0220): ADR-0220. SYCL feature kernels are unconditionally fp64-free; a single fp64 instruction in any lambda blocks the whole TU on Arc A-series. See core/src/sycl/AGENTS.md.
  • Model registry + Sigstore (T6-9, ADR-0211 placeholder, PR #199): --tiny-model-verify flag + registry schema + Sigstore bundle paths. Pairs with ADR-0010 (release signing).
  • Upstream port — feature/motion options from b949cebf (T-NEW-1): PR #197 (b949cebf, MERGED 2026-04-29) ported Netflix's feature/motion several-options commit; PR #213 (open) ports d3647c73 feature/speed extractors (speed_chroma + speed_temporal).
  • Metal float_ms_ssim option parity (ADR-1334): float_ms_ssim_metal exposes enable_db, clip_db, enable_chroma, and enable_lcs matching CPU/SYCL/HIP twins. It emits float_ms_ssim, float_ms_ssim_cb, and float_ms_ssim_cr on the GPU, enforces the >= 176 minimum plane dimension at init, resolves YUV400P to one plane before chroma validation, and uses the exact ceil-subsampled 351x351 YUV420P luma boundary. It wires s->enable_db, s->max_db into vmaf_ms_ssim_emit_scores / vmaf_ssim_emit_score_named. Device-free contracts in core/test/test_metal_ms_ssim_option_semantics, core/test/test_metal_ms_ssim_options_contract.py, and core/test/test_nonfinite_collector_wiring.py protect this against regression.

  • Integer ADM scale-0 masking centre tap (ADR-1402): the fork keeps the 1/15 centre tap of the masking threshold in int32 and clamps |x| - thr * 2^shift to [0, INT32_MAX] in int64, where upstream master narrows the tap to int16 and subtracts in 32 bits (the fork's own Netflix/vmaf PR #1602, second revision, is not merged upstream). The scalar definition is adm_cm_thresh() in core/src/feature/integer_adm_kernels.h and adm_cm_excess_s0() in core/src/feature/adm_cm_accumulator.h; the AVX2, AVX-512, CUDA, HIP, SYCL and Metal twins return its value bit for bit and change together with it. A sync must not restore the (int16_t) cast, the 16-bit sign extension in the vector thresholds, adm_i16() on the SYCL centre term or abs(x) - (thr << shift). adm_avx2.c and adm_avx512.c no longer carry upstream's macros: their scalar parts are the shared kernels. test_integer_adm_cm_threshold, test_integer_adm_simd and test_gpu_adm_tiny_frames guard it; the Netflix golden gate must be re-run on any change. See core/src/feature/AGENTS.md and the "fix/adm-cm-centre-tap-wrap" entry of rebase-notes.

  • Integer ADM enhancement gain limit (ADR-1413): the limited sample is the double product rst * adm_enhn_gain_limit truncated toward zero, as the scalar kernels in core/src/feature/integer_adm_kernels.h store it. The AVX2 and AVX-512 decouple kernels use the truncating conversions (upstream master rounds), and the SYCL twin forms the same value in integers with adm_gain_limit_product() from core/src/feature/adm_gain_limit.h. A sync must not restore _mm256_cvtpd_epi32 / _mm512_cvtpd_epi32 / _mm512_cvtpd_epi64 on the product or a fixed-point limit in the twin. test_integer_adm_simd, test_adm_gain_limit and test_gpu_adm_tiny_frames guard it. See core/src/feature/AGENTS.md and the "fix/adm-decouple-fractional-gain-truncation" entry of rebase-notes.

  • adm_cuda returns the CPU's scores bit for bit (ADR-1416): core/src/feature/cuda/integer_adm_cuda.c includes core/src/feature/integer_adm_kernels.h and takes its CSF weights (adm_csf_factors()), its denominator border and shifts (adm_csf_den_ctx_init(), i4_adm_csf_den_ctx_init()) and its per-scale scores (adm_cm_result(), adm_csf_den_result() and their i4_ forms) from it; it defines none of them itself. integer_adm/adm_csf_den.cu folds one whole row per block through adm_csf_den_round_row_total() (adm_cm_accumulator.h), with the shifts as kernel arguments. A change to those CPU routines reaches the twin through the header; a change to how the CPU folds a denominator row changes that kernel in the same PR. core/test/test_cuda_adm_exact_contract.py and test_adm_cm_row_rounding guard it without a device, test_cuda_adm_parity on one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md.

  • SYCL integer ADM AIM pass (ADR-1362): integer_adm_sycl.cpp computes aim / adm3 on the device and finalises every ADM output in the CPU's float arithmetic (bit-exact with the CPU). The decouple quotient is clamped in int64 before narrowing. Details and the mirror list for upstream integer_adm.c changes: core/src/feature/sycl/AGENTS.md.

  • SYCL float_ssim decimation mirrors the CPU's (ADR-1370): float_ssim_sycl reproduces ssim.c's box low-pass and iqa/decimate.c::iqa_decimate() bit for bit (int64 fixed-point window sum, KBND_SYMMETRIC, picture_copy() scaling) and sizes its planes with the shared iqa/decimate_dim.h. A change on the CPU side of that pipeline changes core/src/feature/sycl/integer_ssim_sycl.cpp in the same PR. See core/src/feature/sycl/AGENTS.md and core/src/feature/iqa/AGENTS.md.

  • HIP float_ssim decimation mirrors the CPU's (ADR-1405): core/src/feature/hip/float_ssim/ssim_decimate.h is the window sum of iqa/decimate.c::iqa_decimate() with ssim.c's box low-pass (int64 fixed-point sum, KBND_SYMMETRIC, picture_copy() scaling), compiled by the kernel and by core/test/test_hip_float_ssim_decimate.c, which holds it against iqa_decimate() byte for byte. A change on the CPU side of that pipeline changes the header in the same PR. See core/src/feature/hip/AGENTS.md.

  • CUDA RC3 CPU parity (ADR-1372, ADR-1373, ADR-1374): both CUDA motion twins run the diff-first SAD kernel of integer_motion_v2/motion_v2_score.cu through integer_motion_sad_cuda.c; an upstream sync must not bring back the blur-each-frame motion_score.cu. psnr_cuda, integer_ssim_cuda, float_ssim_cuda and float_motion_cuda carry the CPU option tables and call the CPU's helpers (psnr_score.h, vmaf_ssim_max_db(), motion_clip()); ssim_score.cu::ssim_terms() mirrors the CPU's l * c * s rounding point for rounding point, and integer_ssim_score builds with --fmad=false (every CUDA fatbin's, ADR-1403) and the CPU's grouping. The integer ADM DWT row and tap arithmetic lives in integer_adm/adm_dwt2_rows.h, and vif_cuda falls back to the CPU below 16 pixels. float_motion_cuda emits the CPU's motion3 (motion_blend_clip()). The motion SAD, PSNR and moment kernels add one atomic per block (per accumulator) and PSNR selects its plane with constant indices (ADR-1392). Details: core/src/feature/cuda/AGENTS.md.

  • CUDA float_ssim is the CPU pipeline on the device (ADR-1399): core/src/feature/cuda/integer_ssim/ssim_score.cu reproduces ssim.c's box low-pass and iqa/decimate.c::iqa_decimate() (exact int64 window sum, one rounding, KBND_SYMMETRIC), iqa/convolve.c's fp32 products added to a double sum in both Gaussian passes, and the ADR-1373 per-pixel combine; the host sizes the planes with the shared iqa/decimate_dim.h. Its score equals the CPU's on every measured frame, and test_cuda_float_ssim_parity asserts equality. A change on the CPU side of that pipeline changes the kernel in the same PR. See core/src/feature/cuda/AGENTS.md and core/src/feature/iqa/AGENTS.md.

  • integer_ssim_cuda returns the CPU's ssim bit for bit (ADR-1424): integer_ssim.c::calc_ssim() adds every pixel's term into one double in raster order, so integer_ssim_vert_combine (core/src/feature/cuda/integer_ssim/integer_ssim_score.cu) stores the terms unreduced and ssim_cuda.c::issim_frame_sum() adds the plane it reads back in index order. Do not reduce the double terms on the device and do not reorder the host loop; the int64 weights may stay a block reduction. A change to ssim_reduce_row_range() or to the order calc_ssim() visits pixels changes the kernel's issim_term() or the host sum in the same PR. core/test/test_cuda_ssim_exact_contract.py guards it without a device, test_cuda_ssim_parity on one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS, feature ssim). See

  • ciede_cuda runs the CPU's arithmetic (ADR-1426): core/src/feature/cuda/integer_ciede/ciede_device.h is ciede.c's get_lab_color() and ciede2000() statement for statement: fp64 where the reference computes in double, float where it stores in float, every float-to-double promotion of a libm argument written out (the kernel is C++). The kernel stores one float per pixel and ciede_frame_sum() (core/src/feature/ciede_frame_sum.h, one definition for the CUDA, SYCL and HIP hosts) adds the read-back plane in raster order. Do not introduce float math functions, a device reduction or another form of the formula. A change to get_lab_color(), ciede2000(), get_r_sub_t() or the order of extract()'s sum in ciede.c changes that header in the same PR. The twin is not bit-identical (glibc's math library against CUDA's); the gate bounds it at 1e-9 through LIBM_TWINS. core/test/test_ciede_device_math.c and core/test/test_cuda_ciede_exact_contract.py guard it without a device, test_cuda_ciede_parity on one. See core/src/feature/cuda/AGENTS.md.

  • ciede_sycl runs the CPU's arithmetic on fp32 pairs (ADR-1436): core/src/feature/ciede_ff_math.h is the same statements as the CUDA twin's ciede_device.h for a device without an fp64 type: every fp64 value is an fp32 pair, every math-library call a function of core/src/feature/ff_math.h, every float of the reference a float rounded from the pair at the reference's statement. Both headers are backend-neutral and shared with ciede_hip (ADR-1448); core/src/feature/sycl/sycl_ciede_math.h and sycl_ff_math.h only name the SYCL primitives they are built on. The kernel stores one float per pixel and ciede_frame_sum() adds the read-back plane in raster order. Do not introduce the device's fp32 math functions, a device reduction, another form of the formula, or a call the compiler does not inline: ciede_pixel() is flattened into the kernel because a call frame is scratch memory (ADR-1395). The header's constants and tables come from scripts/dev/gen_sycl_ff_math.py; the tables are read from device memory. A change to get_lab_color(), ciede2000(), get_r_sub_t() or the order of extract()'s sum in ciede.c changes this header and the CUDA one in the same PR. The twin is not bit-identical (the host's powf); the gate bounds it at 1e-9 through LIBM_TWINS. core/test/test_sycl_ciede_exact_contract.py guards it without a device, test_sycl_ciede_math and test_sycl_ciede_parity on one. See core/src/feature/sycl/AGENTS.md.

  • ciede_hip runs the same fp32-pair statements (ADR-1448): core/src/feature/hip/integer_ciede/ciede_score.hip includes core/src/feature/ciede_ff_math.h through core/src/feature/hip/integer_ciede/ciede_hip_math.h, which names the HIP primitives (core/src/feature/ff_pair.h on plain fp32 operators under the strict FP list, fmaf(), sqrtf(), cbrtf(), expf(0.2f * logf(x))). The device has fp64, but its fp64 math functions cost 17 times the frame time; do not bring them back. The kernel stores one float per pixel and the host adds the plane with ciede_frame_sum(). A change to a shared header changes the SYCL twin too: both are re-measured (A380 and gfx1036) in the same PR. The gate bounds the cell at 1e-9 (LIBM_TWINS), not 0. core/test/test_hip_ciede_exact_contract.py and test_hip_ciede_math guard it without a device, test_hip_ciede_parity on one.

  • float_vif_cuda returns the CPU's scores bit for bit (ADR-1412): the host takes each scale's Gaussian from vif_get_filter(), as float_vif.c does, and hands it to the kernels; no kernel file holds a tap. core/src/feature/float_vif_gpu_common.h (shared with float_vif_hip since ADR-1444; CUDA compiles it through core/src/feature/cuda/float_vif/float_vif_device.h, which maps its operators to the __fmul_rn() family) is vif_pixel_statistic_s() and log2f_approx() operation for operation (vif_sigma_nsq in fp64), float_vif_row_sums adds the terms of a row in one thread, and fvif_sum_rows() adds the rows on the host, both in fp32 as vif_statistic_s() does. A change to vif_get_filter(), to VIF_OPT_FAST_LOG2 / log2f_approx(), to vif_pixel_statistic_s() or to vif_statistic_s() in vif_tools.c changes that header in the same PR. core/test/test_float_vif_device_math.c and core/test/test_cuda_float_vif_exact_contract.py guard it without a device, test_cuda_float_vif_parity on one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md.

  • float_moment_hip adds the CPU's float squares (ADR-1447): the 16-bit kernel of core/src/feature/hip/float_moment/moment_score.hip adds moment_float_square(), one fp32 product of the sample with itself converted to an integer, where moment.c::compute_2nd_moment() forms the square in float; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to how moment.c forms or adds its terms changes the kernel in the same PR. core/test/test_hip_float_moment_exact_contract.py guards it without a device, test_hip_float_moment_parity on one (== while the sum is below 2^53 units, a derived bound past it).
  • CUDA twins declared exact as a group (ADR-1457): scripts/ci/exact_twins.d/{motion,motion_debug,motion_v2,psnr,float_ssim,float_ssim_lcs,float_ms_ssim,float_ms_ssim_lcs,cambi}.cuda make the parity gate compare those cells with tolerance 0, and core/test/test_cuda_exact_twins.c holds motion_cuda, motion_v2_cuda, psnr_cuda, float_ssim_cuda, float_ms_ssim_cuda and cambi_cuda to == on every output. A rebase that changes one of these twins or its CPU extractor keeps them bit-identical; a twin that drifts is fixed, never given a tolerance or taken off the list.
  • float_ssim_cuda frame sums are the CPU's, in the CPU's order (ADR-1464): the pass-2 kernels of core/src/feature/cuda/integer_ssim/ssim_score.cu store every window's terms at its raster position and integer_ssim_cuda.c::float_ssim_frame_sum() / float_ssim_frame_sums_lcs() add them in index order, as iqa/ssim_tools.c::iqa_ssim() adds them. A sync must not bring back a device reduction of the terms: on the frame of core/test/float_ssim_order_frame.h a per-block sum returns the neighbouring float. That header is shared with the HIP and SYCL twin tests and its bytes are fixed. Preserve the kernels, the two host loops, core/test/test_cuda_float_ssim_order.c and core/test/test_cuda_float_ssim_exact_contract.py together.

  • SYCL kernels require sub-group size 16 or 32 (ADR-1468): the default build compiles every kernel ahead of time for the 19 targets of sycl_icpx_aot_targets, and the Xe2 targets do not compile a kernel that requires 8. core/src/feature/sycl/sycl_compat.h rejects another size at compile time (VmafSyclSubGroupSize); a rebase must not bring a raw [[sycl::reqd_sub_group_size(N)]] or sub_group_size<N> into a kernel, nor a size 8. core/test/test_sycl_sub_group_size_contract.py (device-free) and core/test/test_sycl_aot_default_targets.py (suite sycl-aot, compiles every SYCL translation unit for the full default list) guard it; core/test/sycl_aot_targets.py holds the measured sizes per target family and needs an entry for a target added to the list.

  • SYCL float_ssim adds the CPU's terms in the CPU's order (ADR-1463): core/src/feature/sycl/sycl_ssim_terms.h::ssim_double_terms() forms iqa/ssim_accumulate_lane.h's lv and cv as the CPU's doubles in 64-bit integers; float_ssim_sycl stores every window's term unreduced and the host adds them in raster order, as iqa/ssim_tools.c::iqa_ssim() does. No reduction may return to the float twin and no host sum may change its order: either moves the float mean by one step on frames whose terms cancel. A change to ssim_accumulate_lane.h or to the means of ssim_tools.c changes the header in the same PR. core/test/test_sycl_float_ssim_exact_contract.py (device-free) and core/test/test_sycl_float_ssim_parity.c (==, with the constructed pair of core/test/float_ssim_order_frame.h, a file shared byte for byte with the CUDA and HIP tests) guard it. See core/src/feature/sycl/AGENTS.md.
  • SYCL twins declared exact as a group (ADR-1451): scripts/ci/exact_twins.d/{adm,motion,motion_debug,motion_v2,psnr,float_ssim,float_ssim_lcs,cambi}.sycl make the parity gate compare those cells with tolerance 0, and core/test/test_sycl_exact_twins.c holds adm_sycl, motion_sycl, motion_v2_sycl, psnr_sycl, float_ssim_sycl and cambi_sycl to == on every output. A rebase that changes one of these twins or its CPU extractor keeps them bit-identical; a twin that drifts is fixed, never given a tolerance or taken off the list.
  • float_psnr_cuda adds integers (ADR-1455): core/src/feature/cuda/float_psnr/float_psnr_score.cu forms the CPU's term (diff * diff in float, as float_psnr.c does) with __fmul_rn() as an integer in units of 1 / scaler^2 and reduces uint64 values per warp and per block; float_psnr_cuda.c::float_psnr_noise() adds the blocks in uint64 and divides the exact total by scaler^2 and the pixel count. An fp32 block sum is exact only up to 24 bits. A change to how float_psnr.c forms or adds its terms changes the kernel in the same PR. core/test/test_cuda_float_psnr_exact_contract.py guards it without a device, test_cuda_float_psnr_parity (==) on one.

  • vif_cuda reads the CPU's log2 table (ADR-1462): core/src/feature/cuda/integer_vif/vif_statistics.cuh holds the table as the module global vif_cuda_log2_table, log2_lookup() reads it with the CPU's mask, and no vif kernel source evaluates a logarithm. integer_vif_cuda.c::init_fex_cuda() fills it with vif_log2_table_generate()'s values through vmaf_cuda_vif_upload_log2_table() before any frame is submitted. When upstream changes vif_statistics.cuh or filter1d.cu, keep the lookup and do not bring log_generate() back; filter1d.cu itself is untouched by the fork. core/test/test_cuda_vif_log2_contract.py guards it without a device, test_cuda_vif_log2_table on one.

  • float_psnr_sycl adds integers (ADR-1450): core/src/feature/sycl/float_psnr_sycl.cpp forms the CPU's term (diff * diff in float, as float_psnr.c does) as an integer in units of 1 / scaler^2 and reduces uint64 values per sub-group, per work-group and on the host; an fp32 group sum is exact only up to 24 bits. The host divides the exact total by scaler^2 and the pixel count. A change to how float_psnr.c forms or adds its terms changes the kernel in the same PR. core/test/test_sycl_float_psnr_exact_contract.py guards it without a device, test_sycl_float_psnr_parity (==) on one.

  • float_moment_cuda adds the CPU's float squares (ADR-1453): the 16bpc kernel of core/src/feature/cuda/integer_moment/moment_score.cu adds moment_float_square(), one __fmul_rn() product of the sample with itself converted to an integer, where moment.c::compute_2nd_moment() forms the square in float; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to how moment.c forms or adds its terms changes the kernel in the same PR. core/test/test_cuda_float_moment_exact_contract.py guards it without a device, test_cuda_float_moment_parity on one (== while the sum is below 2^53 units, a derived bound past it).

  • float_moment_sycl adds the CPU's float squares (ADR-1449): the kernel of core/src/feature/sycl/integer_moment_sycl.cpp adds moment_float_square(), one fp32 product of the sample with itself converted to an integer, where moment.c::compute_2nd_moment() forms the square in float; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to how moment.c forms or adds its terms changes the kernel in the same PR. core/test/test_sycl_float_moment_exact_contract.py guards it without a device, test_sycl_float_moment_parity on one (== while the sum is below 2^53 units, a derived bound past it).

  • float_vif_hip returns the CPU's scores bit for bit (ADR-1444): the twin compiles core/src/feature/float_vif_gpu_common.h with its default operators, which round once only because every HIP kernel is built with hip_strict_fp_args; float_vif_score.hip defines no operator, holds no tap and reduces nothing per block. The host takes the taps from vif_get_filter() and passes vif_sigma_nsq as a double. A change to the shared header is a change to both twins: core/test/test_hip_float_vif_exact_contract.py and core/test/test_float_vif_device_math.c guard it without a device, test_hip_float_vif_parity on one. See core/src/feature/hip/AGENTS.md.
  • Integer AIM is not clipped, float AIM is (ADR-1417): core/src/feature/integer_adm.c reports aim_num / den (vmaf_adm_scale_ratios()), core/src/feature/adm.c reports MIN(aim_num / aim_den, 1) (vmaf_adm_finalize_scores()), each as its upstream file does. The shipped vmaf_v1.0.16 models read the integer adm3, which the unclipped AIM takes down to adm_min_val. A sync or a cleanup must not unify the two unless upstream does; the Netflix golden gate has no integer AIM above 1 and would not notice. core/test/test_integer_adm_aim_unclipped.c pins both sides.

  • float_adm_hip returns the CPU's scores bit for bit (ADR-1458): it compiles core/src/feature/float_adm_gpu_common.h, the arithmetic of the CUDA twin (next entry), through core/src/feature/hip/float_adm/float_adm_hip_math.h, which keeps the shared header's plain operators: under the strict FP list of the HIP kernels they are the reference's operations, the division included. Do not respell them with the __fmul_rn() family, do not reduce per wave or block, and keep the host on the reference's routines (adm_float_reference.h) with adm_frame_size_check() first in init. core/test/test_hip_float_adm_exact_contract.py guards it without a device, test_hip_float_adm_math (device arithmetic against the host, value by value) and test_hip_float_adm_parity on one. See core/src/feature/hip/AGENTS.md.

  • float_adm_cuda returns the CPU's scores bit for bit (ADR-1420): core/src/feature/float_adm_gpu_common.h (shared with float_adm_hip; core/src/feature/cuda/float_adm/float_adm_device.h gives it the CUDA device spelling) is the decouple, the CSF, the masking threshold and the reduction terms of adm_tools.c operation for operation (the gain limit and the 1/30 and 1/15 constants in fp64, the angle threshold as (cos^2 * |o|^2) * |t|^2), and its division is the reference's, the IEEE fp32 quotient (__fdiv_rn(); see the next entry). float_adm_row_sums adds each row in one thread and the host adds the rows, both in fp32. The weights, the reduced region, the pooling and the angle constant come from adm_tools.c itself through core/src/feature/adm_float_reference.h; keep those exports, and keep the four reductions of adm_tools.c on adm_pool_bands_s(). A change to adm_decouple_s(), adm_csf_s(), adm_cm_thresh3x3_s(), adm_csf_den_scale_s() or adm_cm_s() changes the device header in the same PR. core/test/test_float_adm_device_math.c and core/test/test_cuda_float_adm_exact_contract.py guard it without a device, test_cuda_float_adm_parity on one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md.

  • Float ADM divides (ADR-1442): core/src/feature/adm_options.h does not define ADM_OPT_RECIP_DIVISION and core/src/feature/adm_tools.c has one DIVS(), the plain quotient, with an #error if the macro is defined. Upstream Netflix defines the macro and multiplies by a reciprocal refined from the processor's RCPSS estimate, which made the scores depend on the processor. An upstream sync that touches either file keeps the fork's side of both hunks: no macro, no rcp_s(), no <emmintrin.h> in adm_tools.c. No twin may bring a reciprocal estimate, a probe of the host or a table of it back, and no CUDA flag may relax the division (--use_fast_math, -prec-div=false). core/test/test_float_adm_divides_contract.py scans the reference, every float_adm file of every backend and core/src/meson.build; core/test/test_float_adm_device_math.c checks the value on inputs where the estimate and the quotient differ.

  • Coverage Gate ratchet + per-PR delta gate (ADR-0922): ADR-0922. Absolute floors live in scripts/ci/coverage-check.sh (OVERALL_MIN=70, CRITICAL_MIN=90, PER_FILE_MIN[...]); per-PR drop tolerance lives in scripts/ci/coverage-delta-check.sh (default 0.5pp on overall and per-touched-file). Lowering any floor or loosening the delta tolerance requires a new ADR superseding ADR-0922. The Coverage Gate job in .github/workflows/tests-and-quality-gates.yml invokes both scripts; the delta gate needs actions/checkout with fetch-depth: 0 because it runs git merge-base. See scripts/ci/AGENTS.md §Coverage Gate ratchet for the full coupling.
  • CI action pins — Windows MSVC dev env (ADR-0635): .github/workflows/libvmaf-build-matrix.yml uses TheMrMilchmann/setup-msvc-dev@79dac248… (v4.0.0, Node.js 24) for the Windows GPU build legs. If upstream ADR-0121 is re-implemented or the Windows legs are rebased, do not reintroduce ilammy/msvc-dev-cmd (Node.js 20, deprecated 2026-06-02). The TheMrMilchmann action is a drop-in replacement with identical vcvarsall.bat semantics. Also: both Windows jobs are pinned to windows-2025; do not revert to windows-latest (redirect to windows-2025-vs2026 takes effect 2026-06-15).

  • CPU extractors declare the features they write (ADR-1359): the twin lookup pairs a CPU extractor with a device twin through provided_features. core/src/feature/float_moment.c is an upstream-mirror file whose list the fork changed from upstream's pseudo-name "float_moment" to the four emitted float_moment_* names; an upstream sync must keep the fork's list, or --backend <gpu> --feature float_moment falls back to the CPU again. vmaf_feature_extractor_twin_audit() and test_every_device_twin_is_reachable (core/test/test_feature_extractor.c) fail when any registered device twin is unreachable. See core/src/feature/AGENTS.md.

  • dev-MCP Docker container (ADR-0451): dev/Containerfile installs CUDA through the shared installer's exact --mode=full contract (ADR-1306). build-config.env owns the apt series, release lock, and exact toolkit/nvcc/cudart package versions; do not restore a floating cuda-toolkit-13-4 command in the Containerfile. It also pins the unversioned intel-basekit meta-package (Intel does not publish a intel-basekit-2025.3 apt package), and the digest-pinned rocm/dev-ubuntu-26.04:10.0.0-full image in the rocm-src stage (ADR-1225 / ADR-1231). If SDK versions are bumped (routine security maintenance), update their shared pins in build-config.env and regenerate the mirrors before merging; a ROCm bump additionally means re-validating the rocm-src prune list against its hipcc smoke check. dev/scripts/smoke-probe-loop.sh assumes the golden pair lives at ${VMAF_TESTDATA_PATH}/ref_576x324_48f.yuv / dis_576x324_48f.yuv — do not rename these files. The probe JSON schema fields (ts, host_id, backend_results, mcp_results) are an internal format; update docs/development/dev-mcp.md if the schema changes. This directory does not affect the libvmaf C build or any CI gate.

  • Top-level noxfile.py is a local-dev affordance, not a CI gate (ADR-0914): The repo-root noxfile.py exposes one session per Python package (ai, mcp, vmaf_tune, dev_llm, roi_score, ensemble_kit, python_harness) plus all / lint meta-sessions. CI does not call nox — each package keeps its own python3 -m venv && pip install -e .[dev] && pytest recipe in .github/workflows/tests-and-quality-gates.yml. When adding a new Python package, update both noxfile.py and the CI YAML; missing one drifts the dev experience away from CI. See docs/development/python-test-orchestrator.md. The python_harness session intentionally delegates to tox -c python rather than duplicating the Cython + Netflix golden-data setup that lives in python/tox.ini; do not collapse them.

  • Security support and badge evidence — SECURITY.md describes actual VMAFx release support, not inherited Netflix/libvmaf version strings. Keep the passing worksheet tied to a reviewed source revision and the live project record. Configuration, future releases and agent-authored prose cannot establish historical response times, a human developer's knowledge or a completed external badge. The project website is GitHub Pages; keep the short purpose and participation links in docs/index.md, and verify deployed pages before citing new text.