Rebase-sensitive invariants¶
Cross-package invariants that any upstream-sync or rebase agent must preserve. Referenced from the canonical AGENTS.md harness. Per-subtree detail lives in the AGENTS.md under each subtree; this page is the index. When a rebase touches a cited translation unit, read that subtree harness before resolving conflicts. A subtree AGENTS.md with an AGENTS.d/ next to it is a generated index: read the pages its table names for the paths you touch, and add an invariant as a page there (agents index and topic pages).
Cross-package invariants that any upstream-sync / rebase agent must preserve. Per-subtree details (the load-bearing reasons + load-bearing mechanics) live in the relevant AGENTS.md under that subtree; this list is the index. When a rebase touches the cited TUs, walk the linked AGENTS.md before resolving conflicts.
-
Documentation entry points: keep
README.mdconcise and link to the topic guides for changing build requirements, backend coverage and model defaults.docs/index.mdanddocs/backends/index.mdshould link to backend guides rather than repeat kernel counts or maturity summaries. Keep the repository-root build instructions indocs/getting-started/index.mdand include Meson'score/source directory when showing a configure command. -
Meson test secret environment sanitization (ADR-1333):
scripts/ci/run_meson_test.pydeletes sensitive GitHub credential keys before Meson starts and records its raw parent environment intestlog.txt. Every supported Make, workflow, preflight, bisection, setup-guidance, and Zed entry point must remain on that wrapper.core/meson.buildretains a default test setup usingenvironment().unset()for (GITHUB_PERSONAL_ACCESS_TOKEN,GITHUB_TOKEN,GH_TOKEN,GH_ENTERPRISE_TOKEN,GITHUB_ENTERPRISE_TOKEN,GITHUB_PAT,GH_PAT,GITHUB_AUTH_TOKEN,GITHUB_API_TOKEN,HOMEBREW_GITHUB_API_TOKEN,ACTIONS_ID_TOKEN_REQUEST_TOKEN,ACTIONS_RUNTIME_TOKEN) at the child and JSON-log layer. The regression contract rejects raw supported-entry-point bypasses, alternate setups, and explicit forbidden-name reintroduction. Preserve the runner, callers, setup, andcore/test/test_meson_secret_env_sanitization.pytogether. Raw external Meson/Ninja test commands are outside this bounded guarantee. -
Zed project settings are project-scoped:
.zed/settings.jsonis parsed as Zed'sProjectSettingsContent, so it must not regainagent,agent_servers, provider/model pins, or permission policy. Preserve the currentdocker exec -i vmaf-dev-mcp vmafx-mcpcontext-server entry, the threeStandards:tasks, and the contract inscripts/ci/tests/test_zed_project_config.py. The scoped mechanics live in.zed/AGENTS.md. -
GPU long-tail terminus reached — every registered feature extractor has at least one GPU twin (lpips remains ORT-delegated per ADR-0022). Cross-backend tolerances live in
scripts/ci/cross_backend_parity_gate.py. Governing ADRs: ADR-0182 (batch 1: psnr / ciede / moment), ADR-0188 (batch 2: ssim / ms_ssim / psnr_hvs), ADR-0192 (batch 3: motion_v2 / float-twins / ssimulacra2 / cambi;float_ansnrremoved in commit 70ed8b3ce3 / PR #38). See core/src/feature/AGENTS.md. - Vulkan backend removed (ADR-0726) — the Vulkan backend, its
libvmaf_vulkan.hsurface, thecore/src/vulkan/tree, the Volk-symbol-hiding machinery, and all*_vulkanGLSL kernels (ssim / ms_ssim / motion_v2 / cambi / psnr chroma) no longer exist in the tree. No rebase invariant survives. Treat any lingering Vulkan reference as stale. - MCP embedded scaffold (T5-2a, ADR-0209): ADR-0209. Public header
libvmaf_mcp.h, audit-first-ENOSYSstubs incore/src/mcp/mcp.c,enable_mcp+ 3 transport sub-flags. T5-2b (cJSON + mongoose + transport bodies) is open. See core/AGENTS.md §Rebase-sensitive invariants. - HIP scaffold (T7-10, ADR-0212 placeholder, PR #200) — audit-first AMD HIP backend scaffold. Public
libvmaf_hip.h, 19 registered feature extractors + 3 unregistered legacy stubs,enable_hipmeson option defaultfalse. - SVE2 SIMD ports (T7-38, ADR-0213 placeholder, PR #201) — SSIMULACRA 2 PTLR + IIR-blur SVE2 ports developed against
qemu-aarch64-static. Same bit-exact contract as the existing NEON ports. - GPU-parity CI gate (T6-8, ADR-0214): ADR-0214. Single source of truth for cross-backend tolerances:
scripts/ci/cross_backend_parity_gate.py. Adding a new GPU twin requires (1)FEATURE_METRICSentry, (2)FEATURE_TOLERANCEentry if it relaxes places=4, (3) row indocs/development/cross-backend-gate.md. Declaring a twin bit-identical adds one filescripts/ci/exact_twins.d/<feature>.<backend>(ADR-1428) and edits no shared line; on a conflict in the generateddocs/development/cross-backend-exact-twins.mdtake master's side and runmake docs-fragments-write. See core/AGENTS.md. psnr_hvs_cudareturns the CPU's scores bit for bit (ADR-1397):psnr_hvs_score.custores the 64 termscalc_psnrhvs()sums per block, in the CPU's arithmetic (double masking table and threshold, integer coefficient difference, fatbin built with--fmad=false), andcore/src/feature/psnr_hvs_score.cadds them into one runningfloatin the CPU's order. A change tocalc_psnrhvs()orextract()inthird_party/xiph/psnr_hvs.cchanges the kernel and that file in the same PR.core/test/test_psnr_hvs_twin_exact_sum_contract.pyandtest_psnr_hvs_scoreguard it without a device,test_cuda_psnr_hvs_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md.psnr_hvs_syclandpsnr_hvs_hipreturn the CPU's scores bit for bit (ADR-1401): both store the 64 termscalc_psnrhvs()sums per block and callcore/src/feature/psnr_hvs_score.c, as the CUDA twin does. The HIP kernel (psnr_hvs_score.hip) takes the masking table and the threshold indoubleand is built with-ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt(hip_cu_extra_flags). The SYCL kernel has no fp64: its masking table is a compile-time constant and its threshold comes fromsqrt_prod_rn()incore/src/feature/sycl/sycl_exact_fp.h(integer product and integer square root); it must stay free of scratch memory. A change tocalc_psnrhvs()orextract()inthird_party/xiph/psnr_hvs.cchanges both kernels in the same PR.test_psnr_hvs_twin_exact_sum_contract.pyguards all three twins without a device;test_sycl_psnr_hvs_parity,test_hip_psnr_hvs_parityandtest_sycl_fp_arith_contracton one. See core/src/feature/sycl/AGENTS.md and core/src/feature/hip/AGENTS.md.- FastDVDnet temporal pre-filter (T6-7, ADR-0215 placeholder, PR #203) — 5-frame window pre-filter feeding ssim/ms_ssim.
- psnr chroma GPU twins (T3-15(b), PR #204) —
psnr_cb/psnr_crdevice kernels alongside the existingpsnr_yfrom ADR-0182. (The original Vulkan implementation was removed with the backend in ADR-0726.) - MobileSal saliency extractor (T6-2a, ADR-0218 placeholder, PR #208) — first half of T6-2 (encoder-side ROI bundle). Saliency-weighted VMAF, sidecar emit for
tools/vmaf-roi. - TransNet V2 shot-boundary extractor (T6-3a, PR #210) — ~1M params; feeds
tools/vmaf-perShotCRF predictor. - SYCL SpEED device-resident pipeline (ADR-1358): every SpEED kernel lives in
core/src/feature/sycl/speed_sycl_pipeline.cppand reproducesspeed.coperation for operation; like every SYCL feature TU the SpEED TUs build with contraction off (sycl_strict_fp_args, ADR-1367), divide and take square roots throughdiv_rn()/sqrt_rn(), and never wait on the queue mid-frame.core/test/test_sycl_kernel_source_contract.pyguards the layout;scripts/dev/speed_gpu_parity.py --backend syclre-checks bit parity. See core/src/feature/sycl/AGENTS.md. ssimulacra2_hipreturns the CPU's score bit for bit (ADR-1445):ssimulacra2_device.hipevaluates the six per-pixel terms with the CPU's fp64 expressions (ss2h_terms(), no fp32 pairs) and forms their sums withcore/src/feature/ordered_sum.hin the four kernels of the CUDA twin (ADR-1433): 1024-pixel chunks in raster order, lanes composed in lane order, a checked walk, term-by-term fallback in pixel order. A change tossim_map()/edge_diff_map()inssimulacra2.cchangesss2h_terms()in the same PR.core/test/test_hip_ssimulacra2_exact_contract.pyandcore/test/test_ordered_sum.cguard it without a device,test_hip_ssimulacra2_parity(==) on one. See core/src/feature/hip/AGENTS.md.- SYCL ssimulacra2 / float_ms_ssim single wait (ADR-1363):
ssimulacra2_sycl.cppruns the whole frame on the device and reads one block of per-scale sums incollect(). The exact-fp helpers live incore/src/feature/sycl/sycl_exact_fp.hand need contraction off, which every SYCL feature TU has (ADR-1367).integer_ms_ssim_sycl.cppenqueues every scale insubmit()into its own partials span and waits once.core/test/test_sycl_kernel_source_contract.pyguards all of it. ssimulacra2_syclreturns the CPU's score bit for bit (ADR-1446): a SYCL kernel has no fp64 type, so the six per-sample terms are the CPU's doubles computed in 64-bit integers (core/src/feature/sycl/sycl_ssimulacra2_math.h), and their sums come fromcore/src/feature/ordered_sum.hthrough its_bitsforms (core/src/feature/sycl/sycl_ordered_sum.h): 512-pixel chunks in raster order, lanes composed in lane order, a checked walk, term-by-term fallback in pixel order. The fp32 pair sums are advice for the walk's plan and never a result.ordered_sum.his shared with the CUDA and HIP twins and keeps everydoubleinside#ifndef VMAF_ORDSUM_NO_FP64. A change tossim_map()/edge_diff_map()inssimulacra2.cchanges the math header andreference_terms()ofcore/test/test_sycl_ssimulacra2_math.cin the same PR.core/test/test_sycl_ssimulacra2_exact_contract.pyguards it without a device;test_sycl_ssimulacra2_math,test_sycl_ordered_sum,test_sycl_ssimulacra2_parity(==) andscripts/dev/speed_gpu_parity.py --backend sycl --feature ssimulacra2on one. See core/src/feature/sycl/AGENTS.md.- CUDA CAMBI and SpEED device-resident (ADR-1379, ADR-1380):
cambi_cuda,speed_chroma_cudaandspeed_temporal_cudaread back one result block and wait once per frame, incollect(); a sync must not bring back the host c-values, host pooling, host SpEED linear algebra or a mid-framecuStreamSynchronize. The host constants come fromcambi.c(vmaf_cambi_*helpers incambi_internal.h) andspeed_internal_gpu_configure(), shared with the SYCL twins;speed/speed_score.cukeeps its__f*_rnintrinsics and--fmad=false(every CUDA fatbin's, ADR-1403).core/test/test_cuda_device_resident_contract.pyguards the design. See core/src/feature/cuda/AGENTS.md. - CUDA
speed_chromagate cell (ADR-1430): the twin'slog2stays correctly rounded; the only difference from a glibc CPU is that library'slog2f(13 of 789 measured values, 1.4e-6 at most, none with a correctly roundedlog2fpreloaded). The cell's bound isLIBM_TWINS["speed_chroma"]inscripts/ci/cross_backend_calibration.py(5e-6, sized for scores below 16), andcore/test/test_cuda_speed_chroma_parity.ckeeps its 960x960 textured fixture: a smaller or ramp fixture has a singular covariance and never reaches the scoring path. The HIP twin is listed at the same bound (ADR-1452: 13 of 990 values on a gfx1036); the fixture and the comparison of both tests arecore/test/speed_chroma_twin_parity.h. - No C or C++ translation unit is built with FP contraction (ADR-1461):
core/src/meson.builddeclaresvmaf_strict_fp_argsas a project argument for C and C++ directly after theVMAF strict FP compiler-argument policyblock, above the first build target. Keep both there on a rebase (Meson refusesadd_project_arguments()after a target), and never give a targetvmaf_fp_model_argsalone or any flag that turns contraction back on.core/test/test_strict_fp_compiler_args.pyreads the compile database of the build it runs in;make test-netflix-golden-arm64runs the golden gate on an aarch64 cross build, where a clang build and a GCC build used to differ. See core/AGENTS.md. - SYCL strict FP line on every feature TU (ADR-1367):
core/src/meson.builddefinessycl_strict_fp_argsonce, between theBEGIN/END VMAF SYCL strict FP policymarkers: icpx gets-fp-model=precise -ffp-contract=off -foffload-fp32-prec-div -foffload-fp32-prec-sqrtin that order (precise implies contraction on, so contraction-off must follow it), AdaptiveCpp-ffp-contract=off. Every feature TU takes it throughsycl_feature_tail_args; no TU gets a private FP list.sycl_link_argsalso carriessycl_fp32_prec_argsto every link the icpx driver runs, because the SPIR-V JIT image is generated there; dropping it leaves-Dsycl_icpx_aot_targets=builds with approximate/and sqrt. The MSVC build's explicit device link (ADR-1364) generates every image and takessycl_strict_fp_argswhole.core/test/test_strict_fp_compiler_args.pyexecutes the policy andtest_sycl_fp_arith_contractchecks the device arithmetic. - HIP CAMBI and SpEED device-resident pipelines (ADR-1378, ADR-1384): no host stage of
cambi.c/speed.cand no mid-frame wait; one staged upload, one readback, the wait incollect(). Per-work-item math lives ininteger_cambi/cambi_hip_device.handspeed/speed_hip_device.h, which the host replay tests compile; the SpEED kernel TU keeps-ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt. The init-time helpers arecambi.c's (cambi_internal.h) andspeed_internal_gpu_configure(), shared with SYCL.core/test/test_hip_device_resident_contract.pyguards the layout. See core/src/feature/hip/AGENTS.md. - CUDA ssimulacra2 single readback (ADR-1391):
ssimulacra2_cuda.cenqueues the whole frame insubmit()on the picture stream and reads one block of per-scale sums incollect(); no host compute or host wait mid-frame. The device TUsssimulacra2_deviceandssimulacra2_blurbuild with--fmad=false(as every CUDA fatbin does, ADR-1403), andssimulacra2_device.cucompiles the sharedfeature/ssimulacra2_math.h,ssimulacra2_score.handssimulacra2_eotf_lut.hinto device code through theVMAF_SS2_FUNC/VMAF_SS2_EOTF_LUT_STORAGEhooks, so an upstream change to those helpers must stay valid CUDA device code. The per-pixel SSIM / edge terms are fp64, and their sums are the sums of the CPU's loops (ADR-1433): chunks of 1024 pixels in raster order, integer increments of the running sum's binade composed in pixel order (feature/ordered_sum.h, compiled into device code through theVMAF_ORDSUM_*hooks and tested on the host bycore/test/test_ordered_sum.c), one walk per sum with a term-by-term fallback. The terms must stay non-negative or NaN, and a change tossim_map()/edge_diff_map()inssimulacra2.cchangesss2c_terms()in the same PR.core/test/test_cuda_ssimulacra2_parity.c(==),core/test/test_cuda_ssimulacra2_exact_contract.pyandscripts/dev/speed_gpu_parity.py --backend cuda --feature ssimulacra2re-check parity. See core/src/feature/cuda/AGENTS.md. - SYCL kernels use no scratch memory (ADR-1395): on an Arc A-series GPU under the Linux xe driver, kernels with a private array in memory or spilled registers return wrong values.
test_sycl_kernel_scratchfails on a scratch kernel missing fromcore/src/sycl/scratch_ratchet.txt, whose extractors must matchkScratchExtractorsincore/src/sycl/scratch_check.cpp; the list only shrinks.integer_vif_sycl's SIMD-32 kernels keepVmafSyclKernelShape<32, 256>. See core/src/sycl/AGENTS.md and core/src/feature/sycl/AGENTS.md. float_ms_ssim_cudaper-scale sums are the CPU's, in the CPU's order (ADR-1465):ms_ssim_vert_lcsincore/src/feature/cuda/integer_ms_ssim/ms_ssim_score.custores every window'sl,candsat its raster position andinteger_ms_ssim_cuda.c::ms_ssim_scale_sums()adds the three planes of a scale in index order, asiqa/ssim_tools.c::iqa_ssim()adds them. A sync must not bring back a device reduction of the terms: on the frame ofcore/test/float_ms_ssim_order_frame.hper-block sums return the neighbouringfloatforfloat_ms_ssim_c_scale1. That header is shared with the HIP and SYCL twin tests and its bytes are fixed. Preserve the kernel, the host loop,core/test/test_cuda_float_ms_ssim_order.candcore/test/test_cuda_float_ms_ssim_exact_contract.pytogether.- CUDA device FP policy (ADR-1403): every CUDA fatbin takes
cuda_device_strict_fp_args(--fmad=falseunder nvcc,-ffp-contract=offunder clang CUDA), defined once between theVMAF CUDA device strict FP policymarkers incore/src/meson.build;cuda_cu_extra_flagscarries no floating-point flag. A kernel whose reference fuses writes__fmaf_rn().float_ms_ssim_cudareproducesms_ssim_decimate.c,iqa_convolve()andssim_accumulate_default_scalar()operation for operation and is bit-identical to the CPU.core/test/test_strict_fp_compiler_args.py,core/test/test_cuda_kernel_source_contract.pyandcore/test/test_cuda_float_ms_ssim_parity.cguard it. See core/src/cuda/AGENTS.md and core/src/feature/cuda/AGENTS.md. float_motion_cudaadds its SAD in the CPU's order (ADR-1409):float_motion.c::compute_motion_simd()keeps one fp32 running sum per row and one over the rows. The twin'sfloat_motion_row_sadkernel runs one thread per row with a plain left-to-right loop, and the host finishes throughcore/src/feature/float_motion_sad.h; the scores are the CPU's bit for bit and the parity gate compares them with tolerance 0 (EXACT_TWINS). A change to the CPU's SAD order, or toconvolution_f32_c_s()'s tap order, changes the kernel and the helper in the same PR.core/test/test_cuda_float_motion_parity.c,core/test/test_float_motion_sad.candcore/test/test_cuda_kernel_source_contract.pyguard it.float_motion_sycladds its SAD in the CPU's order (ADR-1411): the same contract as the CUDA twin above.fm_row_sad()incore/src/feature/sycl/float_motion_sycl.cppis one plain left-to-right loop per work-item, launched oversycl::range<1>(height)at sub-group size 8, andcollect()finishes throughcore/src/feature/float_motion_sad.h. No group, sub-group or atomic reduction may return to the TU, and the blur needs the SYCL strict FP line (ADR-1367).core/test/test_sycl_float_motion_parity.c(==) andcore/test/test_sycl_kernel_source_contract.pyguard it; the row kernel must stay free of scratch memory (test_sycl_kernel_scratch, ADR-1395).float_ms_ssim_syclis the CPU's arithmetic (ADR-1414): the decimate spells each tapsycl::fma()asms_ssim_decimate.cfuses it; the window sums and thel/c/sterms come fromcore/src/feature/sycl/sycl_ssim_terms.h, shared withfloat_ssim_sycl(the window sums as exact fp32 pairs,landcas the CPU's doubles in 64-bit integers); every window'sl,candsof every scale is stored unreduced and the host adds them iniqa_ssim()'s raster order (ADR-1466; no reduction may return to the twin); the host rounds each per-scale mean to fp32 and combines asms_ssim.cdoes. A change toms_ssim_decimate.c,iqa/convolve.c,iqa/ssim_tools.c,iqa/ssim_accumulate_lane.horms_ssim.cchanges the header or the twin in the same PR.core/test/test_sycl_ms_ssim_parity.c(==on 18 outputs of 3 frames, and two order pairs: seeded noise andcore/test/float_ms_ssim_order_frame.h, a file shared byte for byte with the CUDA and HIP tests) andcore/test/test_sycl_kernel_source_contract.pyguard it. See core/src/feature/sycl/AGENTS.md.float_vif_syclreturns the CPU's scores bit for bit (ADR-1422): the same contract as the CUDA twin, without an fp64 type. The host takes each scale's Gaussian fromvif_get_filter()and hands it to the kernels by value.core/src/feature/sycl/sycl_float_vif_math.hisvif_pixel_statistic_s()andlog2f_approx()operation for operation; itsone_plus_ratio()evaluates the reference's two fp64 expressions as exact fp32 pairs and replays the fp64 operations in integers next to a rounding boundary.vif_row_sums()adds the terms of a row in one work-item andsum_vif_rows()adds the rows on the host, both in fp32. A change tovif_get_filter(), toVIF_OPT_FAST_LOG2/log2f_approx(), tovif_pixel_statistic_s()or tovif_statistic_s()invif_tools.cchanges that header in the same PR.core/test/test_sycl_float_vif_math.c(host and device),core/test/test_sycl_float_vif_exact_contract.pyandcore/test/test_sycl_float_vif_parity.cguard it; every kernel must stay free of scratch memory (test_sycl_kernel_scratch, ADR-1395). See core/src/feature/sycl/AGENTS.md.vif_syclreturns the CPU's scores bit for bit (ADR-1432):core/src/feature/sycl/sycl_integer_vif_math.hreturns the two integersinteger_vif.c::vif_accumulate_pixel()truncates from its fp64 gain (sigma2_sq - g * sigma12andg * g * sigma1_sq), from one integer division and, for a sample within the fp64 chain's rounding error of an integer, from the reference's fp64 operations replayed in 64-bit integers (core/src/feature/sycl/sycl_soft_double.h, shared withfloat_vif_sycl). The host tail rounds each scale's sums tofloatasvif_store_residuals()does. A change to those lines ofinteger_vif.c(the same lines are inx86/vif_avx2.c,x86/vif_avx512.candarm64/vif_neon.c) changes the header in the same PR. The kernels stay free of fp64, ofsycl::mul_hi()on 64-bit operands (wrong values on an Arc A380) and of scratch memory.core/test/test_sycl_integer_vif_math.c,core/test/test_sycl_vif_exact_gain_contract.pyandcore/test/test_sycl_vif_parity.cguard it. See core/src/feature/sycl/AGENTS.md.integer_ssim_syclreturns the CPU's score bit for bit (ADR-1443):core/src/feature/sycl/sycl_integer_ssim_math.hruns the fp64 operations ofinteger_ssim.c::ssim_reduce_row_range()'s per-pixel term, one for one and in the reference's order, on values held in 64-bit integers (core/src/feature/sycl/sycl_soft_signed.h, onsycl_soft_double.h). The kernel stores the bit pattern of every term unreduced and the host adds the plane incalc_ssim()'s raster order. A change to that expression ininteger_ssim.cchanges the header in the same PR. The twin stays free offloatin its term, of a device reduction of the terms, and of scratch memory (SIMD-16 with the 256-entry register file).core/test/test_sycl_integer_ssim_math.c(host and device),core/test/test_sycl_ssim_exact_contract.pyandcore/test/test_sycl_ssim_parity.cguard it. See core/src/feature/sycl/AGENTS.md.float_adm_syclreturns the CPU's scores bit for bit (ADR-1434):core/src/feature/sycl/sycl_float_adm_math.his the same arithmetic as the CUDA twin's device header, without an fp64 type: the three expressionsadm_tools.cevaluates indouble(the enhancement gain, the 1/30 product and the centre tap's 1/15 product) are exact fp32 pairs, and a result next to an fp32 rounding boundary replays the fp64 operations in 64-bit integers (sycl_soft_double.h). The decouple's quotient is fp32n / d, as the reference'sDIVS()is since ADR-1442; it must not become a product with a reciprocal. The header also holds what one work-item of the decouple, term and row-sum kernels does;float_adm_sycl.cpponly launches them. A row is added by one work-item and the rows by the host, both in fp32. The weights, the region, the pooling and the floor are the reference's own. A change toadm_decouple_s(),adm_csf_s(),adm_cm_thresh3x3_s(),adm_csf_den_scale_s()oradm_cm_s()changes this header and the CUDA one in the same PR.core/test/test_sycl_float_adm_math.candcore/test/test_sycl_float_adm_exact_contract.pyguard it,test_sycl_float_adm_parityon a device; the twin is declared exact byscripts/ci/exact_twins.d/float_adm.sycl. See core/src/feature/sycl/AGENTS.md.- SYCL fp64-less device contract (T7-17, ADR-0220): ADR-0220. SYCL feature kernels are unconditionally fp64-free; a single fp64 instruction in any lambda blocks the whole TU on Arc A-series. See core/src/sycl/AGENTS.md.
- Model registry + Sigstore (T6-9, ADR-0211 placeholder, PR #199):
--tiny-model-verifyflag + registry schema + Sigstore bundle paths. Pairs with ADR-0010 (release signing). - Upstream port — feature/motion options from b949cebf (T-NEW-1): PR #197 (
b949cebf, MERGED 2026-04-29) ported Netflix's feature/motion several-options commit; PR #213 (open) portsd3647c73feature/speedextractors (speed_chroma+speed_temporal). -
Metal
float_ms_ssimoption parity (ADR-1334):float_ms_ssim_metalexposesenable_db,clip_db,enable_chroma, andenable_lcsmatching CPU/SYCL/HIP twins. It emitsfloat_ms_ssim,float_ms_ssim_cb, andfloat_ms_ssim_cron the GPU, enforces the >= 176 minimum plane dimension at init, resolves YUV400P to one plane before chroma validation, and uses the exact ceil-subsampled 351x351 YUV420P luma boundary. It wiress->enable_db, s->max_dbintovmaf_ms_ssim_emit_scores/vmaf_ssim_emit_score_named. Device-free contracts incore/test/test_metal_ms_ssim_option_semantics,core/test/test_metal_ms_ssim_options_contract.py, andcore/test/test_nonfinite_collector_wiring.pyprotect this against regression. -
Integer ADM scale-0 masking centre tap (ADR-1402): the fork keeps the 1/15 centre tap of the masking threshold in int32 and clamps
|x| - thr * 2^shiftto [0, INT32_MAX] in int64, where upstream master narrows the tap to int16 and subtracts in 32 bits (the fork's own Netflix/vmaf PR #1602, second revision, is not merged upstream). The scalar definition isadm_cm_thresh()incore/src/feature/integer_adm_kernels.handadm_cm_excess_s0()incore/src/feature/adm_cm_accumulator.h; the AVX2, AVX-512, CUDA, HIP, SYCL and Metal twins return its value bit for bit and change together with it. A sync must not restore the(int16_t)cast, the 16-bit sign extension in the vector thresholds,adm_i16()on the SYCL centre term orabs(x) - (thr << shift).adm_avx2.candadm_avx512.cno longer carry upstream's macros: their scalar parts are the shared kernels.test_integer_adm_cm_threshold,test_integer_adm_simdandtest_gpu_adm_tiny_framesguard it; the Netflix golden gate must be re-run on any change. See core/src/feature/AGENTS.md and the "fix/adm-cm-centre-tap-wrap" entry of rebase-notes. -
Integer ADM enhancement gain limit (ADR-1413): the limited sample is the double product
rst * adm_enhn_gain_limittruncated toward zero, as the scalar kernels incore/src/feature/integer_adm_kernels.hstore it. The AVX2 and AVX-512 decouple kernels use the truncating conversions (upstream master rounds), and the SYCL twin forms the same value in integers withadm_gain_limit_product()fromcore/src/feature/adm_gain_limit.h. A sync must not restore_mm256_cvtpd_epi32/_mm512_cvtpd_epi32/_mm512_cvtpd_epi64on the product or a fixed-point limit in the twin.test_integer_adm_simd,test_adm_gain_limitandtest_gpu_adm_tiny_framesguard it. See core/src/feature/AGENTS.md and the "fix/adm-decouple-fractional-gain-truncation" entry of rebase-notes. -
adm_cudareturns the CPU's scores bit for bit (ADR-1416):core/src/feature/cuda/integer_adm_cuda.cincludescore/src/feature/integer_adm_kernels.hand takes its CSF weights (adm_csf_factors()), its denominator border and shifts (adm_csf_den_ctx_init(),i4_adm_csf_den_ctx_init()) and its per-scale scores (adm_cm_result(),adm_csf_den_result()and theiri4_forms) from it; it defines none of them itself.integer_adm/adm_csf_den.cufolds one whole row per block throughadm_csf_den_round_row_total()(adm_cm_accumulator.h), with the shifts as kernel arguments. A change to those CPU routines reaches the twin through the header; a change to how the CPU folds a denominator row changes that kernel in the same PR.core/test/test_cuda_adm_exact_contract.pyandtest_adm_cm_row_roundingguard it without a device,test_cuda_adm_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md. -
SYCL integer ADM AIM pass (ADR-1362):
integer_adm_sycl.cppcomputes aim / adm3 on the device and finalises every ADM output in the CPU's float arithmetic (bit-exact with the CPU). The decouple quotient is clamped in int64 before narrowing. Details and the mirror list for upstreaminteger_adm.cchanges: core/src/feature/sycl/AGENTS.md. -
SYCL
float_ssimdecimation mirrors the CPU's (ADR-1370):float_ssim_syclreproducesssim.c's box low-pass andiqa/decimate.c::iqa_decimate()bit for bit (int64 fixed-point window sum,KBND_SYMMETRIC,picture_copy()scaling) and sizes its planes with the sharediqa/decimate_dim.h. A change on the CPU side of that pipeline changescore/src/feature/sycl/integer_ssim_sycl.cppin the same PR. See core/src/feature/sycl/AGENTS.md and core/src/feature/iqa/AGENTS.md. -
HIP
float_ssimdecimation mirrors the CPU's (ADR-1405):core/src/feature/hip/float_ssim/ssim_decimate.his the window sum ofiqa/decimate.c::iqa_decimate()withssim.c's box low-pass (int64 fixed-point sum,KBND_SYMMETRIC,picture_copy()scaling), compiled by the kernel and bycore/test/test_hip_float_ssim_decimate.c, which holds it againstiqa_decimate()byte for byte. A change on the CPU side of that pipeline changes the header in the same PR. See core/src/feature/hip/AGENTS.md. -
CUDA RC3 CPU parity (ADR-1372, ADR-1373, ADR-1374): both CUDA motion twins run the diff-first SAD kernel of
integer_motion_v2/motion_v2_score.cuthroughinteger_motion_sad_cuda.c; an upstream sync must not bring back the blur-each-framemotion_score.cu.psnr_cuda,integer_ssim_cuda,float_ssim_cudaandfloat_motion_cudacarry the CPU option tables and call the CPU's helpers (psnr_score.h,vmaf_ssim_max_db(),motion_clip());ssim_score.cu::ssim_terms()mirrors the CPU'sl * c * srounding point for rounding point, andinteger_ssim_scorebuilds with--fmad=false(every CUDA fatbin's, ADR-1403) and the CPU's grouping. The integer ADM DWT row and tap arithmetic lives ininteger_adm/adm_dwt2_rows.h, andvif_cudafalls back to the CPU below 16 pixels.float_motion_cudaemits the CPU'smotion3(motion_blend_clip()). The motion SAD, PSNR and moment kernels add one atomic per block (per accumulator) and PSNR selects its plane with constant indices (ADR-1392). Details: core/src/feature/cuda/AGENTS.md. -
CUDA
float_ssimis the CPU pipeline on the device (ADR-1399):core/src/feature/cuda/integer_ssim/ssim_score.cureproducesssim.c's box low-pass andiqa/decimate.c::iqa_decimate()(exact int64 window sum, one rounding,KBND_SYMMETRIC),iqa/convolve.c's fp32 products added to adoublesum in both Gaussian passes, and the ADR-1373 per-pixel combine; the host sizes the planes with the sharediqa/decimate_dim.h. Its score equals the CPU's on every measured frame, andtest_cuda_float_ssim_parityasserts equality. A change on the CPU side of that pipeline changes the kernel in the same PR. See core/src/feature/cuda/AGENTS.md and core/src/feature/iqa/AGENTS.md. -
integer_ssim_cudareturns the CPU'sssimbit for bit (ADR-1424):integer_ssim.c::calc_ssim()adds every pixel's term into one double in raster order, sointeger_ssim_vert_combine(core/src/feature/cuda/integer_ssim/integer_ssim_score.cu) stores the terms unreduced andssim_cuda.c::issim_frame_sum()adds the plane it reads back in index order. Do not reduce the double terms on the device and do not reorder the host loop; the int64 weights may stay a block reduction. A change tossim_reduce_row_range()or to the ordercalc_ssim()visits pixels changes the kernel'sissim_term()or the host sum in the same PR.core/test/test_cuda_ssim_exact_contract.pyguards it without a device,test_cuda_ssim_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS, featuressim). See -
ciede_cudaruns the CPU's arithmetic (ADR-1426):core/src/feature/cuda/integer_ciede/ciede_device.hisciede.c'sget_lab_color()andciede2000()statement for statement: fp64 where the reference computes in double, float where it stores in float, every float-to-double promotion of a libm argument written out (the kernel is C++). The kernel stores one float per pixel andciede_frame_sum()(core/src/feature/ciede_frame_sum.h, one definition for the CUDA, SYCL and HIP hosts) adds the read-back plane in raster order. Do not introduce float math functions, a device reduction or another form of the formula. A change toget_lab_color(),ciede2000(),get_r_sub_t()or the order ofextract()'s sum inciede.cchanges that header in the same PR. The twin is not bit-identical (glibc's math library against CUDA's); the gate bounds it at1e-9throughLIBM_TWINS.core/test/test_ciede_device_math.candcore/test/test_cuda_ciede_exact_contract.pyguard it without a device,test_cuda_ciede_parityon one. See core/src/feature/cuda/AGENTS.md. -
ciede_syclruns the CPU's arithmetic on fp32 pairs (ADR-1436):core/src/feature/ciede_ff_math.his the same statements as the CUDA twin'sciede_device.hfor a device without an fp64 type: every fp64 value is an fp32 pair, every math-library call a function ofcore/src/feature/ff_math.h, everyfloatof the reference a float rounded from the pair at the reference's statement. Both headers are backend-neutral and shared withciede_hip(ADR-1448);core/src/feature/sycl/sycl_ciede_math.handsycl_ff_math.honly name the SYCL primitives they are built on. The kernel stores one float per pixel andciede_frame_sum()adds the read-back plane in raster order. Do not introduce the device's fp32 math functions, a device reduction, another form of the formula, or a call the compiler does not inline:ciede_pixel()is flattened into the kernel because a call frame is scratch memory (ADR-1395). The header's constants and tables come fromscripts/dev/gen_sycl_ff_math.py; the tables are read from device memory. A change toget_lab_color(),ciede2000(),get_r_sub_t()or the order ofextract()'s sum inciede.cchanges this header and the CUDA one in the same PR. The twin is not bit-identical (the host'spowf); the gate bounds it at1e-9throughLIBM_TWINS.core/test/test_sycl_ciede_exact_contract.pyguards it without a device,test_sycl_ciede_mathandtest_sycl_ciede_parityon one. See core/src/feature/sycl/AGENTS.md. -
ciede_hipruns the same fp32-pair statements (ADR-1448):core/src/feature/hip/integer_ciede/ciede_score.hipincludescore/src/feature/ciede_ff_math.hthroughcore/src/feature/hip/integer_ciede/ciede_hip_math.h, which names the HIP primitives (core/src/feature/ff_pair.hon plain fp32 operators under the strict FP list,fmaf(),sqrtf(),cbrtf(),expf(0.2f * logf(x))). The device has fp64, but its fp64 math functions cost 17 times the frame time; do not bring them back. The kernel stores one float per pixel and the host adds the plane withciede_frame_sum(). A change to a shared header changes the SYCL twin too: both are re-measured (A380 and gfx1036) in the same PR. The gate bounds the cell at1e-9(LIBM_TWINS), not 0.core/test/test_hip_ciede_exact_contract.pyandtest_hip_ciede_mathguard it without a device,test_hip_ciede_parityon one. -
float_vif_cudareturns the CPU's scores bit for bit (ADR-1412): the host takes each scale's Gaussian fromvif_get_filter(), asfloat_vif.cdoes, and hands it to the kernels; no kernel file holds a tap.core/src/feature/float_vif_gpu_common.h(shared withfloat_vif_hipsince ADR-1444; CUDA compiles it throughcore/src/feature/cuda/float_vif/float_vif_device.h, which maps its operators to the__fmul_rn()family) isvif_pixel_statistic_s()andlog2f_approx()operation for operation (vif_sigma_nsqin fp64),float_vif_row_sumsadds the terms of a row in one thread, andfvif_sum_rows()adds the rows on the host, both in fp32 asvif_statistic_s()does. A change tovif_get_filter(), toVIF_OPT_FAST_LOG2/log2f_approx(), tovif_pixel_statistic_s()or tovif_statistic_s()invif_tools.cchanges that header in the same PR.core/test/test_float_vif_device_math.candcore/test/test_cuda_float_vif_exact_contract.pyguard it without a device,test_cuda_float_vif_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md. float_moment_hipadds the CPU's float squares (ADR-1447): the 16-bit kernel ofcore/src/feature/hip/float_moment/moment_score.hipaddsmoment_float_square(), one fp32 product of the sample with itself converted to an integer, wheremoment.c::compute_2nd_moment()forms the square infloat; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to howmoment.cforms or adds its terms changes the kernel in the same PR.core/test/test_hip_float_moment_exact_contract.pyguards it without a device,test_hip_float_moment_parityon one (==while the sum is below 2^53 units, a derived bound past it).- CUDA twins declared exact as a group (ADR-1457):
scripts/ci/exact_twins.d/{motion,motion_debug,motion_v2,psnr,float_ssim,float_ssim_lcs,float_ms_ssim,float_ms_ssim_lcs,cambi}.cudamake the parity gate compare those cells with tolerance 0, andcore/test/test_cuda_exact_twins.choldsmotion_cuda,motion_v2_cuda,psnr_cuda,float_ssim_cuda,float_ms_ssim_cudaandcambi_cudato==on every output. A rebase that changes one of these twins or its CPU extractor keeps them bit-identical; a twin that drifts is fixed, never given a tolerance or taken off the list. -
float_ssim_cudaframe sums are the CPU's, in the CPU's order (ADR-1464): the pass-2 kernels ofcore/src/feature/cuda/integer_ssim/ssim_score.custore every window's terms at its raster position andinteger_ssim_cuda.c::float_ssim_frame_sum()/float_ssim_frame_sums_lcs()add them in index order, asiqa/ssim_tools.c::iqa_ssim()adds them. A sync must not bring back a device reduction of the terms: on the frame ofcore/test/float_ssim_order_frame.ha per-block sum returns the neighbouringfloat. That header is shared with the HIP and SYCL twin tests and its bytes are fixed. Preserve the kernels, the two host loops,core/test/test_cuda_float_ssim_order.candcore/test/test_cuda_float_ssim_exact_contract.pytogether. -
SYCL kernels require sub-group size 16 or 32 (ADR-1468): the default build compiles every kernel ahead of time for the 19 targets of
sycl_icpx_aot_targets, and the Xe2 targets do not compile a kernel that requires 8.core/src/feature/sycl/sycl_compat.hrejects another size at compile time (VmafSyclSubGroupSize); a rebase must not bring a raw[[sycl::reqd_sub_group_size(N)]]orsub_group_size<N>into a kernel, nor a size 8.core/test/test_sycl_sub_group_size_contract.py(device-free) andcore/test/test_sycl_aot_default_targets.py(suitesycl-aot, compiles every SYCL translation unit for the full default list) guard it;core/test/sycl_aot_targets.pyholds the measured sizes per target family and needs an entry for a target added to the list. - SYCL
float_ssimadds the CPU's terms in the CPU's order (ADR-1463):core/src/feature/sycl/sycl_ssim_terms.h::ssim_double_terms()formsiqa/ssim_accumulate_lane.h'slvandcvas the CPU's doubles in 64-bit integers;float_ssim_syclstores every window's term unreduced and the host adds them in raster order, asiqa/ssim_tools.c::iqa_ssim()does. No reduction may return to the float twin and no host sum may change its order: either moves thefloatmean by one step on frames whose terms cancel. A change tossim_accumulate_lane.hor to the means ofssim_tools.cchanges the header in the same PR.core/test/test_sycl_float_ssim_exact_contract.py(device-free) andcore/test/test_sycl_float_ssim_parity.c(==, with the constructed pair ofcore/test/float_ssim_order_frame.h, a file shared byte for byte with the CUDA and HIP tests) guard it. See core/src/feature/sycl/AGENTS.md. - SYCL twins declared exact as a group (ADR-1451):
scripts/ci/exact_twins.d/{adm,motion,motion_debug,motion_v2,psnr,float_ssim,float_ssim_lcs,cambi}.syclmake the parity gate compare those cells with tolerance 0, andcore/test/test_sycl_exact_twins.choldsadm_sycl,motion_sycl,motion_v2_sycl,psnr_sycl,float_ssim_syclandcambi_syclto==on every output. A rebase that changes one of these twins or its CPU extractor keeps them bit-identical; a twin that drifts is fixed, never given a tolerance or taken off the list. -
float_psnr_cudaadds integers (ADR-1455):core/src/feature/cuda/float_psnr/float_psnr_score.cuforms the CPU's term (diff * diffinfloat, asfloat_psnr.cdoes) with__fmul_rn()as an integer in units of 1 / scaler^2 and reducesuint64values per warp and per block;float_psnr_cuda.c::float_psnr_noise()adds the blocks inuint64and divides the exact total by scaler^2 and the pixel count. An fp32 block sum is exact only up to 24 bits. A change to howfloat_psnr.cforms or adds its terms changes the kernel in the same PR.core/test/test_cuda_float_psnr_exact_contract.pyguards it without a device,test_cuda_float_psnr_parity(==) on one. -
vif_cudareads the CPU's log2 table (ADR-1462):core/src/feature/cuda/integer_vif/vif_statistics.cuhholds the table as the module globalvif_cuda_log2_table,log2_lookup()reads it with the CPU's mask, and no vif kernel source evaluates a logarithm.integer_vif_cuda.c::init_fex_cuda()fills it withvif_log2_table_generate()'s values throughvmaf_cuda_vif_upload_log2_table()before any frame is submitted. When upstream changesvif_statistics.cuhorfilter1d.cu, keep the lookup and do not bringlog_generate()back;filter1d.cuitself is untouched by the fork.core/test/test_cuda_vif_log2_contract.pyguards it without a device,test_cuda_vif_log2_tableon one. -
float_psnr_sycladds integers (ADR-1450):core/src/feature/sycl/float_psnr_sycl.cppforms the CPU's term (diff * diffinfloat, asfloat_psnr.cdoes) as an integer in units of 1 / scaler^2 and reducesuint64values per sub-group, per work-group and on the host; an fp32 group sum is exact only up to 24 bits. The host divides the exact total by scaler^2 and the pixel count. A change to howfloat_psnr.cforms or adds its terms changes the kernel in the same PR.core/test/test_sycl_float_psnr_exact_contract.pyguards it without a device,test_sycl_float_psnr_parity(==) on one. -
float_moment_cudaadds the CPU's float squares (ADR-1453): the 16bpc kernel ofcore/src/feature/cuda/integer_moment/moment_score.cuaddsmoment_float_square(), one__fmul_rn()product of the sample with itself converted to an integer, wheremoment.c::compute_2nd_moment()forms the square infloat; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to howmoment.cforms or adds its terms changes the kernel in the same PR.core/test/test_cuda_float_moment_exact_contract.pyguards it without a device,test_cuda_float_moment_parityon one (==while the sum is below 2^53 units, a derived bound past it). -
float_moment_sycladds the CPU's float squares (ADR-1449): the kernel ofcore/src/feature/sycl/integer_moment_sycl.cppaddsmoment_float_square(), one fp32 product of the sample with itself converted to an integer, wheremoment.c::compute_2nd_moment()forms the square infloat; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to howmoment.cforms or adds its terms changes the kernel in the same PR.core/test/test_sycl_float_moment_exact_contract.pyguards it without a device,test_sycl_float_moment_parityon one (==while the sum is below 2^53 units, a derived bound past it). float_vif_hipreturns the CPU's scores bit for bit (ADR-1444): the twin compilescore/src/feature/float_vif_gpu_common.hwith its default operators, which round once only because every HIP kernel is built withhip_strict_fp_args;float_vif_score.hipdefines no operator, holds no tap and reduces nothing per block. The host takes the taps fromvif_get_filter()and passesvif_sigma_nsqas adouble. A change to the shared header is a change to both twins:core/test/test_hip_float_vif_exact_contract.pyandcore/test/test_float_vif_device_math.cguard it without a device,test_hip_float_vif_parityon one. See core/src/feature/hip/AGENTS.md.-
Integer AIM is not clipped, float AIM is (ADR-1417):
core/src/feature/integer_adm.creportsaim_num / den(vmaf_adm_scale_ratios()),core/src/feature/adm.creportsMIN(aim_num / aim_den, 1)(vmaf_adm_finalize_scores()), each as its upstream file does. The shippedvmaf_v1.0.16models read the integeradm3, which the unclipped AIM takes down toadm_min_val. A sync or a cleanup must not unify the two unless upstream does; the Netflix golden gate has no integer AIM above 1 and would not notice.core/test/test_integer_adm_aim_unclipped.cpins both sides. -
float_adm_hipreturns the CPU's scores bit for bit (ADR-1458): it compilescore/src/feature/float_adm_gpu_common.h, the arithmetic of the CUDA twin (next entry), throughcore/src/feature/hip/float_adm/float_adm_hip_math.h, which keeps the shared header's plain operators: under the strict FP list of the HIP kernels they are the reference's operations, the division included. Do not respell them with the__fmul_rn()family, do not reduce per wave or block, and keep the host on the reference's routines (adm_float_reference.h) withadm_frame_size_check()first ininit.core/test/test_hip_float_adm_exact_contract.pyguards it without a device,test_hip_float_adm_math(device arithmetic against the host, value by value) andtest_hip_float_adm_parityon one. See core/src/feature/hip/AGENTS.md. -
float_adm_cudareturns the CPU's scores bit for bit (ADR-1420):core/src/feature/float_adm_gpu_common.h(shared withfloat_adm_hip;core/src/feature/cuda/float_adm/float_adm_device.hgives it the CUDA device spelling) is the decouple, the CSF, the masking threshold and the reduction terms ofadm_tools.coperation for operation (the gain limit and the 1/30 and 1/15 constants in fp64, the angle threshold as(cos^2 * |o|^2) * |t|^2), and its division is the reference's, the IEEE fp32 quotient (__fdiv_rn(); see the next entry).float_adm_row_sumsadds each row in one thread and the host adds the rows, both in fp32. The weights, the reduced region, the pooling and the angle constant come fromadm_tools.citself throughcore/src/feature/adm_float_reference.h; keep those exports, and keep the four reductions ofadm_tools.conadm_pool_bands_s(). A change toadm_decouple_s(),adm_csf_s(),adm_cm_thresh3x3_s(),adm_csf_den_scale_s()oradm_cm_s()changes the device header in the same PR.core/test/test_float_adm_device_math.candcore/test/test_cuda_float_adm_exact_contract.pyguard it without a device,test_cuda_float_adm_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md. -
Float ADM divides (ADR-1442):
core/src/feature/adm_options.hdoes not defineADM_OPT_RECIP_DIVISIONandcore/src/feature/adm_tools.chas oneDIVS(), the plain quotient, with an#errorif the macro is defined. Upstream Netflix defines the macro and multiplies by a reciprocal refined from the processor'sRCPSSestimate, which made the scores depend on the processor. An upstream sync that touches either file keeps the fork's side of both hunks: no macro, norcp_s(), no<emmintrin.h>inadm_tools.c. No twin may bring a reciprocal estimate, a probe of the host or a table of it back, and no CUDA flag may relax the division (--use_fast_math,-prec-div=false).core/test/test_float_adm_divides_contract.pyscans the reference, everyfloat_admfile of every backend andcore/src/meson.build;core/test/test_float_adm_device_math.cchecks the value on inputs where the estimate and the quotient differ. - Coverage Gate ratchet + per-PR delta gate (ADR-0922): ADR-0922. Absolute floors live in
scripts/ci/coverage-check.sh(OVERALL_MIN=70,CRITICAL_MIN=90,PER_FILE_MIN[...]); per-PR drop tolerance lives inscripts/ci/coverage-delta-check.sh(default 0.5pp on overall and per-touched-file). Lowering any floor or loosening the delta tolerance requires a new ADR superseding ADR-0922. The Coverage Gate job in.github/workflows/tests-and-quality-gates.ymlinvokes both scripts; the delta gate needsactions/checkoutwithfetch-depth: 0because it runsgit merge-base. See scripts/ci/AGENTS.md §Coverage Gate ratchet for the full coupling. -
CI action pins — Windows MSVC dev env (ADR-0635):
.github/workflows/libvmaf-build-matrix.ymlusesTheMrMilchmann/setup-msvc-dev@79dac248…(v4.0.0, Node.js 24) for the Windows GPU build legs. If upstream ADR-0121 is re-implemented or the Windows legs are rebased, do not reintroduceilammy/msvc-dev-cmd(Node.js 20, deprecated 2026-06-02). TheTheMrMilchmannaction is a drop-in replacement with identicalvcvarsall.batsemantics. Also: both Windows jobs are pinned towindows-2025; do not revert towindows-latest(redirect towindows-2025-vs2026takes effect 2026-06-15). -
CPU extractors declare the features they write (ADR-1359): the twin lookup pairs a CPU extractor with a device twin through
provided_features.core/src/feature/float_moment.cis an upstream-mirror file whose list the fork changed from upstream's pseudo-name"float_moment"to the four emittedfloat_moment_*names; an upstream sync must keep the fork's list, or--backend <gpu> --feature float_momentfalls back to the CPU again.vmaf_feature_extractor_twin_audit()andtest_every_device_twin_is_reachable(core/test/test_feature_extractor.c) fail when any registered device twin is unreachable. See core/src/feature/AGENTS.md. -
dev-MCP Docker container (ADR-0451):
dev/Containerfileinstalls CUDA through the shared installer's exact--mode=fullcontract (ADR-1306).build-config.envowns the apt series, release lock, and exact toolkit/nvcc/cudart package versions; do not restore a floatingcuda-toolkit-13-4command in the Containerfile. It also pins the unversionedintel-basekitmeta-package (Intel does not publish aintel-basekit-2025.3apt package), and the digest-pinnedrocm/dev-ubuntu-26.04:10.0.0-fullimage in therocm-srcstage (ADR-1225 / ADR-1231). If SDK versions are bumped (routine security maintenance), update their shared pins inbuild-config.envand regenerate the mirrors before merging; a ROCm bump additionally means re-validating therocm-srcprune list against its hipcc smoke check.dev/scripts/smoke-probe-loop.shassumes the golden pair lives at${VMAF_TESTDATA_PATH}/ref_576x324_48f.yuv/dis_576x324_48f.yuv— do not rename these files. The probe JSON schema fields (ts,host_id,backend_results,mcp_results) are an internal format; updatedocs/development/dev-mcp.mdif the schema changes. This directory does not affect the libvmaf C build or any CI gate. -
Top-level
noxfile.pyis a local-dev affordance, not a CI gate (ADR-0914): The repo-rootnoxfile.pyexposes one session per Python package (ai,mcp,vmaf_tune,dev_llm,roi_score,ensemble_kit,python_harness) plusall/lintmeta-sessions. CI does not call nox — each package keeps its ownpython3 -m venv && pip install -e .[dev] && pytestrecipe in.github/workflows/tests-and-quality-gates.yml. When adding a new Python package, update bothnoxfile.pyand the CI YAML; missing one drifts the dev experience away from CI. Seedocs/development/python-test-orchestrator.md. Thepython_harnesssession intentionally delegates totox -c pythonrather than duplicating the Cython + Netflix golden-data setup that lives inpython/tox.ini; do not collapse them. -
Security support and badge evidence —
SECURITY.mddescribes actual VMAFx release support, not inherited Netflix/libvmaf version strings. Keep the passing worksheet tied to a reviewed source revision and the live project record. Configuration, future releases and agent-authored prose cannot establish historical response times, a human developer's knowledge or a completed external badge. The project website is GitHub Pages; keep the short purpose and participation links indocs/index.md, and verify deployed pages before citing new text.