ADR-1373: CUDA PSNR, SSIM and motion twins take the CPU option tables and the CPU's arithmetic¶
- Status: Accepted
- Date: 2026-09-30
- Deciders: lusoris
- Tags: cuda, gpu-parity, numerics, feature-extractor, fork-local
Context¶
T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26 lists CPU options that stale-branch merges removed from the GPU twins. ADR-1365 closed the SYCL part. For CUDA: psnr_cuda lacked min_sse, enable_mse, reduced_hbd_peak and enable_apsnr; integer_ssim_cuda lacked enable_db / clip_db; float_ssim_cuda lacked enable_lcs, enable_db and clip_db and declared an enable_chroma the CPU float_ssim does not have (its kernel read luma only, so the option did nothing); float_motion_cuda lacked motion_max_val and emitted its debug motion score without motion_fps_weight. Scores stayed correct, because the ADR-1183 gate keeps a feature on the CPU when the twin cannot honour a model option, but those features never ran on the device, and naming the twin with the option failed with unknown option.
Two reviews of the first version of this change (#1637) found more in the same twins:
float_ssim_cudaforced identical windows to exactly 1. The CPU does not:iqa/ssim_tools.ccomputesl * c * swith double numerators over fp32-rounded denominators and returns the frame mean rounded to fp32, so identical flat frames (64x64 at 128, 8-bit; 512 at 10-bit) report 72.247199 dB withenable_db, where the forced 1 gave+inf(and 88 dB against 72.247 withclip_db).integer_ssim_cudagrouped each pixel's term asw * (a * b / den); the CPU computes((w * a) * b) / den. The two round differently wherever the window weight is not a power of two, on 841 of 1952 border pixels at 161x91. NVCC's default contraction also fused the constantc1into the denominator only.motion_v2_cudapublished the unweighted, uncapped SAD, where the CPU publishesMIN(sad * motion_fps_weight, motion_max_val); itsmotion2_v2was weighted but never capped, themotion3_v2stamp blended the unweighted SAD, and a one-frame input emitted nomotion2_v2/motion3_v2where the CPU emits 0 / 0.psnr_cudalacked the CPUpsnr'sVMAF_FEATURE_EXTRACTOR_TEMPORALflag, so with--subsample Nitsapsnr_*summed one frame in N. Its chroma accumulators were zeroed on the readback stream while the chroma kernels ran on the picture stream, the racekernel_template.hdocuments for plane 0.
Constraints: every option must reproduce the CPU semantics exactly; no host round trip in the middle of a frame and one wait per frame, at collect (req); one behaviour, one implementation (HISS-19); public options evolve additively (HISS-14).
Decision¶
We will give the CUDA twins their CPU extractor's option tables (same names, aliases, types, defaults, ranges and flags) and reproduce the CPU's arithmetic, types and rounding points:
- Scalar options on the host, from device-reduced sums, through the CPU's helpers.
psnr_cudaturns each plane's device SSE intopsnr_*,mse_*and (in a newflush)apsnr_*withcore/src/feature/psnr_score.h, which the CPUinteger_psnr.ccalls too; its local copy of the MSE-to-PSNR arithmetic is deleted.psnr_cudaisTEMPORALlike the CPUpsnr, and zeroes every plane's accumulator on the picture stream its kernels run on. The SSIM twins applyenable_db/clip_dbthrough thenonfinite_score.hemitters withvmaf_ssim_max_db().float_motion_cudaapplies the CPU'smotion_clip()(fps weight, thenmotion_max_val) to every score it emits, the debugmotionscore included, and emits the CPU'smotion3with itsmotion_blend_factor/motion_blend_offsetoptions (motion_blend_clip()ofmotion2, frame 0 from the first SAD, the tail fromflush), which the twin did not emit at all (T-GPU-FLOAT-MOTION3-MISSING-2026-09-30, found on the device).motion_v2_cudapublishes the CPU's weighted, capped SAD score and derivesmotion2_v2/motion3_v2from it exactly asinteger_motion_v2.c::flushdoes, one-frame input included. float_ssim_cudacomputes the CPU'sl * c * s. Each pixel followsiqa/ssim_tools.c::ssim_variance_scalarandiqa/ssim_accumulate_lane.h(the helper every CPU SIMD variant shares) operand for operand: variances clamped at zero,landcas double numerators over fp32 denominators,sin fp32, the product in double, each rounding spelled with an intrinsic (__fmul_rn,__fadd_rn,__ddiv_rn, ...) so NVCC cannot fuse it. Block partials are doubles and the host rounds each frame mean to fp32, asiqa_ssim()returns it (as SYCL does since ADR-1370). No window is forced to 1: identical frames score whatever the CPU scores.enable_lcslaunches a second pass-2 kernel that reduces the samel,candsnext to the score.integer_ssim_cudacomputes the CPU's term. Its kernel builds with--fmad=false, as the HIP twin builds with-ffp-contract=off(ADR-0564), and groups each term as the CPU does,((w * a) * b) / den. Every per-pixel term is then the CPU's bit for bit. The frame sum is not: the CPU adds the terms row by row, the kernel per warp, per block and then on the host, so the score can differ by a double rounding of the sum. We keep the parallel reduction and drop any bit-exactness claim.enable_chromastays onfloat_ssim_cuda, accepted and ignored. It was a public option of the twin, so removing it would break--feature float_ssim_cuda=enable_chroma=...(HISS-14). It defaults to false, names no feature, logs that it has no effect when set, and is the one twin-only optiontest_cuda_twin_option_parityallows.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Host options from device sums through the CPU helpers + the CPU's per-pixel SSIM arithmetic and fp32 frame mean (chosen) | Every option on the device path; PSNR bit-exact; one PSNR implementation for CPU, SYCL and CUDA; identical frames match the CPU instead of a guessed 1 | The float SSIM combine and its partials are in double (fp64 on the device); default float_ssim_cuda output moves towards the CPU | — |
| Force identical windows to exactly 1 (the first version, and ADR-1365 for SYCL) | Simple; +inf for identical noise frames | Wrong wherever the CPU is not exactly 1: identical flat frames report +inf against the CPU's 72.247 dB | Refuted by a CPU run (#1637 review) |
Include iqa/ssim_accumulate_lane.h in the device code | One source for CPU and CUDA | The CPU header relies on the host compiler not contracting; making it device-safe means rounding macros in a CPU hot-path header shared by three SIMD variants | The CUDA copy spells each rounding as an intrinsic and cites the header; the contract test pins its rounding points |
| Accumulate the float SSIM convolution in double like the CPU | The moments too would be the CPU's | Roughly 110 fp64 adds per pixel, at 1/64 rate on consumer GPUs | Not required for the defects found; left for a measured follow-up |
| Make the integer SSIM frame sum row-major sequential | Bit-exact with the CPU | A serial pass over every pixel | Throughput; the remaining difference is a double rounding of the sum |
| Keep a CUDA-local copy of the PSNR math | No shared-header dependency | Two implementations of the same arithmetic drift apart | HISS-19 |
| Compute L/C/S on the host from read-back moments | No new kernel | Five full-resolution float planes over PCIe per frame, a mid-frame host dependency | Violates the no-round-trip requirement |
Remove enable_chroma (the first version) | The twin's table equals the CPU's | Breaks a public command line without the ! / Migration: HISS-14 asks for | Keep it as an accepted no-op |
Mark the options VMAF_OPT_FLAG_DEFAULT_ONLY (ADR-1316) | Named requests stop failing | The features still never run on the device | Status quo with better errors |
Consequences¶
- Positive: models that set these options keep the features on the CUDA device.
psnr_cudamatches the CPU bit for bit for every option, also under--subsample;motion_v2_cudamatches the CPU formotion_fps_weight,motion_max_valand one-frame inputs;float_ssim_cudareproduces the CPU'senable_dbresult on identical frames, finite or+inf;integer_ssim_cuda's per-pixel terms equal the CPU's (host replay: 0 mismatches over 9,000,000 terms). Measured on an RTX 4090 on 2026-09-30 with the Netflix pair: every PSNR option gives identicalpsnr_*,mse_*andapsnr_*,--subsample 2included;motion_v2andfloat_motionwithmotion_fps_weight/motion_max_valare identical,motion3included, andfloat_motion'smotion3is within 2.8e-6 of the CPU at the defaults and 1.4e-6 with the blend options (the same run on the final head on 2026-10-01 gave the same numbers);ssimwithenable_dbis within 7.3e-13 dB of the CPU andfloat_ssimwithin 6.9e-6 dB; identical flat 64x64 frames give 72.247198959355487 dB on both sides, as the host replay predicted. - Negative: default
float_ssim_cudaandinteger_ssim_cudaoutputs move towards the CPU by an fp32 / double rounding. The float SSIM combine now runs in fp64; on an RTX 4090 at 4K its cost does not show above the frame upload (4.2 ms per frame against 4.6 for amasterbuild, 4.1 withenable_lcs; (t(200) - t(2)) / 198, median of 7, interleaved on a shared host).integer_ssim_cudastill differs from the CPU by a double rounding of the frame sum; on identical frames with a side below 12 pixels that can put one side at+infand the other near 156 dB withenable_dband noclip_db(host replay: 72 of 57,600 such frames of 1 to 24 pixels per side). - Neutral / follow-ups: the HIP and Metal twins still lack these options (the row stays open for them;
float_ms_ssim_cudaenable_chromastays withT-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06). The SYCL, HIP and Metal twins share some of the defects (both SYCL SSIM twins force identical windows to 1,psnr_syclis notTEMPORAL, and the SYCL, HIP and Metalmotion_v2twins publish the raw SAD):T-GPU-TWIN-PARITY-GAPS-OUTSIDE-CUDA-2026-09-30. Guarded bytest_cuda_twin_option_parity(option tables device-free; CPU parity with a device, flat and single-pixel frames and--subsample 2included), the motion_v2 cases oftest_cuda_motion_tiny_frames, and the option, SSIM and stream cases oftest_cuda_kernel_source_contract.py.
References¶
- req: "there shouldnt be any gpu cpu rountrips" (user, relayed in the RC3 CUDA port brief, 2026-09-30).
- req: RC3 CUDA port brief (2026-09-30): "psnr enable_mse/enable_apsnr/reduced_hbd_peak/min_sse through the shared core/src/feature/psnr_score.h; integer ssim enable_db/clip_db; float_ssim enable_lcs/enable_db/clip_db and drop the phantom enable_chroma; float_motion motion_max_val; the debug motion score missing motion_fps_weight".
- req: #1637 review fix pass (2026-09-30): "Reproduce the CPU arithmetic (types and rounding points) instead of forcing 1"; "Keep it as an accepted no-op (documented as ignored, log once) ... Prefer the no-op".
- Research-1372.
- ADR-1365, ADR-1370, ADR-1183, ADR-1221, ADR-1302, ADR-1193, ADR-0564.