Skip to content

ADR-1373: CUDA PSNR, SSIM and motion twins take the CPU option tables and the CPU's arithmetic

  • Status: Accepted
  • Date: 2026-09-30
  • Deciders: lusoris
  • Tags: cuda, gpu-parity, numerics, feature-extractor, fork-local

Context

T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26 lists CPU options that stale-branch merges removed from the GPU twins. ADR-1365 closed the SYCL part. For CUDA: psnr_cuda lacked min_sse, enable_mse, reduced_hbd_peak and enable_apsnr; integer_ssim_cuda lacked enable_db / clip_db; float_ssim_cuda lacked enable_lcs, enable_db and clip_db and declared an enable_chroma the CPU float_ssim does not have (its kernel read luma only, so the option did nothing); float_motion_cuda lacked motion_max_val and emitted its debug motion score without motion_fps_weight. Scores stayed correct, because the ADR-1183 gate keeps a feature on the CPU when the twin cannot honour a model option, but those features never ran on the device, and naming the twin with the option failed with unknown option.

Two reviews of the first version of this change (#1637) found more in the same twins:

  • float_ssim_cuda forced identical windows to exactly 1. The CPU does not: iqa/ssim_tools.c computes l * c * s with double numerators over fp32-rounded denominators and returns the frame mean rounded to fp32, so identical flat frames (64x64 at 128, 8-bit; 512 at 10-bit) report 72.247199 dB with enable_db, where the forced 1 gave +inf (and 88 dB against 72.247 with clip_db).
  • integer_ssim_cuda grouped each pixel's term as w * (a * b / den); the CPU computes ((w * a) * b) / den. The two round differently wherever the window weight is not a power of two, on 841 of 1952 border pixels at 161x91. NVCC's default contraction also fused the constant c1 into the denominator only.
  • motion_v2_cuda published the unweighted, uncapped SAD, where the CPU publishes MIN(sad * motion_fps_weight, motion_max_val); its motion2_v2 was weighted but never capped, the motion3_v2 stamp blended the unweighted SAD, and a one-frame input emitted no motion2_v2 / motion3_v2 where the CPU emits 0 / 0.
  • psnr_cuda lacked the CPU psnr's VMAF_FEATURE_EXTRACTOR_TEMPORAL flag, so with --subsample N its apsnr_* summed one frame in N. Its chroma accumulators were zeroed on the readback stream while the chroma kernels ran on the picture stream, the race kernel_template.h documents for plane 0.

Constraints: every option must reproduce the CPU semantics exactly; no host round trip in the middle of a frame and one wait per frame, at collect (req); one behaviour, one implementation (HISS-19); public options evolve additively (HISS-14).

Decision

We will give the CUDA twins their CPU extractor's option tables (same names, aliases, types, defaults, ranges and flags) and reproduce the CPU's arithmetic, types and rounding points:

  1. Scalar options on the host, from device-reduced sums, through the CPU's helpers. psnr_cuda turns each plane's device SSE into psnr_*, mse_* and (in a new flush) apsnr_* with core/src/feature/psnr_score.h, which the CPU integer_psnr.c calls too; its local copy of the MSE-to-PSNR arithmetic is deleted. psnr_cuda is TEMPORAL like the CPU psnr, and zeroes every plane's accumulator on the picture stream its kernels run on. The SSIM twins apply enable_db / clip_db through the nonfinite_score.h emitters with vmaf_ssim_max_db(). float_motion_cuda applies the CPU's motion_clip() (fps weight, then motion_max_val) to every score it emits, the debug motion score included, and emits the CPU's motion3 with its motion_blend_factor / motion_blend_offset options (motion_blend_clip() of motion2, frame 0 from the first SAD, the tail from flush), which the twin did not emit at all (T-GPU-FLOAT-MOTION3-MISSING-2026-09-30, found on the device). motion_v2_cuda publishes the CPU's weighted, capped SAD score and derives motion2_v2 / motion3_v2 from it exactly as integer_motion_v2.c::flush does, one-frame input included.
  2. float_ssim_cuda computes the CPU's l * c * s. Each pixel follows iqa/ssim_tools.c::ssim_variance_scalar and iqa/ssim_accumulate_lane.h (the helper every CPU SIMD variant shares) operand for operand: variances clamped at zero, l and c as double numerators over fp32 denominators, s in fp32, the product in double, each rounding spelled with an intrinsic (__fmul_rn, __fadd_rn, __ddiv_rn, ...) so NVCC cannot fuse it. Block partials are doubles and the host rounds each frame mean to fp32, as iqa_ssim() returns it (as SYCL does since ADR-1370). No window is forced to 1: identical frames score whatever the CPU scores. enable_lcs launches a second pass-2 kernel that reduces the same l, c and s next to the score.
  3. integer_ssim_cuda computes the CPU's term. Its kernel builds with --fmad=false, as the HIP twin builds with -ffp-contract=off (ADR-0564), and groups each term as the CPU does, ((w * a) * b) / den. Every per-pixel term is then the CPU's bit for bit. The frame sum is not: the CPU adds the terms row by row, the kernel per warp, per block and then on the host, so the score can differ by a double rounding of the sum. We keep the parallel reduction and drop any bit-exactness claim.
  4. enable_chroma stays on float_ssim_cuda, accepted and ignored. It was a public option of the twin, so removing it would break --feature float_ssim_cuda=enable_chroma=... (HISS-14). It defaults to false, names no feature, logs that it has no effect when set, and is the one twin-only option test_cuda_twin_option_parity allows.

Alternatives considered

Option Pros Cons Why not chosen
Host options from device sums through the CPU helpers + the CPU's per-pixel SSIM arithmetic and fp32 frame mean (chosen) Every option on the device path; PSNR bit-exact; one PSNR implementation for CPU, SYCL and CUDA; identical frames match the CPU instead of a guessed 1 The float SSIM combine and its partials are in double (fp64 on the device); default float_ssim_cuda output moves towards the CPU —
Force identical windows to exactly 1 (the first version, and ADR-1365 for SYCL) Simple; +inf for identical noise frames Wrong wherever the CPU is not exactly 1: identical flat frames report +inf against the CPU's 72.247 dB Refuted by a CPU run (#1637 review)
Include iqa/ssim_accumulate_lane.h in the device code One source for CPU and CUDA The CPU header relies on the host compiler not contracting; making it device-safe means rounding macros in a CPU hot-path header shared by three SIMD variants The CUDA copy spells each rounding as an intrinsic and cites the header; the contract test pins its rounding points
Accumulate the float SSIM convolution in double like the CPU The moments too would be the CPU's Roughly 110 fp64 adds per pixel, at 1/64 rate on consumer GPUs Not required for the defects found; left for a measured follow-up
Make the integer SSIM frame sum row-major sequential Bit-exact with the CPU A serial pass over every pixel Throughput; the remaining difference is a double rounding of the sum
Keep a CUDA-local copy of the PSNR math No shared-header dependency Two implementations of the same arithmetic drift apart HISS-19
Compute L/C/S on the host from read-back moments No new kernel Five full-resolution float planes over PCIe per frame, a mid-frame host dependency Violates the no-round-trip requirement
Remove enable_chroma (the first version) The twin's table equals the CPU's Breaks a public command line without the ! / Migration: HISS-14 asks for Keep it as an accepted no-op
Mark the options VMAF_OPT_FLAG_DEFAULT_ONLY (ADR-1316) Named requests stop failing The features still never run on the device Status quo with better errors

Consequences

  • Positive: models that set these options keep the features on the CUDA device. psnr_cuda matches the CPU bit for bit for every option, also under --subsample; motion_v2_cuda matches the CPU for motion_fps_weight, motion_max_val and one-frame inputs; float_ssim_cuda reproduces the CPU's enable_db result on identical frames, finite or +inf; integer_ssim_cuda's per-pixel terms equal the CPU's (host replay: 0 mismatches over 9,000,000 terms). Measured on an RTX 4090 on 2026-09-30 with the Netflix pair: every PSNR option gives identical psnr_*, mse_* and apsnr_*, --subsample 2 included; motion_v2 and float_motion with motion_fps_weight / motion_max_val are identical, motion3 included, and float_motion's motion3 is within 2.8e-6 of the CPU at the defaults and 1.4e-6 with the blend options (the same run on the final head on 2026-10-01 gave the same numbers); ssim with enable_db is within 7.3e-13 dB of the CPU and float_ssim within 6.9e-6 dB; identical flat 64x64 frames give 72.247198959355487 dB on both sides, as the host replay predicted.
  • Negative: default float_ssim_cuda and integer_ssim_cuda outputs move towards the CPU by an fp32 / double rounding. The float SSIM combine now runs in fp64; on an RTX 4090 at 4K its cost does not show above the frame upload (4.2 ms per frame against 4.6 for a master build, 4.1 with enable_lcs; (t(200) - t(2)) / 198, median of 7, interleaved on a shared host). integer_ssim_cuda still differs from the CPU by a double rounding of the frame sum; on identical frames with a side below 12 pixels that can put one side at +inf and the other near 156 dB with enable_db and no clip_db (host replay: 72 of 57,600 such frames of 1 to 24 pixels per side).
  • Neutral / follow-ups: the HIP and Metal twins still lack these options (the row stays open for them; float_ms_ssim_cuda enable_chroma stays with T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06). The SYCL, HIP and Metal twins share some of the defects (both SYCL SSIM twins force identical windows to 1, psnr_sycl is not TEMPORAL, and the SYCL, HIP and Metal motion_v2 twins publish the raw SAD): T-GPU-TWIN-PARITY-GAPS-OUTSIDE-CUDA-2026-09-30. Guarded by test_cuda_twin_option_parity (option tables device-free; CPU parity with a device, flat and single-pixel frames and --subsample 2 included), the motion_v2 cases of test_cuda_motion_tiny_frames, and the option, SSIM and stream cases of test_cuda_kernel_source_contract.py.

References

  • req: "there shouldnt be any gpu cpu rountrips" (user, relayed in the RC3 CUDA port brief, 2026-09-30).
  • req: RC3 CUDA port brief (2026-09-30): "psnr enable_mse/enable_apsnr/reduced_hbd_peak/min_sse through the shared core/src/feature/psnr_score.h; integer ssim enable_db/clip_db; float_ssim enable_lcs/enable_db/clip_db and drop the phantom enable_chroma; float_motion motion_max_val; the debug motion score missing motion_fps_weight".
  • req: #1637 review fix pass (2026-09-30): "Reproduce the CPU arithmetic (types and rounding points) instead of forcing 1"; "Keep it as an accepted no-op (documented as ignored, log once) ... Prefer the no-op".
  • Research-1372.
  • ADR-1365, ADR-1370, ADR-1183, ADR-1221, ADR-1302, ADR-1193, ADR-0564.