ADR-1382: HIP PSNR, SSIM and float-motion twins take the CPU option tables¶
- Status: Accepted
- Date: 2026-09-30
- Deciders: lusoris
- Tags: hip, gpu-parity, numerics, feature-extractor, fork-local
Context¶
T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26 lists CPU options the GPU twins lack. ADR-1365 closed the SYCL part; for HIP: psnr_hip lacked min_sse, enable_mse, reduced_hbd_peak and enable_apsnr; integer_ssim_hip lacked enable_db / clip_db; float_ssim_hip lacked enable_db, clip_db and enable_lcs; float_motion_hip lacked motion_max_val and emitted its debug VMAF_feature_motion_score without motion_fps_weight (the CPU emits motion_clip(score)). Scores stayed correct because the ADR-1183 gate kept any feature with such an option on the CPU, but those features never ran on the device, and naming the twin with the option failed with unknown option. The row asks ports to call psnr_score.h and vmaf_ssim_max_db() instead of copying the math, and to make identical SSIM windows score exactly 1 as ADR-1365 does.
The HIP twins differ from the SYCL ones in two ways that shape the port. HIP kernels may use fp64, and integer_ssim_hip already evaluates the CPU's per-pixel expression in double with contraction off, so its per-pixel terms are the CPU's own doubles. float_ssim_hip builds with hipcc's default -ffp-contract=fast, which fuses products across statements, so named temporaries alone do not keep its numerator and denominator mirrored. HIP-wide fp-contract settings are a separate port (the SYCL counterpart is #1630).
Decision¶
We give the four HIP twins the CPU option tables (same names, aliases, types, defaults, ranges and flags) and implement every option with the CPU's semantics:
- Scalar options on the host from device-reduced sums.
psnr_hipturns each plane's device-reduced integer SSE intopsnr_*,mse_*and theapsnr_*aggregates with the helpers ofcore/src/feature/psnr_score.hthat the CPU extractor calls (vmaf_psnr_peak(),vmaf_psnr_max(),vmaf_psnr_from_mse(),vmaf_psnr_aggregate()), accumulating the APSNR totals incollect()and publishing them in a newflush().integer_ssim_hip/float_ssim_hippassenable_dbandvmaf_ssim_max_db()to the sharednonfinite_score.hemitters.float_motion_hipsends every score it emits, the debugmotionscore and the flush tail included, through the CPU'smotion_clip()(fps weight, thenmotion_max_val). enable_lcson the device.float_ssim_hipselectscalculate_ssim_hip_vert_combine_lcs, a pass-2 variant that also computes the per-pixel L, C and S ofiqa/ssim_tools.cin the CPU's own types (fp32 inputs and clamped variances, double L and C, fp32 S with the flat-window covariance clamp) and reduces them per block in double; the host sums the blocks and emitsfloat_ssim_{l,c,s}with the score in CPU order.- Identical frames follow the CPU.
integer_ssim_hipkeeps the CPU's double expression operand for operand and returns the window weight when numerator and denominator factors are equal (always, for 8- and 10-bit input): the CPU's raster-order sum absorbs the ulp its quotient leaves, the twin's per-block tree does not, and with the rule identical frames from 3x3 up score exactly 1 on both (a host replay of the CPUcalc_ssim(); 1x1 and 2x2 remain,T-HIP-INTEGER-SSIM-TINY-IDENTICAL-DB-2026-09-30).float_ssim_hipdoes not force anything: its pass 2 forms each pixel's term as the CPU'sssim_accumulate_default_scalar()does,l * c * sin double from the CPU-typed factors (ssim_lcs(), contraction off), sums one double per block, and the host rounds the frame mean to fp32 asiqa_ssim()returns it. The CPU itself is not exactly 1 on every identical frame (72.247 dB = 1 - 2^-24 on a flat 64x64 frame, from its fp32 luminance denominator), so a forced 1 would disagree with it; the reproduced arithmetic agrees wherever the vertical moments agree. - The rest of the CPU contract of these twins, found in review.
motion_v2_hipstores its SAD as the CPU does,MIN(score * motion_fps_weight, motion_max_val), and foldsmotion2_v2/motion3_v2from the stored value; a one-frame run emits both as 0.psnr_hipcarriesVMAF_FEATURE_EXTRACTOR_TEMPORALlike the CPUpsnr, so--subsample N > 1does not drop frames from theapsnr_*totals.motion_hipdefaultsdebugto false and emitsVMAF_integer_feature_motion_sad_scoreevery frame, the CPUmotion's set.scripts/ci/cross_backend_parity_gate.pygains thehipbackend and afloat_ssim_lcscell (float_ssimwithenable_lcs=true, 5e-5), andtest_hip_twin_option_parityholdsfloat_ssimand its L / C / S to that gate tolerance. - One HIP error translator. The four twins drop their private copies of the
hipError_tto errno mapping and callvmaf_hip_rc_to_errno()(core/src/hip/common.h), whichkernel_template.cnow defines; the mappings were identical (HISS-19).
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Host options + device LCS variant + CPU per-pixel float SSIM + exact-1 integer windows (chosen) | Every option on the device path; PSNR bit-exact through the CPU's own helpers; both SSIM twins compute each pixel's term the CPU's way; identical frames agree with the CPU | New flush for psnr_hip; float_ssim_hip pass 2 works in double (as its enable_lcs variant already did) | — |
Force an identical float_ssim window to exactly 1 (ADR-1365's SYCL form) | Identical frames always +inf | The CPU reports 72.247 dB on a flat identical frame, so the twin disagrees with its reference exactly where enable_db magnifies the difference; keeps the combined-formula residual (T-SYCL-FLOAT-SSIM-COMBINED-FORMULA-RESIDUAL-2026-09-29) | The CPU is the reference; HIP has correctly rounded sqrtf and division, so the product form needs no equal-operand guard |
Port ADR-1365's fp32 SSIM form to integer_ssim_hip too | Same code shape as SYCL | Throws away a double per-pixel term that is already the CPU's exactly | The HIP kernel is closer to the CPU as it is |
Build ssim_score.hip with -ffp-contract=off | No pragma | Changes the moment passes too; overlaps the pending fp-contract port | Out of scope here; the pragma is local to the per-pixel formula |
| Compute L/C/S on the host from read-back moments | No new kernel | Five full-resolution planes over PCIe per frame, a host round trip mid-frame | Violates device residency (maintainer: no GPU-CPU round trips) |
| Keep a HIP-local copy of the PSNR math | No include of a CPU header | Duplicate arithmetic drifts; the row asks for psnr_score.h | HISS-19 |
Mark the options VMAF_OPT_FLAG_DEFAULT_ONLY (ADR-1316) | Named requests stop failing | The feature still never runs on the device | Status quo with better errors |
Consequences¶
- Positive: models that set these options keep
psnr,ssim,float_ssimandfloat_motionon the HIP device;psnr_hipis expected to match the CPU bit for bit with every option (integer SSE, same host helpers), the SSIM and motion options within the twins' existing tolerances, and identical frames exactly.float_motion_hip's debug score now carries the fps weight like the CPU. - Negative: not measured on AMD hardware in this change;
test_hip_twin_option_paritycarries the per-option device checks and skips without a device.float_ssim_hip's default score moves from the combined Wang formula to the CPU's product form (the SYCL measurements put the formula residual at up to 7.8e-5 and the product form within 9e-7 of the CPU); bit-exact identical-frame dB also needs the vertical moments to match the CPU's, which the pending fp-contract work decides.float_motion_hipstill lacksmotion3and five options (T-HIP-FLOAT-MOTION-MOTION3-OPTIONS-2026-09-30). - Neutral / follow-ups: the CUDA and Metal parts of the row stay open. Guarded by
test_hip_twin_option_parity(the option-table and unknown-option checks need no device),test_gpu_psnr_option_parity_contract.py,test_gpu_option_alias_contract.py(float_motion_hipmotion_max_valaliasmmxv), and the option cases oftest_hip_kernel_source_contract.py.
References¶
- req: RC3 port brief (2026-09-30): "psnr options via core/src/feature/psnr_score.h; integer_ssim_hip enable_db/clip_db; float_ssim_hip enable_lcs; float_motion motion_max_val; debug motion fps weight", and "there shouldnt be any gpu cpu rountrips".
- ADR-1365 and Research-2127 — the SYCL port this mirrors.
- Research-1377 — why the integer SSIM term needs the identical-window rule.
- ADR-1183, ADR-1193, ADR-1221, ADR-1302, ADR-0564.