ADR-1365: SYCL PSNR, SSIM and float-motion twins take the CPU option tables¶
- Status: Accepted
- Date: 2026-09-29
- Deciders: lusoris
- Tags: sycl, gpu-parity, numerics, feature-extractor, fork-local
Context¶
T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26 lists CPU options that stale-branch merges removed from the GPU twins. For SYCL: psnr_sycl lacked min_sse, enable_mse, reduced_hbd_peak and enable_apsnr; integer_ssim_sycl and float_ssim_sycl lacked enable_db / clip_db, and float_ssim_sycl also enable_lcs; float_motion_sycl lacked motion_max_val. Scores stayed correct because the ADR-1183 gate computes a feature on the CPU when the twin cannot honour a model option, but those features never ran on the device, and naming the twin with the option failed with unknown option. RC3 owns the device path (ADR-1341, ADR-1352).
Constraints: every option must reproduce the CPU semantics exactly; SYCL kernels stay fp64-free (ADR-0220); no host round trip in the middle of a frame; one behaviour, one implementation (HISS-19). Two facts shaped the design. PSNR, the SSIM dB transform and the motion cap act on per-frame scalars the device already reduces, but enable_lcs needs per-pixel luminance, contrast and structure terms. And the CPU contract for a perfect score with enable_db is +inf, or the clip_db ceiling (ADR-1221): the fp32 SSIM twins did not return exactly 1 for identical windows, so turning on enable_db would have reported a finite dB value where the CPU reports the ceiling.
Decision¶
We will give the four SYCL twins the CPU option tables (same names, aliases, types, defaults, ranges and flags) and implement every option with the CPU's semantics:
- Scalar options on the host, from device-reduced sums. The PSNR option math (peak,
min_sseceiling, MSE to PSNR withuncapped, clip-aggregate APSNR) moves verbatim frominteger_psnr.cintocore/src/feature/psnr_score.h, which both the CPU extractor andpsnr_syclcall;psnr_syclgains aflushthat publishesapsnr_*. The SSIM twins convert with the existingnonfinite_score.hemitters and a newvmaf_ssim_max_db()for theclip_dbceiling.float_motion_syclapplies the CPU'smotion_clip()(fps weight, thenmotion_max_val) to every score it emits, the debugmotionscore included, and honoursmotion_force_zero, which it declared but ignored. enable_lcson the device.float_ssim_syclselects a second vertical-pass kernel that computes the per-pixel L, C and S terms ofiqa/ssim_tools.cin fp32 from the same moments and reduces them per work-group next to the SSIM value; the host sums the partials in double. The default kernel is unchanged in structure.- Identical windows score exactly 1. Both SSIM twins hold every product in a named temporary (icpx does not contract across statements, ADR-1358), sum the two variances as a pair, and return 1 when numerator and denominator are equal, so the numerator and denominator run the same operations mirrored. We accept that the default linear scores move by up to 1.1e-8 (2.2e-8 on identical frames, which reach exactly 1).
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Host options + device LCS kernel variant + mirrored SSIM arithmetic (chosen) | Every option on the device path; PSNR bit-exact; one PSNR implementation; identical frames follow ADR-1221 | Touches the upstream-mirror integer_psnr.c; default SSIM twin outputs move by up to 1.1e-8 | — |
Keep a SYCL-local copy of the PSNR math (as for uncapped before) | No change to the CPU file | Two implementations of the same arithmetic drift; the CUDA / HIP / Metal ports would add three more | HISS-19 |
| Compute L/C/S on the host from read-back moments | No new kernel | Five full-resolution float planes per frame over PCIe, a mid-frame host dependency | Violates device residency; far slower |
| Keep the SSIM formula and document the identical-frame dB gap | Default outputs bit-identical to before | With enable_db, identical frames report tens of dB below the CPU's ceiling, or finite instead of +inf | Breaks the CPU contract the option exists to reproduce |
| Detect identical frames on the host and substitute 1 | No kernel change | An extra full-frame compare; windows that are identical inside otherwise different frames still miss 1 | Treats the symptom; costs a pass |
Mark the options VMAF_OPT_FLAG_DEFAULT_ONLY (ADR-1316) | Named requests stop failing | The feature still never runs on the device | Status quo with better errors; RC3 asked for the device path |
Consequences¶
- Positive: models that set these options keep the four features on the SYCL device.
psnr_syclmatches the CPU bit for bit for every option; the SSIM and motion options agree within the twins' existing tolerances, and identical frames match the CPU exactly. Two latentfloat_motion_sycldivergences are fixed (motion_force_zero, unweighted debug score). - Negative: the default outputs of
integer_ssim_sycl/float_ssim_syclchange by up to 1.1e-8;integer_psnr.cnow calls a fork header, which an upstream sync has to respect (see the rebase notes). The dB form of SSIM magnifies an fp32 difference by4.34 / (1 - ssim), so near-perfect content can exceed the linear tolerance in dB; this is the twins' fp32 precision, not new error. - Neutral / follow-ups: the CUDA, HIP and Metal twins still lack these options (the row stays open for them); when ported they should call
psnr_score.handvmaf_ssim_max_db()rather than copy the math. Guarded bytest_sycl_twin_option_parityand theenable_lcscase intest_sycl_init_unwind.
References¶
- req: task brief for the SYCL part of
T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26(2026-09-29): "Implement each option on the device (or in the per-frame host scoring from device-reduced sums — no mid-frame round-trips) with exactly the CPU semantics". - Research-2127 — per-option parity on Arc B580 and UHD 770, mutation checks.
- ADR-1183 — the option gate; ADR-1193 —
uncapped; ADR-1221 — perfect-score dB contract; ADR-1302 — non-finite publication.