ADR-1441: float_ssim_hip forms its window sums through the arithmetic float_ms_ssim_hip shares with the CPU and returns the CPU's score bit for bit¶
- Status: Accepted
- Date: 2026-10-01
- Deciders: lusoris
- Tags:
hip,ssim,gpu-parity,numerics,testing,ci,rc3,fork-local
Context¶
float_ssim and float_ms_ssim run the same core on the CPU, iqa_ssim(): iqa_convolve() filters five planes with the separable 11-tap Gaussian, and ssim_accumulate_default_scalar() forms l, c and s per window. The convolution multiplies each tap in fp32, adds the eleven products in fp64 and rounds to fp32 once per pass.
float_ssim_hip formed l, c and s in the CPU's types since ADR-1382 and decimates as the CPU does since ADR-1405. Its two convolution passes added the products in an fp32 running sum, which rounds at every tap. Measured on a gfx1036 at --precision max against the CPU on origin/master 80c5a0332 (the RC3 sweep, #1772): 27 of 178 frames had the CPU's score, the others were up to 4.8e-7 away. With enable_lcs, float_ssim_l was identical on every frame and float_ssim_c / float_ssim_s were up to 5.4e-7 off: the window means were right, the window sums of squares and products were not.
float_ms_ssim_hip had the same defect and removed it in ADR-1403: its arithmetic lives in core/src/feature/hip/integer_ms_ssim/ms_ssim_arith.h, which carries the fp64 sum as an exact fp32 pair (a two-sum, six fp32 operations per add) and is compiled by the kernels and by a host test that holds it against the CPU.
Decision¶
We will compute float_ssim_hip's window sums and its l / c / s through that header: vmaf_hip_ms_ssim_horizontal() for pass 1, vmaf_hip_ms_ssim_vertical() for pass 2 and vmaf_hip_ms_ssim_lcs() for the terms. The kernel keeps no tap table, no window sum and no SSIM term of its own. float_ssim and float_ssim_lcs: hip are declared exact twins.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| The shared header's exact fp32 pair sums (this ADR) | The CPU's values; one implementation of the window arithmetic for both twins, already held against the CPU by test_hip_ms_ssim_arith | Six fp32 operations per add instead of one: 17 % more time per 1920x1080 frame and 7 % per 3840x2160 frame at the default scale, 32 % at scale=1 | Chosen |
| fp64 accumulators | Simplest to read | ADR-1403 measured them on this device: 299 instead of 173 ms per 3840x2160 float_ms_ssim frame | Slower than the pair |
| Keep fp32 sums and the 5e-5 tolerance | No cost | 151 of 178 frames differ from the CPU; the twin stays the one SSIM-family HIP twin that is not exact | RC3 asks for the bits |
| A second copy of the pair arithmetic in the float_ssim kernel | No cross-directory include | Two copies of arithmetic that must both follow the CPU | One behaviour, one implementation |
Consequences¶
- Positive: measured on a gfx1036 at
--precision max, every output of every frame equals--backend cpu: 178 of 178 forfloat_ssim(27 before) and 712 of 712 withenable_lcs(Netflix 576x324 at 8, 10, 12 and 16 bits and as 10-bit 4:2:2, both 1920x1080 checkerboard pairs, Sparks 480x270 at 10 bits, 48 frames of BBB 3840x2160, full-range noise at four depths, a bright 16-bit 1080p pair). Also withscale=1,scale=3,enable_dbandenable_dbplusclip_dbon five fixtures. - Positive: the kernel has no window arithmetic of its own; a change to
iqa_convolve()or to the accumulate routine is one edit in the header for both twins. - Negative: the twin is slower. Steady state on the gfx1036, medians of 11 interleaved pairs of runs: 1.72 to 2.02 ms per 1920x1080 frame (+17 %) and 4.86 to 5.18 ms per 3840x2160 frame (+7 %) at the default scale (4 and 8); with
enable_lcs1.98 to 2.31 ms at 1080p (+17 %); withscale=117.7 to 23.4 ms at 1080p (+32 %) and 82.3 to 109.6 ms at 3840x2160 (+33 %).T-HIP-FLOAT-SSIM-EXACT-THROUGHPUT-2026-10-01. - Negative: stored
float_ssim_hipscores change by up to 4.8e-7. - Neutral / follow-ups: the per-frame sum of the terms is fp64 in block order and the mean is rounded to fp32, as for
float_ms_ssim; the rounding absorbs the order unless the sum lies within its own rounding error of an fp32 boundary (by estimate a few means in a million; the sweep in #1772 states it forfloat_ms_ssim). Guards:test_hip_float_ssim_parityand its 960x540 registration (==on every score of every case; the first case fails on the old twin),test_hip_ms_ssim_arith(the header against the CPU, no device) and three planted regressions intest_hip_kernel_source_contract.py.
References¶
req(maintainer brief for the second HIP lane, 2026-10-01): "Not identical: find the cause [...], fix it the way the CUDA / SYCL twin of the same feature was fixed if one was (read that ADR and reuse its helpers [...])" and "note timing regressions above 5 % and stop to report if a fix costs more than 20 %."- ADR-1403 (the shared arithmetic), ADR-1382, ADR-1405, ADR-1397 (the exact cell), ADR-0214.
docs/state.md:T-HIP-FLOAT-SSIM-NOT-CPU-ARITHMETIC-2026-10-01,T-HIP-FLOAT-SSIM-EXACT-THROUGHPUT-2026-10-01.