Research-1407: contraction off in every HIP kernel — parity and cost per twin on a gfx1036¶
- Status: Active
- Workstream: RC3 HIP lane,
T-HIP-FP-CONTRACT-DEFAULT-2026-09-29; decision in ADR-1407 - Last updated: 2026-10-01
Question¶
hipcc contracts a * b + c into a fused multiply-add for device code by default, and all but three HIP kernels were built that way. What does each twin's output and cost do when every kernel is built with -ffp-contract=off, and can a test tell on the device whether a kernel was contracted?
Sources¶
core/src/meson.build(hip_kernel_sources, the formerhip_cu_extra_flags,hip_strict_fp_args).- ADR-1367 and Research-1367, the SYCL precedent.
- Host:
ryzen-4090-arc, AMD gfx1036 iGPU, ROCm 7.2.4 (hipccHIP 7.2.53211, AMD clang 22.0.0git), Linux 7.2.8. - Builds:
meson setup build-hip core -Denable_hip=true -Denable_hipcc=true -Dhip_gfx_targets=gfx1036 -Denable_cuda=false -Denable_sycl=false --buildtype=release -Db_lto=false, once atorigin/master0e4368cdb ("before") and once with the list ("after").
Findings¶
A device probe separates the builds¶
test_hip_fp_arith_contract runs one kernel that computes a * b + c, a / b and sqrtf(|a|) for 1 048 576 random operand triples (random sign and mantissa, exponents -20 to 20) and 2 197 boundary combinations, and compares each result with the correctly rounded host value (the operation in fp64, rounded once).
| Probe compiled with | a * b + c differing | a / b | sqrtf |
|---|---|---|---|
hip_strict_fp_args | 0 | 0 | 0 |
| hipcc defaults | 126 577 | 0 | 0 |
-ffp-contract=off -fno-hip-fp32-correctly-rounded-divide-sqrt | 0 | 303 181 | 158 765 |
So contraction is the only thing hipcc's defaults change on this device, and the division flag, though a default, is worth pinning.
Parity per twin, before and after¶
Max abs diff against --backend cpu --threads 16 at --precision max, with the number of frames on which every output is bit-identical. 576x324 is the Netflix pair (48 frames), 4K is BBB 3840x2160 (22 frames). float_ssim at 4K runs with scale=1, the only scale the twin implemented at this commit.
| Twin | 576x324 before | 576x324 after | 4K before | 4K after |
|---|---|---|---|---|
vif_hip | 5.36e-7 (0/48) | 5.36e-7 (0/48) | 2.98e-7 (0/22) | 2.98e-7 (0/22) |
adm_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
motion_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
motion_v2_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
psnr_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
float_psnr_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
float_moment_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
integer_ssim_hip | 2.32e-14 (0/48) | 2.32e-14 (0/48) | 5.58e-13 (0/22) | 5.58e-13 (0/22) |
cambi_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
ssimulacra2_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
speed_chroma_hip | 1.19e-6 (47/48) | 1.19e-6 (47/48) | 1.43e-6 (20/22) | 1.43e-6 (20/22) |
speed_temporal_hip | 0 (48/48) | 0 (48/48) | 0 (22/22) | 0 (22/22) |
float_adm_hip | 2.50e-5 (0/48) | 2.53e-6 (0/48) | 6.04e-6 (0/22) | 1.28e-5 (0/22) |
float_vif_hip | 2.72e-5 (0/48) | 3.82e-5 (0/48) | 5.74e-6 (0/22) | 7.02e-6 (0/22) |
float_motion_hip | 3.01e-6 (1/48) | 3.12e-6 (1/48) | 2.37e-5 (1/22) | 2.36e-5 (1/22) |
float_ssim_hip | 1.79e-7 (0/48) | 1.19e-7 (7/48) | 7.21e-6 (0/22) | 4.83e-6 (0/22) |
integer_ms_ssim_hip | 6.89e-8 (0/48) | 5.53e-8 (0/48) | 5.82e-7 (0/22) | 1.22e-6 (0/22) |
ciede_hip | 1.133e-5 (0/48) | 1.134e-5 (0/48) | 1.65e-6 (0/22) | 1.43e-6 (0/22) |
psnr_hvs_hip | 8.37e-5 (0/48) | 8.37e-5 (0/48) | 1.10e-2 (0/22) | 1.10e-2 (0/22) |
The first twelve twins produce the same output before and after on every frame (the after build against the before build: 48 of 48 and 22 of 22 identical). Their kernels are integer, or were already built strict, or have no product feeding an add. vif_hip's 5.4e-7 is therefore not a contraction effect.
Mean abs diff for the seven twins whose output changes:
| Twin | 576x324 mean before -> after | 4K mean before -> after |
|---|---|---|
float_adm_hip | 1.66e-7 -> 4.34e-8 | 2.03e-7 -> 1.84e-7 |
float_vif_hip | 3.07e-6 -> 3.18e-6 | 1.52e-6 -> 1.39e-6 |
float_motion_hip | 9.09e-7 -> 9.16e-7 | 2.07e-6 -> 2.07e-6 |
float_ssim_hip | 1.18e-7 -> 5.84e-8 | 2.68e-6 -> 1.90e-6 |
integer_ms_ssim_hip | 2.45e-8 -> 1.97e-8 | 3.42e-7 -> 4.03e-7 |
ciede_hip | 1.02e-5 -> 1.02e-5 | 1.31e-6 -> 7.84e-7 |
psnr_hvs_hip moves by at most 1.3e-6 between the builds, far below its residual against the CPU's running fp32 sum (T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30).
The picture matches SYCL's (ADR-1367): float_adm improves tenfold at 576x324, float_vif's worst frame gets worse while its mean stays, and the rest move inside their existing residual. cross_backend_parity_gate.py --backends cpu hip on the Netflix pair passes its 17 runnable cells before and after; the motion cell aborts on a metric name in both.
Cost at 4K¶
(t(22) - t(2)) / 20 ms per frame, whole CLI. The host ran other agents' builds throughout (load average 12 to 26), so the light twins were re-measured with the two builds interleaved, seven repetitions each, and the heavier ones with five; the rest are the median of three runs per build, not interleaved.
| Twin | Before | After | Reps |
|---|---|---|---|
float_psnr_hip | 4.49 | 4.44 | 7, interleaved |
psnr_hip | 8.62 | 8.44 | 7, interleaved |
motion_hip | 12.92 | 13.08 | 7, interleaved |
motion_v2_hip | 13.98 | 13.81 | 7, interleaved |
float_motion_hip | 20.38 | 19.97 | 7, interleaved |
adm_hip | 76.76 | 72.96 | 5, interleaved |
float_adm_hip | 77.95 | 78.75 | 5, interleaved |
float_vif_hip | 82.94 | 86.49 | 5, interleaved |
float_ssim_hip (scale=1) | 93.60 | 84.76 | 5, interleaved |
vif_hip | 141.17 | 139.15 | 5, interleaved |
integer_ms_ssim_hip | 164.61 | 157.06 | 5, interleaved |
speed_chroma_hip | 5.46 | 5.50 | 3 |
float_moment_hip | 8.76 | 8.00 | 3 |
speed_temporal_hip | 14.26 | 14.60 | 3 |
psnr_hvs_hip | 18.26 | 18.60 | 3 |
ciede_hip | 73.60 | 74.68 | 3 |
cambi_hip | 99.49 | 98.79 | 3 |
integer_ssim_hip | 136.9 | 139.1 | 3 |
ssimulacra2_hip | 3765 | 3649 | 3 |
float_vif_hip is the one consistent change: its five after samples (84.7 to 87.2) all sit above its five before samples (81.6 to 84.1), 4.3% on the median. Every other twin's difference is inside its own run-to-run spread. No twin is 10% slower, ADR-1367's bar for an exemption.
Wrong frames during the measurement were the device, not the flag¶
Five runs showed frames far from the CPU: vif_hip 2.1e-3 on one 4K frame, float_moment_hip 1.5e4 on one frame, and adm_hip 1.5e-2 on one or two frames in three runs. A repeat of each was clean, and an interleaved repeat of adm_hip at 4K gave 15 of 15 clean runs with the list and 14 of 15 without. That is the gfx1036 losing a run of a stream's commands (T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01), at a higher rate than that row's one frame in 10^4 while the host was loaded. The tables above are from clean runs.
Alternatives explored¶
See the decision matrix in ADR-1407.
Open questions¶
float_vif_hip: a systematic 3e-6 mean offset and a worst frame at 3.8e-5 remain with strict arithmetic; SYCL's ADR-1367 attributes the same offset tolog2and to the CPU's fp641 + x / y. Not isolated here.- The CUDA twins' counterpart is
T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29.