ADR-1407: Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root¶
- Status: Accepted
- Date: 2026-10-01
- Deciders: lusoris
- Tags: hip, gpu, numerics, build, rc3, fork-local
Context¶
hipcc (amdclang++) compiles device code with -ffp-contract=fast: a * b + c becomes one fused multiply-add, across statements too. The CPU reference build (gcc/clang, x86-64 baseline, no -mfma in scalar TUs) does not contract, so a contracted kernel rounds differently from the extractor it mirrors. core/src/meson.build turned contraction off for three kernels only, through a per-kernel table (hip_cu_extra_flags, ADR-0594): ssimulacra2_blur, integer_ssim_score and speed_pipeline. The other eighteen kernels contracted, and float_ssim carries #pragma clang fp contract(off) in the part that mirrors the CPU (ADR-1382). Two floating-point behaviours for one kind of kernel is a HISS-19 defect, and every new exact kernel had to remember to join the table (T-HIP-FP-CONTRACT-DEFAULT-2026-09-29).
ADR-1367 settled the same question for SYCL: one strict line for every feature TU, a device test, and an exemption only for a twin that gets more than 10% slower at 4K with no parity gain. The row asks for the HIP measurement on an AMD device, which ADR-1367's host did not have.
Decision¶
Every HIP kernel is compiled with one list, defined once in core/src/meson.build between the VMAF HIP strict FP policy markers:
hip_strict_fp_args = ['-ffp-contract=off', '-fhip-fp32-correctly-rounded-divide-sqrt']
The second flag is hipcc's default; it is passed explicitly so that a toolchain default cannot move a score. The per-kernel table is removed: a kernel listed in hip_kernel_sources gets the policy, and there is no way to exempt one. A kernel that needs a fused operation writes fmaf() / fma(). The existing #pragma clang fp contract(off) lines stay; they are redundant now and harmless.
Two tests pin it. test_hip_strict_fp_policy.py reads the build files without a device: the list holds both flags, every HSACO compile gets it, no per-kernel table exists, no kernel source turns contraction back on, and the device probe is compiled with the same list. test_hip_fp_arith_contract runs a probe kernel compiled with that list on the device and compares 1 048 576 random and 2 197 boundary a * b + c, a / b and sqrtf(|a|) results with correctly rounded host values. Its host side is the SYCL test's source (test_sycl_fp_arith_contract.c) built for the HIP probe, so both backends are held to one set of operands and references.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Keep the per-kernel table | No other twin's output changes | Two FP behaviours for one kind of kernel; every new exact kernel must remember to join it | HISS-19; the row asks for one shared list |
| Contraction off only, leave division and square root to the toolchain default | One flag | The default is correct today, but the probe shows what a changed default costs: 303 181 of 1 048 576 divisions and 158 765 square roots differ with -fno-hip-fp32-correctly-rounded-divide-sqrt | Pinning the default is free |
Exempt float_vif_hip, whose worst frame gets worse and which is 4% slower | Keeps its old numbers | Its mean barely moves (3.07e-6 to 3.18e-6 at 576x324, better at 4K), it stays inside the 5e-5 gate, and 4% is under ADR-1367's 10% bar | Nothing crosses the bar |
#pragma clang fp contract(off) in every kernel source | No build change | Eighteen files to edit and every future one to remember; a pragma covers only the block it sits in | The flag covers everything |
| One strict list for every kernel (chosen) | One behaviour; a device test enforces it; the build test forbids exemptions | Seven twins' outputs change (each inside its ADR-0214 tolerance) | Chosen |
Consequences¶
Measured on ryzen-4090-arc (gfx1036, ROCm 7.2.4, hipcc 7.2.53211, AMD clang 22.0.0git), origin/master 0e4368cdb against the same tree with the list.
- The device test separates the three builds. With the list: 0 mismatches. With hipcc's defaults: 126 577 of 1 048 576 multiply-adds differ. With
-ffp-contract=off -fno-hip-fp32-correctly-rounded-divide-sqrt: 303 181 divisions and 158 765 square roots differ. - Per-twin max abs diff against
--backend cpuat--precision max, before -> after (576x324: Netflix pair, 48 frames; 4K: BBB 3840x2160, 22 frames; mean in brackets): - better:
float_adm2.50e-5 -> 2.53e-6 (1.66e-7 -> 4.34e-8) at 576x324;float_ssim1.79e-7 -> 1.19e-7, 7 of 48 frames now bit-identical, and at 4Kscale=17.21e-6 -> 4.83e-6;ciedeat 4K 1.65e-6 -> 1.43e-6 (1.31e-6 -> 7.84e-7);float_ms_ssim6.89e-8 -> 5.53e-8 at 576x324;float_motionat 4K 2.37e-5 -> 2.36e-5. - unchanged output, bit for bit:
vif,adm,motion,motion_v2,psnr,float_psnr,float_moment,ssim,cambi,ssimulacra2,speed_chroma,speed_temporal. Of theseadm,motion,motion_v2,psnr,float_psnr,float_moment,cambi,ssimulacra2andspeed_temporalequal the CPU on every frame before and after. - worse max, inside tolerance:
float_vif2.72e-5 -> 3.82e-5 at 576x324 (mean 3.07e-6 -> 3.18e-6) and 5.74e-6 -> 7.02e-6 at 4K (mean 1.52e-6 -> 1.39e-6);float_admat 4K 6.04e-6 -> 1.28e-5 (mean 2.03e-7 -> 1.84e-7);float_ms_ssimat 4K 5.82e-7 -> 1.22e-6 (mean 3.42e-7 -> 4.03e-7);float_motionat 576x324 3.01e-6 -> 3.12e-6;ciedeat 576x324 1.133e-5 -> 1.134e-5.psnr_hvschanges below its own residual (8.37e-5 and 1.10e-2 against the CPU's running fp32 sum, before and after). The same twins moved the same way on SYCL (ADR-1367:float_adm2.50e-5 -> 2.53e-6,float_vif2.71e-5 -> 3.81e-5): what is left is not in the operations the list controls (transcendentals, summation order, fp32 where the CPU uses fp64). - ADR-0214 gate (
cross_backend_parity_gate.py --backends cpu hip, Netflix pair): the 17 runnable cells pass before and after. Themotioncell aborts on a metric name before and after (T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29). - 4K cost, ms per frame,
(t(22) - t(2)) / 20, runs of the two builds interleaved, load average 12 to 26 from other jobs:float_psnr4.49 -> 4.44,motion12.92 -> 13.08,motion_v213.98 -> 13.81,float_motion20.38 -> 19.97,psnr8.62 -> 8.44 (median of 7);float_vif82.94 -> 86.49,adm76.76 -> 72.96,vif141.17 -> 139.15,float_ms_ssim164.61 -> 157.06,float_ssimatscale=193.60 -> 84.76,float_adm77.95 -> 78.75 (median of 5);ciede73.60 -> 74.68,psnr_hvs18.26 -> 18.60,ssimulacra23765 -> 3649,cambi99.49 -> 98.79,speed_chroma5.46 -> 5.50,speed_temporal14.26 -> 14.60,ssim136.9 -> 139.1,float_moment8.76 -> 8.00 (median of 3, not interleaved).float_vifis 4.3% slower in every interleaved sample; no other twin moves outside its run-to-run spread, and none is 10% slower. - Negative: seven twins' outputs change (
float_adm,float_vif,float_motion,float_ssim,float_ms_ssim,ciede,psnr_hvs), so a stored per-build HIP output comparison needs a re-run; no fork snapshot undertestdata/is a HIP output. The gfx1036 dropped a run of commands in at least five of about 150 runs during the measurement (onevif, onefloat_momentand threeadmruns, each with one or two wrong frames, on both builds: in an interleaved repeat 15 of 15admruns with the list and 14 of 15 without were clean); that isT-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01, and those runs were repeated. - Neutral / follow-ups: the CUDA twins still contract in all but two kernels (
T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29, its own policy ADR in review).float_vif_hip's residual against the CPU is the next thing to look at once both backends compile strict.
References¶
- req: RC3 HIP lane brief (2026-10-01): "T-HIP-FP-CONTRACT-DEFAULT-2026-09-29: HIP kernels contract a*b+c into an FMA everywhere except two; turn contraction off through one shared flag list (the SYCL precedent is ADR-1367 ...), pin it with a test, measure parity and ms/frame before/after per twin."
- ADR-1367 — the SYCL policy and the 10% bar.
- ADR-0594 — the per-kernel table this replaces; ADR-0564, ADR-1384 — its other two entries.
- ADR-0214 — the tolerances.
- Research-1407 — the full before / after table.