Research-1403: FMA contraction off for every CUDA kernel — which kernels change, what each twin does against the CPU, what it costs¶
Question¶
T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29: nvcc fuses a * b + c into one FMA by default, and only six of the 21 CUDA fatbins were built with --fmad=false. The CPU reference and, since ADR-1367, every SYCL twin do not contract. If every CUDA kernel is built without contraction: which kernels' code changes, which twins move against --backend cpu, does any twin rely on a fused operation, and what does a 3840x2160 frame cost before and after?
Sources¶
core/src/meson.build(cuda_cu_sources, the fatbincustom_target),core/src/feature/cuda/**/*.cu, the CPU extractors they mirror.- RTX 4090 (sm_89, driver 615.71.09), nvcc 13.4.92, gcc 16.2.1, glibc 2.44, Ryzen 9 9950X3D.
meson setup <dir> core -Denable_cuda=true -Denable_sycl=false --buildtype=release -Db_lto=false. Before:masterbd061d95a. After: the same plus this change. (Two earlier rounds measured the same numbers for every twin they did not concern: on2c3acf1c9, before ADR-1397 gavepsnr_hvs_scorethe flag and its CPU-order sum, and onbd17c352f, before ADR-1399 rewrotessim_score.cuwith explicit rounding intrinsics. Thefloat_ms_ssimsteps in finding 3 and the fp64 cost are from the first round, the timings in finding 5 from the second.) - Fixtures, 8-bit 4:2:0,
--precision max: the Netflix pairsrc01_hrc00/01_576x324(48 frames), the checkerboard pairscheckerboard_1920_1080_10_3_0_0against_1_0and_10_0(3 frames each), and the first 50 frames of BBB 3840x2160 (testdata/bbb/ref/dis_3840x2160_200f.yuv). - Every twin is requested by its registered name (
--backend cuda --feature <name>_cuda), the CPU extractor with--backend cpu --threads 16.
Findings¶
1. Fifteen fatbins keep their machine code, six change¶
cuobjdump -sass of every fatbin, before against after, for each shipped architecture (sm_80, sm_86, sm_89, sm_90, sm_100, sm_120); FMA counts are the sm_89 cubin's.
| Fatbin | Twin | Machine code | fp32 FMA | fp64 FMA |
|---|---|---|---|---|
speed_score, ssimulacra2_blur, ssimulacra2_device, float_adm_score, integer_ssim_score, psnr_hvs_score | speed_*, ssimulacra2, float_adm, ssim, psnr_hvs | already --fmad=false; fatbins byte-identical | unchanged | unchanged |
adm_dwt2, adm_cm, adm_csf, adm_csf_den | adm | identical on all six | 0 / 746 / 52 / 0 | 0 |
cambi_score, moment_score, motion_v2_score, psnr_score | cambi, float_moment, motion, motion_v2, psnr | identical on all six | 0 | 0 |
ssim_score | float_ssim | identical on all six: every rounding is an explicit intrinsic (ADR-1399) | 110 | 48 |
filter1d | vif | changes | 370 -> 370 | 130 -> 120 |
float_psnr_score | float_psnr | changes | 10 -> 9 | 0 |
float_motion_score | float_motion | changes | 77 -> 9 | 0 |
float_vif_score | float_vif | changes | 330 -> 123 | 0 |
ciede_score | ciede | changes | 938 -> 834 | 0 |
ms_ssim_score | float_ms_ssim | changes | 131 -> 8 (flag only) | 34 -> 31 (flag only) |
For the nine identical ones the only difference in the fatbin is the PTX text: mul.f32 / mul.f64 become mul.rn.f32 / mul.rn.f64 and the header gains ptxasOptions = --fmad false. nvcc had formed no FMA from their source: the FMAs counted in adm_cm, adm_csf and ssim_score are in the PTX of neither build, they are the back end's own expansion of division and square root. The FMAs that remain in the changed kernels after the flag are of that kind (division, square root, libdevice math functions), plus the explicit ones.
2. Per-twin agreement with the CPU, before and after¶
Identical frame outputs / all frame outputs (every output of the twin, every frame), then the largest and the mean absolute difference over them. "Same" means the twin's own value is unchanged on every frame.
| Twin | Fixture | Identical before | Identical after | Max before | Max after | Mean before | Mean after |
|---|---|---|---|---|---|---|---|
vif, motion, motion_v2, psnr, psnr_hvs, float_psnr, float_moment, float_ssim, cambi, speed_temporal | all four | all | all | 0 | 0 | 0 | 0 |
float_ms_ssim | Netflix | 0/48 | 48/48 | 6.89e-8 | 0 | 2.45e-8 | 0 |
| checkerboard 1 px | 0/3 | 3/3 | 2.14e-6 | 0 | 1.47e-6 | 0 | |
| checkerboard 10 px | 0/3 | 3/3 | 4.42e-6 | 0 | 3.76e-6 | 0 | |
| BBB 4K | 0/50 | 50/50 | 5.82e-7 | 0 | 2.49e-7 | 0 | |
ciede | Netflix | 0/48 | 0/48 | 1.14e-5 | 1.14e-5 | 1.02e-5 | 1.03e-5 |
| checkerboards | 0/6 | 0/6 | 8.55e-7 | same | 8.55e-7 | same | |
| BBB 4K | 0/50 | 0/50 | 1.40e-6 | 1.49e-6 | 1.17e-6 | 7.70e-7 | |
float_motion | Netflix | 2/144 | 2/144 | 3.01e-6 | 3.12e-6 | 9.11e-7 | 9.18e-7 |
| checkerboards | 4/18 | 4/18 | 1.35e-4 | 1.36e-4 | 9.69e-5 | 9.76e-5 | |
| BBB 4K | 2/150 | 2/150 | 2.37e-5 | 2.36e-5 | 8.66e-7 | 8.64e-7 | |
float_vif | Netflix | 0/192 | 0/192 | 2.72e-5 | 3.81e-5 | 3.07e-6 | 3.18e-6 |
| checkerboard 1 px | 0/12 | 0/12 | 1.02e-6 | 1.04e-6 | 4.44e-7 | 4.72e-7 | |
| checkerboard 10 px | 10/12 | 10/12 | 1.12e-12 | 1.14e-12 | 1.04e-13 | 1.05e-13 | |
| BBB 4K | 0/200 | 0/200 | 5.75e-6 | 7.03e-6 | 1.94e-6 | 1.83e-6 | |
adm | all four | 176/728 | same | 2.09e-7 | same | 2.8e-8 | same |
float_adm | all four | 138/728 | same | 1.28e-5 | same | 7.2e-8 | same |
ssim (integer) | all four | 0/104 | same | 1.06e-11 | same | 5.1e-13 | same |
ssimulacra2 | all four | 8/104 | same | 7.28e-11 | same | 2.0e-12 | same |
speed_chroma | all four | 306/312 | same | 1.43e-6 | same | 1.9e-8 | same |
float_ssim_cuda runs at every scale since ADR-1399 and was bit-identical before this change; on the previous base (bd17c352f, scale 1 only, plain fp32 window sums) the flag moved it from 1.79e-7 to 1.19e-7 on the Netflix pair. speed_chroma's six differing outputs are the CPU's glibc log2f (ADR-1380), not the twin.
Against the ADR-0214 tolerances (FEATURE_TOLERANCE in scripts/ci/cross_backend_parity_gate.py; the gate runs the Netflix pair): every twin is inside its tolerance on the Netflix pair before and after. Run as the gate itself (--backends cpu cuda): all 18 cells pass before and after (adm, cambi, ciede, float_adm, float_moment, float_motion, float_ms_ssim, float_ms_ssim_lcs, float_psnr, float_ssim, float_ssim_lcs, float_vif, motion, motion_v2, psnr, psnr_hvs, ssimulacra2, vif). One is outside it on larger frames, before and after alike, for a reason the flag does not touch: float_motion on the checkerboards (1.36e-4 against 5e-5; finding 4).
3. float_ms_ssim_cuda relied on contraction in two ways; with the reference's arithmetic it is bit-identical¶
With the flag alone the twin gets closer on small frames and further at 4K. The steps, each on top of the previous (max / mean absolute difference):
| Step | Netflix | Checkerboard 1 px | Checkerboard 10 px | BBB 4K |
|---|---|---|---|---|
| before | 6.89e-8 / 2.45e-8 | 2.14e-6 / 1.47e-6 | 4.42e-6 / 3.76e-6 | 5.82e-7 / 2.49e-7 |
--fmad=false only | 5.53e-8 / 1.97e-8 | 1.05e-6 / 5.38e-7 | 2.97e-6 / 2.26e-6 | 1.23e-6 / 7.57e-7 |
+ decimate taps as __fmaf_rn() | 6.80e-8 / 2.26e-8 | 1.57e-6 / 8.49e-7 | 3.16e-6 / 2.47e-6 | 1.13e-6 / 7.24e-7 |
| + window sums in fp64 | 2.73e-8 / 8.28e-9 | 3.19e-8 / 1.65e-8 | 3.45e-8 / 2.25e-8 | 3.49e-8 / 1.07e-8 |
| + reference operand types, fp32 means | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 0 |
| window sums as an fp32 pair instead of fp64 | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 0 |
What differed from the CPU, in the order the steps remove it:
ms_ssim_decimate.caccumulates each low-pass tap with one fused multiply-add (vmaf_fmaf_exact(),fmaddin the SIMD twins). The kernel wroteacc += sample * tapand matched only because nvcc fused it.iqa_convolve()forms each Gaussian-window product in fp32 and adds the products in fp64, then rounds to fp32 once per pass. The kernel kept an fp32 running sum. Fusing each step, which keeps the product exact, hid most of the difference; without fusion it shows, most at 4K.ssim_accumulate_default_scalar()divides an fp64 numerator by an fp32 denominator forlandc, and computessas an fp32 quotient; its constantsC1,C2,C3are fp32. The kernel promoted all of it to fp64.iqa_ssim()returns each per-scale mean as a float andms_ssim.craises the floats to the Wang exponents. The host kept fp64 means, so the 15enable_lcsoutputs were up to 1.3e-6 off as well.
The fp64 window sums cost 3.4 ms per 3840x2160 frame: 10.73 -> 14.09 ms (7 alternating repetitions of 100 frames; 10.49 -> 14.44 with 5 repetitions of 20 frames). An fp32 pair (running sum plus the exact error of each addition, six fp32 operations per term) carries the 11-term sum to about 48 bits, gives the same 104 frame scores, and costs nothing measurable (10.54 -> 10.57, 7 repetitions of 100 frames). It is what ships. It is not the reference's operation literally: the pair and the 53-bit sum can round to different fp32 values when the exact sum lies within about 2^-18 ulp of a rounding boundary, which then moves one pixel's statistic by one ulp. The fp64 l / c / s block sums differ from the CPU's raster order by more than that already; the fp32 rounding of the mean absorbs both.
4. What the remaining differences are¶
float_motion: the CPU keeps one fp32 running sum of the absolute differences of a whole plane (float_sad_line_c()per row,compute_motion_simd()over the rows) and divides in fp32. The kernel sums per 16x16 block and the host adds the blocks in fp64. Replayed on the host for the 1920x1080 checkerboard: blur and differences in fp32 with separate multiply and add, then the fp32 running sum, reproduces the CPU's score bit for bit (18.805618286132812 and 18.8588924407959 for frames 1 and 2); the fp64 sum of the same differences is 1.17e-4 and 1.35e-4 lower, and the twin is within 6e-7 of that. So the blur matches and the whole gap is the CPU's own rounding.T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01.float_vif: not isolated. The kernel calls the devicelog2fwhere the CPU calls glibc's on an fp64-promoted argument, and it reduces per block. The SYCL twin shows the same two numbers on the Netflix pair before and after ADR-1367 (2.71e-5 and 3.81e-5), so contraction-free CUDA and SYCL now agree with each other.T-CUDA-FLOAT-VIF-RESIDUAL-2026-10-01.ciede: per-pixelpow,sqrt,sin,atan2on the device against libm; tolerance 5e-3.
5. Cost¶
3840x2160, BBB, one twin per run, (t(N) - t(2)) / (N - 2) ms per frame, alternating pairs with the order swapped every pair; median before and after, then the median and the quartiles of the paired difference. The host is shared; twins whose machine code is identical show what it resolves.
15 pairs, N = 102, load average 9 to 16:
| Twin | Code | Before | After | Paired difference | After slower in |
|---|---|---|---|---|---|
vif | changed | 3.42 | 3.44 | -0.07 (-0.30 .. +0.21) | 5/15 |
float_psnr | changed | 2.62 | 2.61 | -0.07 (-1.03 .. +0.13) | 4/15 |
float_motion | changed | 2.73 | 3.03 | +0.21 (-0.24 .. +0.44) | 10/15 |
float_vif | changed | 2.53 | 2.65 | +0.12 (-0.02 .. +0.22) | 11/15 |
ciede | changed | 2.42 | 2.46 | +0.04 (-0.05 .. +0.59) | 11/15 |
float_ms_ssim | changed | 11.26 | 10.87 | -0.12 (-1.34 .. +0.62) | 6/15 |
psnr_hvs | identical | 12.18 | 12.09 | -0.08 (-0.16 .. +0.10) | 7/15 |
psnr | identical | 2.84 | 3.12 | +0.07 (-0.12 .. +0.30) | 11/15 |
adm | identical | 3.50 | 3.47 | +0.04 (-0.11 .. +0.14) | 8/15 |
speed_chroma | identical | 2.38 | 2.44 | +0.04 (-0.08 .. +0.16) | 8/15 |
The two largest, repeated with 21 pairs and N = 200, load average 9:
| Twin | Code | Before | After | Paired difference | After slower in |
|---|---|---|---|---|---|
float_motion | changed | 2.82 | 2.89 | +0.10 (-0.22 .. +0.27) | 14/21 |
float_vif | changed | 2.43 | 2.56 | 0.00 (-0.04 .. +0.05) | 12/21 |
psnr | identical | 2.61 | 2.38 | -0.06 (-0.35 .. +0.07) | 8/21 |
float_moment | identical | 2.53 | 2.45 | -0.03 (-0.11 .. +0.24) | 10/21 |
No paired difference is distinguishable from the controls; the other thirteen twins run the same code as before. At 4K these twins spend their time reading and uploading the frame, not in the kernels (what the flag adds to float_motion, about 70 separate multiplies and adds per pixel, is a few microseconds per 4K frame on this GPU). A kernel-only measurement needs a quiet host or an in-process harness; vmaf_bench --gpu-profile is SYCL-only.
6. The clang CUDA path¶
-Denable_nvcc=false did not configure on master: the shared fatbin command concatenates nvcc_ccbin_flags and nvcc_host_includes, which only the nvcc branch assigned. With both assigned (empty) and -ffp-contract=off, clang 22.1.8 builds all 21 kernels against CUDA 13.4 (it warns that the SDK is newer than it knows). Its PTX has mul.rn.f32 / add.rn.f32 throughout and fma.rn.f32 only where the source asks for it.
Run on the RTX 4090 against the nvcc build of the same tree, on the four fixtures (BBB: 22 frames): 18 of the 19 twins give the nvcc build's value on every output of every frame (3 952 outputs), among them the ones that are not bit-identical to the CPU (adm, float_adm, float_vif, float_motion, ssim, ssimulacra2, speed_chroma); float_ssim and float_ms_ssim are bit-identical to the CPU under both compilers. The one that differs calls device math functions: ciede (6 of 76 outputs equal, 1.13e-5 from the CPU against nvcc's 1.14e-5). The first clang build of float_ms_ssim_cuda matched the CPU on 9 of 48 frames (3.4e-8): clang compiles sqrtf() to sqrt.approx.f32 where nvcc emits sqrt.rn.f32. With __fsqrt_rn() the clang build is bit-identical to the CPU on all 104 frames as well, and the nvcc machine code is unchanged.
Reproduce¶
meson setup build-cuda core -Denable_cuda=true -Denable_sycl=false --buildtype=release -Db_lto=false
ninja -C build-cuda
# which fatbins hold FMAs
for f in build-cuda/src/*.fatbin; do printf '%s %s\n' "$f" "$(cuobjdump -sass -arch sm_89 "$f" | grep -c FFMA)"; done
# one twin against the CPU
python3 scripts/dev/speed_gpu_parity.py --backend cuda --vmaf "$PWD/build-cuda/tools/vmaf" \
--netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb --feature float_ms_ssim
build-cuda/test/test_cuda_float_ms_ssim_parity
python3 -m pytest core/test/test_strict_fp_compiler_args.py core/test/test_cuda_kernel_source_contract.py