ADR-1419: float_motion_hip stores its differences transposed and adds each row in the CPU's order¶
- Status: Accepted
- Date: 2026-10-01
- Deciders: lusoris
- Tags:
hip,gpu-parity,numerics,float-motion,testing,ci,rc3,fork-local
Context¶
ADR-1409 established that the score of float_motion.c is a property of the order of its additions: the absolute differences of one row are added, left to right, into one float, the row sums top to bottom into a second float, and the total is divided by the pixel count in float. It made float_motion_cuda reproduce that order, ADR-1411 did the same for SYCL, and T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01 stayed open for HIP and Metal.
float_motion_hip summed each 16x16 block on the device and the blocks in double on the host. Measured on a gfx1036 at --precision max against --backend cpu, with the blur already built without FMA contraction (ADR-1407): 3.1e-6 on the Netflix 576x324 pair and 1.36e-4 on the 1920x1080 checkerboard pairs, above the 5e-5 tolerance of ADR-0214; 2.2e-4 with motion_add_scale1. Only the frames whose score is 0 matched.
The HIP twin differs from the other two in two ways. It implements every option of the CPU extractor (ADR-1404): the half-size term of motion_add_scale1 and the chroma planes of motion_add_uv are sums of the same kind. And the device it is measured on is an integrated GPU with two compute units. The kernel ADR-1409 and ADR-1411 use, one thread per row that reads both blurred planes, takes 145 ms per 3840x2160 frame there, where the whole twin took 18: the threads of a wave walk different rows, so every step of a wave reads one cache line per row from memory the GPU shares with the host.
Decision¶
We will make float_motion_hip return the CPU extractor's scores bit for bit by storing every absolute difference first and adding each row afterwards.
Differences in a parallel kernel, stored transposed. The blur kernel already visits every sample; from the second frame on it also stores |cur - prev| of that sample. The differences go into a plane whose rows are grouped in sets of 64 and in which, within a group, the samples of one column are adjacent (vmaf_hip_float_motion_diff_index()). For motion_add_scale1 a second parallel kernel (float_motion_hip_scale1_diff) scales both blurred frames with the CPU's bilinear scaler and stores the differences of the half-size plane the same way.
Row sums in the CPU's order. float_motion_hip_row_sum runs one thread per row, 64 threads per block, one block per group. A thread adds the width differences of its row, left to right, into one fp32 accumulator. At every step the threads of a block read 64 consecutive floats, so the kernel streams through memory instead of gathering. The same kernel adds the scale-0 and the scale-1 differences of every plane.
Rows on the host. collect() builds each plane's score with vmaf_hip_float_motion_plane_score(): the fp32 mean of vmaf_float_motion_score_from_row_sads() (ADR-1409's helper), plus the fp32 scale-1 mean, and adds the planes in double, as motion.c::vmaf_image_sad_c() and float_motion.c::motion_score_pair() do.
One header. The difference, the layout index, the row sum, the bilinear sample and the plane score live in core/src/feature/hip/float_motion/float_motion_rows.h, plain C and HIP C++. The kernels and the extractor compile it, and a device-free test compiles the same lines against motion.c::compute_motion().
Gate. float_motion lists hip next to cuda and sycl in EXACT_TWINS (scripts/ci/cross_backend_calibration.py).
Alternatives considered¶
Times are per 3840x2160 frame of --backend hip --feature float_motion on the gfx1036, steady state inside the process, with other jobs loading the host. The twin took 18 ms before.
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Parallel differences stored transposed, then one thread per row (this ADR) | The CPU's bits with every option; 20 ms; consecutive loads in the serial kernel | One more device plane per picture plane (width x height floats, rounded up to 64 rows), and a second one with motion_add_scale1 | Chosen |
| One thread per row reading both blurred planes (the CUDA and SYCL kernel) | No extra plane; the kernel of ADR-1409 unchanged | 145 ms per frame, and 229 ms with motion_add_scale1, where the scaler also runs inside the serial loop | Eight times slower than the twin was: the lanes of a wave read 32 or 64 different cache lines at every step |
| Differences stored in picture order, one thread per row | No transposed index | Half the loads of the option above, still one cache line per lane and step | The gathers are the cost, not the subtraction |
| Read the differences back and add them on the host | Trivially exact; the host's cache handles a row sum well | 33 MB of readback per plane and frame at 3840x2160, and the sum runs on the thread that drives every extractor | The rows are independent, so the device can add them once it reads them in order |
| Keep the per-block sum, keep the tolerance | No code change | 1.36e-4 at 1080p is above the tolerance; any tolerance accepts a twin defect below it | Rejected by ADR-1409 |
Consequences¶
- Positive:
float_motion_hipequals--backend cpuat--precision maxon every value ofmotion,motion2andmotion3measured on a gfx1036: the Netflix 576x324 pair at 8 bits (48 frames) and 10 bits (3), both 1920x1080 checkerboard pairs (3 each) and BBB 3840x2160 (20 frames), 231 values per option set, for the default options,motion_add_scale1,motion_add_uv, both together,motion_filter_size3 and 1, and the fps weight, blend and cap options: 1617 of 1617 values, where 237 of 1617 matched before (the zero frames, and the capped scores of the last set). The gate cell is an equality, so a twin defect of any size fails it. - Negative: the stored differences and the second kernel cost time on the gfx1036. Medians of seven interleaved runs of both builds: 2.50 to 2.78 ms per 1920x1080 frame and 18.09 to 19.76 per 3840x2160 frame with the default options, 3.03 to 3.74 and 21.02 to 24.68 with
motion_add_scale1, 19.05 to 20.52 at 3840x2160 withmotion_add_uv.--model version=vmaf_float_v0.6.1reads 49.98 and 50.11 ms at 1920x1080 and 192.66 and 199.17 at 3840x2160 (five runs each), inside the spread of the runs. The twin holds one more float plane per picture plane on the device, two withmotion_add_scale1. It now reproduces a value that is further from the exact sum than the one it returned before; that is the CPU extractor's value and the one the models were trained on. - Neutral / follow-ups:
float_motion_metalkeeps its per-block sum (T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01stays open for it). Any storedfloat_motion_hipoutput changes in its low digits. A discrete AMD GPU is unmeasured; it has the lanes to run the gathering kernel fast, and the transposed layout costs it nothing. Guards:test_hip_float_motion_rows(the header againstcompute_motion(), no device, with and withoutmotion_add_scale1, odd sizes),test_hip_float_motion_parityand its 960x540 variant (==onmotion,motion2andmotion3of four frames, with the default options and withmotion_add_scale1+motion_add_uv; 10 of 12 scores differ on the old twin) andtest_hip_kernel_source_contract.py(a wave reduction, an fp64 row sum, an untransposed read, an fp64 scale sum and a host sum of the twin's own are each rejected).
References¶
req(maintainer brief for the HIP lane, 2026-10-01): "T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM, HIP part: reproduce the CPU's per-row fp32 running sums like PR #1699 does for CUDA (core/src/feature/float_motion_sad.h)."req(maintainer brief for the HIP lane, 2026-10-01): "The gfx1036 is slow: correctness evidence first, timings second."- ADR-1409 (the decision this applies to HIP, and the host helper), ADR-1411 (SYCL), ADR-1397 (the exact cell), ADR-1407 (contraction off), ADR-1404 (the twin's options), ADR-0214.
docs/state.md:T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01(HIP part closed by this decision).