ADR-1411: float_motion_sycl adds its SAD in the CPU's order and returns the CPU's scores bit for bit¶
- Status: Accepted
- Date: 2026-10-01
- Deciders: lusoris
- Tags:
sycl,gpu-parity,numerics,float-motion,testing,ci,rc3,fork-local
Context¶
ADR-1409 established that the score of float_motion.c is a property of the order of its additions: compute_motion_simd() adds the absolute differences of one row, left to right, into one float, adds the row sums top to bottom into a second float, and divides by the pixel count in float. It made float_motion_cuda reproduce that order and left the SYCL, HIP and Metal twins as T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01.
float_motion_sycl reduced the differences of each 32x4 work-group on the device (a sub-group reduction, then the sub-groups of the group) and added the groups in double on the host. Measured on an Arc A380 at --precision max against --backend cpu, with the blur already built without FMA contraction (ADR-1367): 3.1e-6 on the Netflix 576x324 pair, 2.4e-5 on BBB 3840x2160 and 1.36e-4 on both 1920x1080 checkerboard pairs, above the 5e-5 tolerance of ADR-0214. Those are the numbers the CUDA twin had with the same reduction: the blur matched, the reduction did not.
A SYCL kernel has two constraints the CUDA kernel does not: it may not use the fp64 type (ADR-0220), and on Arc A-series GPUs under the xe driver it may not use scratch memory (ADR-1395). A sequential fp32 sum needs neither.
Decision¶
We will make float_motion_sycl return the CPU extractor's scores bit for bit, by the design of ADR-1409.
Row sums on the device. launch_float_motion_row_sad() (core/src/feature/sycl/float_motion_sycl.cpp) runs one work-item per row. fm_row_sad() adds |cur[j] - prev[j]| for j = 0 .. w - 1 into one fp32 accumulator and stores it; the frame's readback is h floats. The blur kernel writes the blurred plane only. The TU contains no group, sub-group or atomic reduction.
Sub-group size 8. The kernel requests it. The lanes of a hardware thread are rows, each loop step is two gathers, and the pass is bound by reading the two blurred planes again; narrow sub-groups put more hardware threads on it.
Rows on the host. collect() calls vmaf_float_motion_score_from_row_sads() (core/src/feature/float_motion_sad.h), the helper ADR-1409 introduced. The twin keeps no sum of its own.
Gate. float_motion lists sycl next to cuda in EXACT_TWINS (scripts/ci/cross_backend_calibration.py): the CPU, CUDA and SYCL cells are compared with tolerance 0 at --precision max. HIP and Metal keep places=4.
Alternatives considered¶
Kernel times are per 3840x2160 frame on an Arc A380: the frame time of the vmaf tool minus the 3.51 ms it takes with the row kernel left out. The old combined blur-and-reduce kernel spent 0.34 ms on its SAD.
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| One work-item per row, sub-group size 8 (this ADR) | The CPU's bits on every frame; the simplest kernel; no dependence on how work-items map to lanes; 0.70 ms | 0.36 ms more than the old reduction, because both blurred planes are read a second time | Chosen |
| The same kernel at the compiler's sub-group size (32), or 16 | No attribute | 1.12 ms and 0.82 ms: 68 or 135 hardware threads for 2160 rows | Slower, same result |
One sub-group of 16 consecutive pixels per row, summed lane by lane with select_from_group() | Contiguous loads, one thread per row | 1.10 ms: the shuffle per pixel costs more than the gathers it saves; correctness depends on the sub-group being exactly the work-group | Slower and more fragile |
| One work-group per row, 16-wide vector loads, a chain that is uniform across the lanes | Contiguous loads, scalar adds | 0.80 ms; seven of eight lanes compute a value that is thrown away | Not faster than the simple kernel |
| Keep the per-group sum, raise the tolerance | No code change | Accepts any twin error below the CPU's own rounding error, which grows with the frame area | Rejected by ADR-1409 |
| Read all differences back and add them on the host | Trivially exact | 33 MB of readback per 3840x2160 frame | The rows are independent, so the device can add them |
| Add the rows inside the blur kernel | No second pass over the planes | The sum of a row is sequential across the 32-pixel work-groups of the blur; no kernel that blurs in parallel can produce it | Not possible with a tiled blur |
Consequences¶
- Positive:
float_motion_syclequals--backend cpuat--precision maxon every frame ofmotionandmotion2measured on an Arc A380: the Netflix 576x324 pair (48 frames; before 3.1e-6), both 1920x1080 checkerboard pairs (3 frames each; before 1.36e-4), BBB 3840x2160 (200 frames; before 2.4e-5), the Netflix pair at 10, 12 and 16 bits and as 4:2:2 10-bit, and the Netflix pair withmotion_fps_weightandmotion_max_valset. The same holds against a GCC build of the CPU extractor and against the CPU extractor of the icx build itself. The gate cell is an equality, so a twin defect of any size fails it. - Negative: 0.38 ms more per 3840x2160 frame on the Arc A380 through the
vmaftool: 3.85 ms before, 4.23 after (medians of 25 paired 200-frame runs, host load 7 to 11; the untouchedfloat_psnr_syclread 3.37 and 3.38 in the same session). On the Netflix 576x324 pair: 0.15 ms before, 0.22 after (15 paired 48-frame runs). The twin now reproduces a value that is further from the exact sum than the one it returned before; that is the CPU extractor's value and the one the models were trained on. - Neutral / follow-ups:
float_motion_hipandfloat_motion_metalkeep their per-block sums (T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01stays open for them). The twin still provides nomotion3(T-GPU-FLOAT-MOTION3-MISSING-2026-09-30). Any storedfloat_motion_sycloutput changes in its low digits (by at most the differences above). The row kernel uses no scratch memory (test_sycl_kernel_scratch, 110 kernels audited on the A380, the ratchet list unchanged). Guards:test_sycl_float_motion_parityand its 960x540 variant compare every frame with==at 8, 10 and 12 bits (they fail on the old twin by 2.0e-5 and 6.1e-5), andtest_sycl_kernel_source_contract.pyrejects a group reduction, a strided row loop, a blocked launch, a host sum of the twin's own and an fp64 row total.
References¶
req(maintainer brief for the SYCL exactness lane, 2026-10-01): "results before speed; a twin reproduces the CPU bit for bit, tuning comes afterwards".req(maintainer, popup answer recorded in ADR-1397, 2026-10-01): "2, we tune afterwards, thats fixing now, what is the speed worth if the results are wrong".- ADR-1409 (the decision this applies to SYCL, and the host helper), ADR-1397 (the exact cell), ADR-1367 (contraction off), ADR-1395, ADR-0220, ADR-0214, ADR-1365 (the twin's options).
docs/state.md:T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01(SYCL part closed by this decision).