ADR-1404: float_motion_hip emits motion3 and implements every CPU float_motion option on the device¶
- Status: Accepted
- Date: 2026-10-01
- Deciders: lusoris
- Tags: hip, gpu-parity, motion, feature-extractor, fork-local
Context¶
The CPU float_motion extractor emits VMAF_feature_motion3_score for every frame and takes nine options. float_motion_hip emitted motion and motion2 only and declared four of the options (debug, motion_force_zero, motion_fps_weight, motion_max_val). Scores stayed correct, because the ADR-1183 gate keeps a feature with an undeclared option on the CPU, but --backend hip --feature float_motion dropped motion3 from the output without a warning, a model reading motion3 or setting motion_blend_factor, motion_blend_offset, motion_filter_size, motion_add_scale1 or motion_add_uv never ran on the device, and naming the twin with one of them failed with unknown option (T-HIP-FLOAT-MOTION-MOTION3-OPTIONS-2026-09-30, the HIP part of T-GPU-FLOAT-MOTION3-MISSING-2026-09-30).
motion3 and the two blend options are host arithmetic on the SAD the twin already reduces; the CUDA twin got them that way in #1637. The other three options change what the device computes: the blur filter, a second SAD over both blurred frames scaled to half size, and the same chain on the U and V planes. The state row allows either kernel work for these or an explicit -ENOTSUP, the posture motion_hip takes for motion_add_uv (ADR-0989).
Decision¶
float_motion_hip takes the CPU option table unchanged (names, aliases, types, defaults, ranges, order) and implements all of it:
motion3on the host, asfloat_motion.cforms it:motion_blend()from the sharedmotion_blend_tools.hof the fps-weighted score, then themotion_max_valcap; index 0 from the first SAD alone, then the blendedmotion2, the tail fromflush(), 0 for a one-frame run and undermotion_force_zero.motion_filter_sizeas a kernel argument. The blur selectsFILTER_5_s,FILTER_3_s(3) orFILTER_5_NO_OP_s(1) asmotion_blur_plane()does. The 3-tap and no-op filters run as five taps with zero outer weights, which adds exact zeros, so the tile, its halo and the launch geometry are the same for every filter. The minimum frame is the CPU's: 2x2 for the 3-tap filter, 3x3 otherwise.motion_add_scale1as a second kernel.float_motion_hip_scale1_sadscales both blurred frames with the bilinear scaler ofmotion.c(motion_scale_bilinear(), contraction off) and reduces|cur - prev|per block; the host adds the scale-1 mean to the scale-0 mean.motion_add_uvas the same kernels per plane. The state holds one staging buffer and one blur ping-pong per plane; a frame uploads the planes it needs in one call and launches the blur (and scale-1) kernel once per plane with the plane's geometry (ceiling chroma extents fromvmaf_chroma_extent()); all partials come back in one read-back. 4:0:0 input is refused at init, as on the CPU.
The two 8-bit / 16-bit kernel bodies become one template, so the tile load, the blur and the block reduction exist once.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
Host motion3 and blend options; filter, scale 1 and chroma on the device (chosen) | Every CPU float_motion request runs on the device; one option table; still one upload call, one read-back and one wait per frame | A second kernel and per-plane state; chroma triples the uploads when asked for | — |
Host motion3 and blend options only; -ENOTSUP for motion_add_scale1 / motion_add_uv, default-only motion_filter_size (ADR-1316) | Small change, the CUDA twin's current scope | Three options still pin the feature to the CPU; a direct request for the twin fails at init | The kernel work is small: 26 ms a frame at 4K with both options against 75 ms on 16 CPU threads |
| Compute scale 1 or chroma on the host from read-back blurred planes | No new kernel | A full-resolution float plane per frame over the bus, a host pass per frame | Violates device residency |
Round each plane's mean to fp32 as vmaf_image_sad_c() does | Mirrors the CPU's types | The CPU's fp32 running sum cannot be reproduced in parallel, so the rounding would not make the twin exact; the double sums are closer to the true value | No gain in parity |
| Separate 3-tap kernel with a 1-pixel halo | Fewer loads for motion_filter_size=3 | A second tile geometry and entry-point pair for a rarely used option | The zero-weight taps cost nothing measurable |
Consequences¶
- Positive: on a gfx1036 (
ryzen-4090-arc, ROCm 7.2.4),--precision max, against--backend cpu:motion3within 2.78e-6 on the Netflix 576x324 pair and 5.71e-6 on BBB 3840x2160, the same asmotion2;motion_blend_factor=0.5:motion_blend_offset=21.39e-6;motion_filter_size=35.65e-6 and=14.72e-7;motion_add_scale13.17e-6;motion_add_uv2.73e-6; all options together 9.84e-6 (Netflix) and 8.45e-6 (4K);motion_force_zeroand a one-frame run identical.motionandmotion2with the previous options are bit-identical to the build before the change on every fixture.--backend hip --feature float_motion=motion_add_uv=truenow listsfloat_motion_hipinfeature_backends; before it warned and ran the CPU extractor. - Negative:
float_motion_hipcarries three plane states, of which the default configuration uses one. The 1080p checkerboard pairs stay at 1.35e-4 from the CPU on a score of 18.86, as before: the CPU adds 2 million fp32 terms in one running sum. - Neutral / follow-ups: the SYCL and Metal parts of
T-GPU-FLOAT-MOTION3-MISSING-2026-09-30stay open, and so domotion_filter_size,motion_add_scale1andmotion_add_uvon the CUDA, SYCL and Metalfloat_motiontwins. Guarded bytest_hip_twin_option_parity(device: every option against the CPU, the 2x2 boundary, the 4:0:0 and below-minimum refusals; device-free: the option table), six planted regressions intest_hip_kernel_source_contract.py, andtest_vmaf_feature_backend_hip(the CLI runs the twin formotion_filter_size=3).
References¶
- req: RC3 HIP lane brief (2026-10-01): "T-HIP-FLOAT-MOTION-MOTION3-OPTIONS-2026-09-30: float_motion_hip emits no motion3 score and lacks five CPU float_motion options."
- ADR-1382 — the first four options and
motion_clip(). - ADR-1183, ADR-1359 — why an undeclared option kept the feature on the CPU.
- ADR-0219, ADR-0989 — host-side
motion3andmotion_add_uvon the integer twins. - Research-1372 — the CUDA twin's host-side
motion3.