ADR-1372: CUDA motion differences the frames before the blur, in the kernel motion_v2 already had¶
- Status: Accepted
- Date: 2026-09-30
- Deciders: lusoris
- Tags: cuda, gpu-parity, numerics, motion, fork-local
Context¶
Since the port of Netflix a4a1492d (PR #532) the CPU motion extractor computes its SAD as sum |H(V(prev - cur))|: the frames are differenced first, then the 5-tap Gaussian runs vertically (rounded, >> bpc) and horizontally (rounded, >> 16). motion_cuda kept the algorithm from before the port: integer_motion/motion_score.cu blurred each frame into a uint16 ping-pong and summed |H(V(cur)) - H(V(prev))|. The two are the same sum only without rounding. The SYCL twin with the same order was 2.0e-4 off at 17x17 and 1.3e-5 on the Netflix 576x324 pair until ADR-1371 (T-CUDA-MOTION-BLUR-THEN-DIFF-2026-09-29). motion_v2_cuda already ran the CPU's order: its kernel differences first, and the CPU motion_v2 uses the same arithmetic as the CPU motion.
Two smaller gaps sat in the same extractor. Its debug VMAF_integer_feature_motion_score was the raw normalised SAD, where the CPU emits its SAD score, weighted by motion_fps_weight and capped at motion_max_val. And each batch readback (ADR-0845) synchronised the readback stream twice, once before queueing the copies and once after, although the copies already queue behind every frame's event.
Constraints: integer motion must be bit-exact with the CPU; one behaviour, one implementation (HISS-19); no GPU-CPU round trip and no host wait in the middle of a frame (req); each extractor must be correct on its own, without leaning on the engine's once-per-frame context barrier (ADR-1199), which exists for pictures an external producer fills.
Decision¶
- One diff-first kernel for both CUDA motion twins.
motion_cudaandmotion_v2_cudacallvmaf_cuda_motion_sad_submit()(core/src/feature/cuda/integer_motion_sad_cuda.{h,c}), which loads and launches the kernel ofinteger_motion_v2/motion_v2_score.cu. A block stagesprev - curwith reflect-101 borders in shared memory, filters it vertically and horizontally with the CPU's rounding, and adds|h|into one uint64 accumulator. The vertical sum is int32 for 8-bit input and int64 above, picked by the host.integer_motion/motion_score.cuand its module are deleted. motion_cudakeeps raw frames, not blurred ones. A device ping-pongraw[2]holds the packed luma of the current and previous frame (1 or 2 bytes per pixel instead of the old uint16 blurred planes), filled by a device-to-device copy before the kernel. The first frame only copies.- Ordering on the device, not the host. Each frame's stream waits on the previous frame's event before it overwrites one ping-pong slot and reads the other.
motion_cuda's batch readback queues its copies behind the chained events and synchronises once.submit()never waits on the host. - The debug score is the CPU's.
motion_cudaemitsMIN(sad * motion_fps_weight, motion_max_val)forVMAF_integer_feature_motion_score, asinteger_motion.c::extractdoes.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Reuse the motion_v2 kernel through one host helper (chosen) | Bit-exact at every size; the kernel already existed and matched the CPU motion_v2; one implementation for both twins; no uint16 blurred planes | motion_cuda now reads two raw planes per tile | — |
Fix the order inside motion_score.cu | Smaller diff in the motion TU | Two kernels with the same arithmetic drift apart (HISS-19); SYCL chose one shared kernel for the same reason (ADR-1371) | HISS-19 |
| Keep the per-frame blur and widen the motion gate | No code change | The integer twin stays inexact; the discrepancy is arithmetic, not precision | Nothing to tolerate |
Rely on the engine's per-frame cuCtxSynchronize for the ping-pong hazards | No event wait | The extractor would race the moment the ADR-1199 barrier is narrowed or removed | Device-side ordering costs nothing |
| Drop ADR-0845's eight-frame readback batch for a per-frame readback | Simpler collect path | Changes a measured throughput decision outside this fix | Out of scope; the batch keeps one wait per eight frames |
Consequences¶
- Positive:
motion_cudacomputes the CPUmotionSAD exactly, somotion2/motion3and the debug score equal the CPU's;motion_v2_cuda's SAD is unchanged (same kernel, same arithmetic); its option handling is aligned with the CPU by ADR-1373. One module per motion twin pair instead of two. The batch readback waits once instead of twice. Tile loads of padding threads are clamped into the plane (cuda_tile_index.h), so planes smaller than the 20x20 tile can no longer read before a buffer. - Negative: the
motion_cudakernel reads two planes per tile. On an RTX 4090 (2026-09-30, shared host) the 4K wall time per frame stayed within the noise: 5.1 against 4.8 ms for amasterbuild at (t(200) - t(2)) / 198, median of 5, and 2.6 against 3.1 at (t(22) - t(2)) / 20, median of 3; reading and uploading the frames dominate. The SAD kernels use 28 (8-bit) and 40 (16-bit) registers with no local memory, 6 blocks of 256 threads per SM. - Neutral / follow-ups: verified on an RTX 4090 on 2026-09-30:
integer_motion2/integer_motion3equal the CPU (0.0) on the Netflix pair and on 50 frames of a 3840x2160 clip, where amasterbuild was 1.26e-5 and 6.93e-5 off (T-CUDA-MOTION-BLUR-THEN-DIFF-2026-09-29, closed). The same run found thatmotion_force_zerocrashedmotion_cudaon its first frame, an engine defect fixed alongside (T-GPU-MOTION-FORCE-ZERO-FIRST-FRAME-SEGV-2026-09-30). Guarded bytest_cuda_motion_tiny_frames(==against the scalar CPU, 3x3 to 1283x723, 8, 10 and 16 bits) and the motion cases oftest_cuda_kernel_source_contract.py. The HIP and Metal twins areT-HIP-MOTION-BLUR-THEN-DIFF-2026-09-29andT-METAL-MOTION-BLUR-THEN-DIFF-2026-09-29.
References¶
- req: "there shouldnt be any gpu cpu rountrips" (user, relayed in the RC3 CUDA port brief, 2026-09-30).
- req: RC3 CUDA port brief (2026-09-30): "Implementation matching the CPU reference (integer paths bit-exact ...), no host round trips and no mid-frame host waits (one wait per frame, at collect)".
- Research-1372 — root cause, launch-grid traces, what was and was not verifiable without a device.
- ADR-1371 (the SYCL decision this ports), ADR-0845, ADR-0358, ADR-0219, ADR-1199.