Skip to content

ADR-1237: CAMBI Anti-Dithering AVX2 Vectorization, SpEED SIMD QR Dispatch, and Threaded GPU Flush Alignment

  • Status: Accepted
  • Date: 2026-09-08
  • Deciders: Kilian, Claude Opus 5
  • Tags: perf, simd, x86, hip, cambi, speed

Context

Issue #1245 initiated a comprehensive benchmark and performance tuning pass across all four backends (CPU, CUDA, SYCL, and HIP) on workstation silicon (AMD Ryzen 9 9950X3D 32-thread CPU, NVIDIA RTX 4090, Intel Arc A380, and AMD Radeon Graphics gfx1036).

Profiling and harness validation surfaced several hot spots and pipeline bottlenecks:

  1. CAMBI anti-dithering filter: in core/src/feature/cambi.c, anti_dithering_filter() applies a 2×2 box filter on 16-bit luma to suppress spatial dithering for all 8-bit inputs (enc_bitdepth < 10). This loop was purely scalar, processing pixels sequentially despite CAMBI's other stages having AVX2 paths.
  2. SpEED host-side QR multiply: core/src/feature/speed_internal.c defines si_mat_mul() for small-matrix QR factorisation in GPU SpEED twins. Per ADR-1196, si_mat_mul was left scalar pending resolution of GPU motion flush ordering (T-GPU-MOTION-FLUSH-DOUBLE-EMIT-2026-09-06). With the flush mechanism stabilized, the scalar multiplication loop was ready for bit-exact SIMD dispatch.
  3. Threaded GPU flush & extractor pooling: in core/src/libvmaf.c, batch_extractor_skip() and read_pictures_should_skip() only excluded CUDA and SYCL from CPU thread pooling. When --threads N was passed with HIP, HIP feature extractors were erroneously scheduled into worker threads via .extract (which GPU extractors do not implement), returning -EINVAL (-22) and leaving gpu_pending uncollected.
  4. Harness robustness (testdata/bench_all.sh): the candidate search order placed /opt/intel/oneapi-2025.3/setvars.sh ahead of /opt/intel/oneapi/setvars.sh (2026.0), causing symbol resolution failures when binaries were linked against 2026.0. In addition, Test 2 targeted a non-existent 1080p fixture name rather than the canonical checkerboard files in python/test/resource/yuv/.

Decision

We make the following changes:

  1. Vectorize anti_dithering_filter with AVX2: implement anti_dithering_filter_avx2() in core/src/feature/x86/cambi_avx2.c and dispatch to it in cambi.c when VMAF_X86_CPU_FLAG_AVX2 is active. The kernel processes 16 uint16_t lanes per iteration using zero-extended 32-bit addition, arithmetic shift, and _mm256_packus_epi32 + lane permute, guaranteeing bit-exact parity with scalar arithmetic.
  2. Wire si_mat_mul to SIMD matmul: dispatch si_mat_mul in core/src/feature/speed_internal.c through speed_matmul_avx512 / speed_matmul_avx2 when CPU flags match, falling back to speed_matmul_scalar. These kernels share the exact same accumulation order and compile under -ffp-contract=off to ensure bit-exact arithmetic.
  3. Align GPU extractor pooling and threaded flush: update not_pooled in batch_extractor_skip() and gpu in read_pictures_should_skip() to include VMAF_FEATURE_EXTRACTOR_HIP and VMAF_FEATURE_EXTRACTOR_METAL. In flush_context_threaded(), drain gpu_pending for non-CUDA/SYCL extractors before flushing, mirroring flush_context_serial().
  4. Fix testdata/bench_all.sh: prefer /opt/intel/oneapi/setvars.sh before legacy versions, and update Test 2 to use checkerboard_1920_1080_10_3_0_0.yuv and checkerboard_1920_1080_10_3_1_0.yuv.

Alternatives considered

Option Pros Cons Why not chosen
Vectorize anti_dithering_filter in AVX2 ~3.7× kernel speedup, bit-exact, zero API changes Requires x86 AVX2 path Chosen: clean, isolated, bit-exact performance win.
Use OpenMP for CAMBI preprocessing Multi-core scaling Thread creation overhead on small stencils, non-deterministic scheduling High overhead for 576p/1080p frames, breaks single-thread predictable execution.
Keep si_mat_mul scalar No code changes Leaves 4× matrix multiply performance on the table Unnecessary now that GPU flush ordering is stabilized.
Dispatch si_mat_mul via speed_matmul_avx2/avx512 4–5× speedup on matrix operations, bit-identical Adds header includes in speed_internal.c Chosen: leverages already-tested bit-exact SIMD libraries.

Consequences

  • Positive:
  • Measurable throughput increase in CAMBI on 8-bit luma inputs with zero drift in scores.
  • Faster matrix factorisation in SpEED GPU twins.
  • Seamless operation of --threads N across all GPU backends including HIP.
  • testdata/bench_all.sh executes reliably out of the box on multi-backend workstations.
  • Negative:
  • None. All numerical outputs are bit-identical to the scalar references.
  • Neutral / follow-ups:
  • Documented benchmark baseline in research digest and PR body.

References

  • Issue #1245: Benchmark + tuning performance pass
  • ADR-0108: Deep-dive deliverables rule
  • ADR-0964: SpEED internal API for GPU twins
  • ADR-1196: matrix_mul SIMD dispatch
  • ADR-1197: Threaded GPU flush ownership