ADR-1237: CAMBI Anti-Dithering AVX2 Vectorization, SpEED SIMD QR Dispatch, and Threaded GPU Flush Alignment¶
- Status: Accepted
- Date: 2026-09-08
- Deciders: Kilian, Claude Opus 5
- Tags: perf, simd, x86, hip, cambi, speed
Context¶
Issue #1245 initiated a comprehensive benchmark and performance tuning pass across all four backends (CPU, CUDA, SYCL, and HIP) on workstation silicon (AMD Ryzen 9 9950X3D 32-thread CPU, NVIDIA RTX 4090, Intel Arc A380, and AMD Radeon Graphics gfx1036).
Profiling and harness validation surfaced several hot spots and pipeline bottlenecks:
- CAMBI anti-dithering filter: in
core/src/feature/cambi.c,anti_dithering_filter()applies a 2×2 box filter on 16-bit luma to suppress spatial dithering for all 8-bit inputs (enc_bitdepth < 10). This loop was purely scalar, processing pixels sequentially despite CAMBI's other stages having AVX2 paths. - SpEED host-side QR multiply:
core/src/feature/speed_internal.cdefinessi_mat_mul()for small-matrix QR factorisation in GPU SpEED twins. Per ADR-1196,si_mat_mulwas left scalar pending resolution of GPU motion flush ordering (T-GPU-MOTION-FLUSH-DOUBLE-EMIT-2026-09-06). With the flush mechanism stabilized, the scalar multiplication loop was ready for bit-exact SIMD dispatch. - Threaded GPU flush & extractor pooling: in
core/src/libvmaf.c,batch_extractor_skip()andread_pictures_should_skip()only excluded CUDA and SYCL from CPU thread pooling. When--threads Nwas passed with HIP, HIP feature extractors were erroneously scheduled into worker threads via.extract(which GPU extractors do not implement), returning-EINVAL(-22) and leavinggpu_pendinguncollected. - Harness robustness (
testdata/bench_all.sh): the candidate search order placed/opt/intel/oneapi-2025.3/setvars.shahead of/opt/intel/oneapi/setvars.sh(2026.0), causing symbol resolution failures when binaries were linked against 2026.0. In addition, Test 2 targeted a non-existent 1080p fixture name rather than the canonical checkerboard files inpython/test/resource/yuv/.
Decision¶
We make the following changes:
- Vectorize
anti_dithering_filterwith AVX2: implementanti_dithering_filter_avx2()incore/src/feature/x86/cambi_avx2.cand dispatch to it incambi.cwhenVMAF_X86_CPU_FLAG_AVX2is active. The kernel processes 16uint16_tlanes per iteration using zero-extended 32-bit addition, arithmetic shift, and_mm256_packus_epi32+ lane permute, guaranteeing bit-exact parity with scalar arithmetic. - Wire
si_mat_multo SIMD matmul: dispatchsi_mat_mulincore/src/feature/speed_internal.cthroughspeed_matmul_avx512/speed_matmul_avx2when CPU flags match, falling back tospeed_matmul_scalar. These kernels share the exact same accumulation order and compile under-ffp-contract=offto ensure bit-exact arithmetic. - Align GPU extractor pooling and threaded flush: update
not_pooledinbatch_extractor_skip()andgpuinread_pictures_should_skip()to includeVMAF_FEATURE_EXTRACTOR_HIPandVMAF_FEATURE_EXTRACTOR_METAL. Inflush_context_threaded(), draingpu_pendingfor non-CUDA/SYCL extractors before flushing, mirroringflush_context_serial(). - Fix
testdata/bench_all.sh: prefer/opt/intel/oneapi/setvars.shbefore legacy versions, and update Test 2 to usecheckerboard_1920_1080_10_3_0_0.yuvandcheckerboard_1920_1080_10_3_1_0.yuv.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
Vectorize anti_dithering_filter in AVX2 | ~3.7× kernel speedup, bit-exact, zero API changes | Requires x86 AVX2 path | Chosen: clean, isolated, bit-exact performance win. |
| Use OpenMP for CAMBI preprocessing | Multi-core scaling | Thread creation overhead on small stencils, non-deterministic scheduling | High overhead for 576p/1080p frames, breaks single-thread predictable execution. |
Keep si_mat_mul scalar | No code changes | Leaves 4× matrix multiply performance on the table | Unnecessary now that GPU flush ordering is stabilized. |
Dispatch si_mat_mul via speed_matmul_avx2/avx512 | 4–5× speedup on matrix operations, bit-identical | Adds header includes in speed_internal.c | Chosen: leverages already-tested bit-exact SIMD libraries. |
Consequences¶
- Positive:
- Measurable throughput increase in CAMBI on 8-bit luma inputs with zero drift in scores.
- Faster matrix factorisation in SpEED GPU twins.
- Seamless operation of
--threads Nacross all GPU backends including HIP. testdata/bench_all.shexecutes reliably out of the box on multi-backend workstations.- Negative:
- None. All numerical outputs are bit-identical to the scalar references.
- Neutral / follow-ups:
- Documented benchmark baseline in research digest and PR body.
References¶
- Issue #1245: Benchmark + tuning performance pass
- ADR-0108: Deep-dive deliverables rule
- ADR-0964: SpEED internal API for GPU twins
- ADR-1196:
matrix_mulSIMD dispatch - ADR-1197: Threaded GPU flush ownership