ADR-1143: CUDA and Intel SYCL Backend Gap Closure¶
- Status: Accepted
- Date: 2026-09-02
- Deciders: Kilian, Antigravity Agent
- Tags: cuda, sycl, ci, dispatch, python, docs
Context¶
An architectural audit of the CUDA and Intel SYCL backends in VMAFx identified 8 operational gaps across backend dispatch, runtime configuration, build artifacts, and documentation:
- Python Harness Backend Selection Ignored (
GAP-CI-VMAF-FORCE-BACKEND-IGNORED):.github/workflows/tests-and-quality-gates.ymlinvoked pytest withVMAF_FORCE_BACKEND=cudaandVMAF_FORCE_BACKEND=sycl, but neitherExternalProgramCallernorcompat/python-vmaf/__init__.pyreadVMAF_FORCE_BACKEND. Furthermore, evaluating arbitrary Python tests against GPU backends would fail on Netflix CPU-golden assertions (assertAlmostEqualbit-exact float checks) rather than comparing against calibrated GPU ULP tolerances. - Unwired Dispatch Environment Getter (
GAP-GPU-DISPATCH-ENV-UNWIRED-IN-BACKENDS):core/src/gpu_dispatch_env.cppdefined the thread-safe cached gettervmaf_gpu_dispatch_env_get(ADR-0858), butcore/src/cuda/dispatch_strategy.candcore/src/sycl/dispatch_strategy.cppbypassed it with rawgetenv()calls, redundantpthread_once/InitOnceExecuteOnceguards, and// NOLINT(concurrency-mt-unsafe). - Stale TensorRT EP Documentation (
GAP-DOCS-INDEX-TENSORRT-EP-UNIMPLEMENTED):docs/index.md:92claimed TensorRT Execution Provider was supported in Tiny-AI, but no TensorRT integration exists incore/src/dnn/ort_backend.c. - Dead CUDA Source Files (
GAP-BUILD-DEAD-CUDA-ADM-DECOUPLEandGAP-BUILD-UNCOMPILED-CUDA-RESOLUTION-DISPATCH):core/src/feature/cuda/integer_adm/adm_decouple.cuwas superseded by inline decoupling inadm_csf.cuviaadm_decouple_inline.cuh, andcore/src/feature/cuda/resolution_dispatch.c/.hwere completely uncompiled and orphaned. - Dishonest CUDA Graph Dispatch Stub (
GAP-CUDA-DISPATCH-STRATEGY-GRAPH-STUB):core/src/cuda/dispatch_strategy.cparsedgraphand returnedVMAF_CUDA_DISPATCH_GRAPH_CAPTURE, misleading users even though graph capture is not implemented for driver API CUDA extractors. - Undocumented CUDA Zero-Copy Import Gap (
GAP-CUDA-NO-ZERO-COPY-IMPORT):core/src/cuda/picture_cuda.cstages picture transfers through pinned host memory rather than importing external memory or Linux DMA-BUFs directly. - SYCL DMA-BUF Windows Stub (
GAP-SYCL-DMABUF-IMPORT-WIN32-ENOSYS):core/src/sycl/dmabuf_import.cppreturned-ENOSYSon_WIN32silently without explaining that DMA-BUF is a Linux kernel primitive.
Decision¶
We resolve all 8 gaps in a unified change:
- Python Harness & CI Integration:
ExternalProgramCaller.call_vmafexec_multi_featuresandcall_vmafexecincompat/python-vmaf/__init__.pyreadVMAF_FORCE_BACKEND(and fallbackVMAF_BACKEND), appending--backend <val>to the childvmafcommand line.- Scoped
.github/workflows/tests-and-quality-gates.ymlGPU pytest legs away from the 5 Netflix CPU golden assertion files (quality_runner_test.py,feature_extractor_test.py,vmafexec_test.py,vmafexec_feature_extractor_test.py,result_test.py). - Covered with 4 new unit tests in
python/test/python_harness_coverage_test.py. - GPU Dispatch Environment Integration:
core/src/cuda/dispatch_strategy.candcore/src/sycl/dispatch_strategy.cppinvokevmaf_gpu_dispatch_env_getdirectly.- Removed local once-initialization blocks, duplicate mutexes, and
// NOLINTsuppressions. - Linked
gpu_dispatch_env_cpp23_libintolibvmaf_feature_static_libincore/src/meson.buildso all consumers, intermediate libraries, and tests link cleanly across all backend configurations. - Tested in
core/test/test_gpu_dispatch_runtime.cwith new SYCL dispatch override test cases. - Honest CUDA Graph Dispatch:
- In
vmaf_cuda_select_strategy, ifgraphstrategy is parsed, log a clear warning:libvmaf: CUDA graph dispatch requested for '%s' but graph capture is not implemented; falling back to directand returnVMAF_CUDA_DISPATCH_DIRECT. - Dead File Cleanup:
- Removed
core/src/feature/cuda/integer_adm/adm_decouple.cu,core/src/feature/cuda/resolution_dispatch.c, andcore/src/feature/cuda/resolution_dispatch.h. - SYCL Win32 Informative Error Logging:
- In
core/src/sycl/dmabuf_import.cpp, log an informative error explaining that DMA-BUF is a Linux kernel primitive before returning-ENOSYSon_WIN32. - Documentation & State Tracking:
- Corrected
docs/index.md(TensorRT EP marked as roadmap pointing todocs/ai/roadmap.md). - Updated
docs/backends/cuda/overview.md,docs/backends/index.md, anddocs/usage/env-vars.md. - Recorded Deferred and Confirmed not-affected rows in
docs/state.mdand updateddocs/rebase-notes.md.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Implement full CUDA graph capture now | Enables graph execution for CUDA extractors | High complexity (>800 LOC), requires persistent buffer reuse and graph node rebuild across dynamic pitches and frame sizes | Out of scope for gap closure; honest warning + fallback is the correct maintainable posture |
Keep raw getenv with NOLINT in CUDA & SYCL | Zero churn in dispatch code | Duplicates once-cached environment logic; leaves mt-unsafe warnings and dead gpu_dispatch_env.cpp | Defeats the purpose of the unified gpu_dispatch_env subsystem introduced in ADR-0858 |
| Run full python test suite on GPU in CI | Exercises maximum tests with GPU backend | Fails immediately against Netflix CPU golden assertions that check bit-exact float equality | Violates AGENTS.md §8: Netflix golden assertions are CPU-only ground truth; GPU parity belongs in tolerance gates |
Consequences¶
- Positive:
VMAF_FORCE_BACKENDis now functional across all Python harness callers.- CUDA and SYCL backends share a unified, thread-safe, once-cached dispatch environment parser.
- Test and documentation surfaces are truthful and eliminate misleading claims about TensorRT and CUDA graph capture.
- 3 dead files removed from the codebase.
- Negative:
VMAF_CUDA_DISPATCH=graphexplicitly emits a warning and falls back to direct execution rather than attempting graph capture.- Neutral / follow-ups:
- Full CUDA graph capture and CUDA Linux DMA-BUF external memory import remain tracked in
docs/state.md(Deferred).
References¶
- User requirement: Close CUDA + Intel (SYCL) bucket of backend-gap inventory.
- ADR-0181: GPU dispatch strategy abstraction.
- ADR-0214: Cross-backend GPU parity CI gate.
- ADR-0483: GPU dispatch parse deduplication.
- ADR-0858: C++23 isolated static library for
gpu_dispatch_env.