ADR-1179: Fix Intel Arc SYCL Crashes and Default Model Resolution Divergence¶
- Status: Accepted
- Date: 2026-09-05
- Deciders: maintainers
- Tags: sycl, gpu, cambi, speed, model, default, arc, fp64, adr-0220
Context¶
Running vmaf --backend sycl --model version=vmaf_v1.0.16_3d0h (the default model on master per ADR-1169) on Intel Arc GPUs (e.g. Arc A380 / ACM-G11) failed with multiple distinct fatal crashes and execution errors:
- SIGSEGV in
integer_cambi_sycl.cpp: The SYCL CAMBI extractor attempted to bisect and re-evaluate contrast limits oninit(), but leftbuffers.v_band_baseandbuffers.v_band_sizeuninitialized, causing invalid device pointer offsets and buffer overrun. Furthermore, the histogram buffer was sized only tonum_binswithout consideringv_band_size. - SIGABRT in
speed_chroma_sycl.cppandspeed_temporal_sycl.cpp: Intel Arc A-series GPUs lack hardware double-precision support (aspect::fp64). Both kernels allocateddoubleaccumulators andsycl::local_accessor<double, 1>work-group arrays, triggering an uncaught SYCL exceptionterminate_handler: Device does not have aspect fp64. This directly violated ADR-0220 ("SYCL fp64-less device contract"). - Exit 255 (
problem generating pooled VMAF score): The default modelvmaf_v1.0.16_3d0hconfigures CAMBI withcambi_high_res_speedup: 1080. The CPU CAMBI extractor registers this parameter asVMAF_OPT_FLAG_FEATURE_PARAM, serializing the feature name ascambi_hrs_1080_cmxv_17_vlt_0.06. However,options_cambi_syclomittedcambi_high_res_speedup(hrs), causingvmaf_feature_name_dict_from_provided_featuresto omit_hrs_1080from the SYCL feature dictionary. When the prediction engine attempted to fetchcambi_hrs_1080_...,predict_load_feature_scorereturned-EAGAIN(-11), aborting execution. - Sub-ULP Float Drift in Host SIMD Build under icx: When building with the oneAPI toolchain (
icx),x86_avx2_static_libandx86_avx512_static_liblacked_x86_simd_strict_fp_extra(-fp-model=precise), causingicxto apply-fp-model=fast=1and driftfloat_motionscores outside the Netflix golden assertion threshold by 0.00007.
Decision¶
We fix all six items (the four failure points, a test-harness flush, and a regression test) at their respective architectural boundaries:
- Cambi Initialization & Histogram Sizing: Expose
vmaf_cambi_init_tvi_and_vltincore/src/feature/cambi.candcambi_internal.h. Ininteger_cambi_sycl.cpp, invoke this helper oninit()to properly populate TVI, VLT, and contrast tables. Size device histogram buffers withMAX(num_bins, v_band_size). - FP64 Elimination in Speed Extractors: Replace all
doubleaccumulators, work-group local accessors, and reduction buffers withfloatinspeed_chroma_sycl.cppandspeed_temporal_sycl.cpp, adhering strictly to ADR-0220. cambi_high_res_speedupParity: Addcambi_high_res_speedup(alias"hrs") tooptions_cambi_syclwithVMAF_OPT_FLAG_FEATURE_PARAM. Add resolution checks againstCAMBI_HIGH_RES_SPEEDUP_THRESHOLD_*, implement window adjustment incambi_sycl_adjust_window, and update scale decimation logic to ensure identical feature dictionary keys and numerical equivalence with CPU CAMBI.- NULL Picture Flush Tolerance: Guard against NULL pictures in
core/test/test_integer_cambi_sycl.cto prevent test harness segfaults during EOS flush. - Strict-FP on SIMD Static Libraries under icx: Add
_x86_simd_strict_fp_extra(-fp-model=precise) tox86_avx2_static_libandx86_avx512_static_libincore/src/meson.build, ensuring that host SIMD floating-point routines compiled withicxmaintain bit-exact parity with GCC reference builds. - Dedicated Regression Test: Add
python/test/sycl_default_model_test.pywith an automated 2-frame SYCL probe that verifies that--backend syclwith the default model executes without crashing and outputs valid pooled and per-framevmafscores.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
Fall back to CPU for cambi or speed when SYCL default model is invoked | Avoids modifying SYCL kernels | Defeats the purpose of GPU acceleration; causes silent fallback violating ADR-0214 | Violates project charter against silent narrowing of scope. |
Loosen Netflix golden assertions for float_motion under icx | Easy one-line change in Python tests | Breaks the cardinal project invariant: NEVER modify Netflix golden data assertions | Hard rule: golden assertions are ground truth. |
| Re-implement full TVI table generation in SYCL | Decouples from cambi.c | Duplicate math code prone to drift; violates DRY | Sharing vmaf_cambi_init_tvi_and_vlt guarantees bit-exact table parity. |
Consequences¶
- Positive:
vmaf --backend syclruns out of the box with the default model (vmaf_v1.0.16_3d0h) on Intel Arc GPUs.- No crash (SIGSEGV, SIGABRT) on the three verified pairs (576x324 src01, 1080p 1-px and 10-px checkerboards) on an Arc A380.
- Pooled
vmafSYCL vs CPU: 82.814061 vs 82.816062 at 576x324 (delta 2.0e-3), identical at six decimals on both 1080p checkerboards;integer_adm3/integer_motion3identical at six decimals,speed_chroma_uvwithin 2e-6. GPU twins are not bit-exact to the CPU: the residualcambidrift (2.7e-3 pooled at 576x324) is tracked indocs/state.mdas T-SYCL-CAMBI-PARITY-DRIFT-2026-09-05. - Netflix CPU golden test suite (
make test-netflix-golden) remains 100% green (271 passed, 12 skipped, 0 failed). - Negative:
- None. Code complexity is minimal and reuses existing C helpers.
- Neutral / follow-ups:
python/test/sycl_default_model_test.pyensures CI runners with Level Zero hardware continuously gate this path against regressions.
References¶
- ADR-0220: SYCL fp64-less device contract
- ADR-0214: GPU-parity CI gate
- ADR-1169: Default VMAF model vmaf_v1.0.16_3d0h
- Netflix/vmaf issue discussions regarding CAMBI high resolution speedup