ADR-1299: Compute MS-SSIM chroma on SYCL¶
Instead of accepting enable_chroma and ignoring it.
- Status: Accepted
- Date: 2026-09-23
- Deciders: Lusoris
- Tags:
sycl,gpu,feature-extractor,ms-ssim,parity - Supersedes: ADR-0526 in part — its
enable_chromareasoning only; theenable_lcsdecision stands.
Context¶
ADR-0526 added enable_chroma to the SYCL MS-SSIM twin and recorded that "the enable_chroma option is structural (MS-SSIM is luma-only by construction) and is accepted for option-table symmetry without semantic effect."
That premise is false. MS-SSIM is not luma-only by construction, and this repository is the proof: core/src/feature/float_ms_ssim.c has looped n_planes and emitted float_ms_ssim_cb / float_ms_ssim_cr since PR #939. The option was therefore not "structural" — it was unimplemented, and the ADR recorded the absence as a property of the metric.
The consequence was a silent one. configure_ms_ssim computed s->n_planes = enable_chroma ? 3 : 1 and nothing ever read n_planes — the identifier occurred exactly twice in the translation unit, once as the field and once in that assignment. Setting the option changed nothing, raised nothing, and provided_features advertised only float_ms_ssim, so a model asking for the chroma features had them served by the CPU twin under the ADR-0530 name-based fallback. Scores were right; the acceleration was absent; nothing said so.
Two further facts surfaced while closing this:
-
The CPU reference was broken below the exact 351x351 4:2:0 luma floor.
init()infloat_ms_ssim.cchecked the 5-level pyramid minimum against luma only, so a 4:2:0 input frommin_dimthrough2 * min_dim - 2passed and then died mid-run inside upstreamms_ssim.c, which printserror: scale below 1x1!to stdout and returns 1. Reproduced on this repository's own primary Netflix fixture:--width 576 --height 324 --pixel_format 420 --feature float_ms_ssim=enable_chroma=trueemitted that line, logged a bare "problem with feature extractor", and wrote no output file at all. 576x324 gives 288x162 chroma, and 162 < 176. -
docs/metrics/ms-ssim.mdclaimed the SYCL twin implemented chroma "fully (3 planes)". PR #1520 corrected that text to "accepted but a no-op" and filed a tracking row. Correcting the sentence was right; stopping there was not. A missing backend leg is not a documentation defect.
Decision¶
Implement it.
MsSsimPlaneGeometryholds the per-plane dimensions, staging buffers and pyramid. Geometry is per plane because a 4:2:0 chroma plane is a different size from luma and every kernel takes those dimensions as a row pitch; feeding a chroma plane through luma geometry reads at the wrong pitch, which is a wrong number and an out-of-bounds read rather than a crash.- The reduction workspace stays shared and plane-0 sized. No chroma plane is wider or taller than luma in any supported pixel format (
picture.c:146-149), so plane-0 sizing dominates. Planes run sequentially for the same reason the five scales already do: the horizontal intermediates and the partials are one workspace, andcompute_scale_lcswaits on its readback before returning. - Every plane is staged in
submit_fex_sycl, notcollect_fex_sycl, because libvmaf's double-buffered GPU dispatch releases theVmafPictureonce submit returns. provided_featuresadvertisesfloat_ms_ssim_cbandfloat_ms_ssim_cr, so the features are served by the GPU twin rather than silently by the CPU one.- The per-scale
l/c/sbreakdown stays luma-only, matching the CPU twin'sp == 0guard. - Both twins now reject an input whose chroma planes fall under the pyramid minimum, naming the actual chroma size and the luma resolution that would satisfy it.
Alternatives considered¶
| Option | Why not |
|---|---|
| Implement chroma on SYCL (chosen) | The option now does what it says, and the feature is accelerated rather than quietly served by the CPU. |
| Keep the no-op, document it accurately | What PR #1520 did. It makes the record honest and leaves the backend leg missing; a user enabling chroma on SYCL still gets no GPU chroma. |
Reject enable_chroma on SYCL, as CUDA does | Honest and cheap — CUDA has no such option and errors with -EINVAL. But it removes a capability the hardware can serve, and the CPU twin already defines the semantics. |
| Give every GPU twin the option and implement none | The status quo across HIP and Metal. It is the shape that produced this defect. |
Consequences¶
- Positive:
enable_chroma=trueon SYCL computes chroma on the GPU. Verified on an Intel Arc A380 against the CPU twin with a textured-chroma 4:2:0 fixture:float_ms_ssim_cb0.946694,float_ms_ssim_cr0.986886, identical to the CPU to all six emitted digits across three frames, with GPU time rising 15.32 ms to 18.57 ms per frame-pair as the extra planes are dispatched. - Positive: chroma MS-SSIM no longer dies mid-run when either 4:2:0 luma axis is from 176 through 350 pixels; such input is refused at init with the resolution it needs.
- Negative: three planes cost roughly 21% more GPU time when enabled. Default is unchanged (
enable_chroma=false). - Negative:
n_dispatches_per_frameonVmafFeatureCharacteristicsis a static field and cannot vary with an option, so the scheduler's cost model still reflects the luma-only dispatch count. Recorded rather than worked around. - Open: HIP (
integer_ms_ssim_hip.cassignsn_planes = 1uon both arms of itsif), CUDA (noenable_chromaat all) and Metal keep the same gap. This ADR closes SYCL only; the rest are tracked indocs/state.md.