ADR-1369: SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence¶
- Status: Accepted
- Date: 2026-09-29
- Deciders: lusoris
- Tags:
sycl,gpu,performance,psnr,psnr-hvs,motion-v2,rc3,fork-local
Context¶
Every SYCL run uploads the luma of both pictures once per frame into the double-buffered shared frame (vmaf_sycl_shared_frame_upload, called from read_pictures_sycl_prep before any extractor submits). Several twins then uploaded the same frame again on their own. Profiled at 3840x2160 on an Arc B580 (Research-1369): psnr_hvs_sycl converted all three planes of both pictures to float on the host (14.5 ms) and uploaded 99.5 MB (7.2 ms of DMA) for a 6.3 ms kernel; motion_v2_sycl repacked and re-uploaded the reference luma; psnr_sycl repacked and uploaded its own copy of the chroma, which psnr_hvs then uploaded again. The maintainer's requirement is that planes do not make GPU-CPU round trips ("there shouldnt be any gpu cpu rountrips"). The shared frame is luma-only by design (core/src/sycl/AGENTS.md), because most runs (the default model) never read chroma.
Uploading into a slot also had no ordering against the kernels that last read it. The CLI's double-buffer dispatch waits for a frame's kernels in the next frame's collect(), before the upload that reuses the slot, but an extractor that skips a frame (n_subsample) has nothing collected, so a later upload could overwrite a slot its kernels were still reading. Moving psnr_hvs onto the shared slots would have extended that to it.
Decision¶
We will make the SYCL state the single owner of per-frame plane uploads, and have twins read those planes instead of uploading their own:
- Opt-in shared chroma planes in
core/src/sycl/common.cpp:vmaf_sycl_shared_chroma_init()(idempotent per geometry, aftervmaf_sycl_shared_frame_init()) allocates Cb / Cr of ref and dis for both slots;vmaf_sycl_shared_chroma_upload()puts the current frame's chroma into the compute slot once per frame — the first chroma-reading twin of a frame uploads, later ones return at once — by packing each plane into a pinned (host USM) staging buffer and issuing one DMA per plane on the copy queue, folded intolast_upload_event.vmaf_sycl_get_shared_plane()returns a compute-slot plane (0 = luma). Luma-only runs never allocate or upload chroma. vmaf_sycl_queue_after_upload()gives a twin on its own queue the same device-side input barriers the combined graph gets (last upload, last de-tile), without a host wait.- A device-side slot fence in
vmaf_sycl_shared_frame_upload(): the upload into a slot waits on the device for markers (ext_oneapi_get_last_event()of the primary and combined queues) taken when that slot stopped being the compute slot; the copy-queue barrier is only submitted when a marker is still running. - psnr_sycl reads chroma from the shared planes; motion_v2_sycl runs the ADR-1371 SAD pipeline on the shared luma and lets its
cur_copykeep the frame as the next frame's "prev"; psnr_hvs_sycl reads all three planes from the state as raw samples and scores every active plane in one dispatch, two work-items per 8x8 block, with per-block float arithmetic unchanged. - A psnr kernel rework that keeps the result bit-identical: one atomic per work-group instead of one per pixel, 32-bit squares. (A motion_v2 kernel that computed the vertical taps once per tile column in exact 32-bit arithmetic, 13.2 to 6.3 ms per 4K frame on the UHD 770, was dropped when ADR-1371 moved motion and motion_v2 onto one shared pipeline; it is a follow-up for that pipeline.)
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Shared chroma planes, uploaded lazily by the first twin that needs them (this ADR) | One upload per plane per frame for every twin; luma-only runs unchanged; no change to libvmaf.c or the extractor interface | Adds state and API to common.cpp; the zero-copy import path still has no chroma | Chosen |
Copy chroma straight from the picture (one memcpy, or ext_oneapi_memcpy2d for pitched rows) | No packing copy | The pictures are pageable: the driver stages on the calling thread, and the pitched 2-D copy took 5.1 ms per 576x324 frame on the B580 and 168 ms on the UHD 770 | Measured, replaced by pinned staging |
Always upload chroma with luma in vmaf_sycl_shared_frame_upload | Simplest ordering | Every SYCL run, the default model included, would pay 8.3 MB more DMA and 1 ms of host time per 4K frame for planes it never reads | Cost on the common path |
Upload chroma in read_pictures_sycl_prep when an extractor asked for it | Upload before any submit | Extractors initialise lazily inside the first frame's dispatch, after prep has run, so frame 0 would need a second path anyway | More code for the same result |
| Keep per-twin uploads, only drop the host float conversion in psnr_hvs | Smallest diff | Still uploads luma twice or three times per frame when twins run together | Leaves the round trips the maintainer asked to remove |
| motion_v2 reads "prev" from the other shared slot instead of its own copy | Saves the 8 MB device-to-device copy | The next frame's upload overwrites that slot while this frame's kernel may still run (the hazard the fence addresses would become the normal case) | Correctness |
| Host-side fence (wait for readers before uploading) | Simple | Serialises upload after compute on the host, losing the upload/compute overlap the double buffer exists for | Throughput |
| Allocate the CLI's picture pool in host USM so uploads stop blocking the host | Removes the remaining 3-4 ms of host time per 4K frame (driver staging of pageable memory) | Changes the generic picture pool and prepare_picture_pool, which the CLI read-ahead work also touches | Follow-up, recorded in Research-1369 |
| A psnr_hvs kernel with a parallel per-block float order | Much faster on the UHD 770 | Per-block scores stop being bit-identical to the current twin (the ADR-1361 gate would still hold) | Bit-identity kept; open question in Research-1369 |
Consequences¶
- Positive: at 3840x2160 on the B580 the per-frame cost psnr_hvs adds on top of the shared luma upload drops from about 29 ms (host conversion, private float upload, kernel) to about 3 ms (shared chroma upload and a 1.3 ms kernel); motion_v2 and psnr lose their private uploads and host repacking (motion_v2 about 1.6 ms of host copy and DMA per 4K frame). Twins run together share one chroma upload. Scores are bit-identical to the previous twins on the Arc B580 and the UHD 770; psnr and motion_v2 equal the CPU, psnr_hvs keeps its ADR-1361 distance. psnr_hvs at 9 and 11 bits now reads the raw sample the CPU reads (it used to see 16 times it). motion_v2 and luma-only psnr_hvs no longer dereference the NULL pictures the zero-copy import path passes to
submit(). - Negative:
common.cppowns more state (one aggregate member, the fence markers). On the zero-copy VA import path, which imports luma only, psnr and psnr_hvs with chroma now fail the frame with-EINVALand an error log; they used to crash there. psnr_hvs remains slow on the UHD 770 (49 ms per 4K frame). - Neutral / follow-ups: the pinned picture pool; other twins that stage chroma per extractor (
ciede_sycl,float_ssim_sycl/ssim_sycl,float_ms_ssim_syclwithenable_chroma,float_psnr_sycl) can move onto the shared planes the same way; the CUDA and HIP twins of these three features carry the same overheads (RC3 rows indocs/state.md).test_sycl_shared_planespins the helper contract and the three twins' per-frame CPU parity with changing chroma;test_sycl_init_unwindwraps the new init call.
References¶
- Source:
req— "benching = run it and find performance if possible, not just take a snapshot"; "there shouldnt be any gpu cpu rountrips". - Research-1369 — the profile before and after, ablations, parity evidence, CUDA/HIP review.
- ADR-1361 — psnr_hvs gate; ADR-0220 — fp64-free kernels; ADR-1121 — zero-copy import path; ADR-0458 — no per-step waits.