ADR-1384: Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic¶
- Status: Accepted
- Date: 2026-09-30
- Deciders: lusoris
- Tags: hip, gpu, speed, performance, numerics, rc3, fork-local
Context¶
speed_chroma_hip and speed_temporal_hip kept the ADR-0567 split: the host copied and filtered each plane (picture_copy, speed_internal_filter_and_downscale), the device computed means and covariance, the host read the covariance back and waited on the stream, ran the eigenvalue problem, the regularity decision and the QR factorisation, re-uploaded, and read the entropies back again to score. Every frame round-tripped through the host several times and the scores differed from the CPU in the last bits (T-HIP-SPEED-HOST-RESIDUAL-2026-09-29).
ADR-1358 moved the whole chain onto the device for SYCL and made it bit-identical to the CPU: every stage reproduces speed.c operation for operation in fp32, the fp64 expressions of the reference are reproduced with exact fp32 pairs, and division, square root and log2 are correctly rounded. The RC3 brief asks for the same on HIP, with no GPU-CPU round trip, one wait per frame at collect, and one implementation per behaviour (HISS-19). HIP brings its own rounding hazards: hipcc contracts a * b + c into FMAs by default (-ffp-contract=fast-honor-pragmas), and without OCML_BASIC_ROUNDED_OPERATIONS the __fsqrt_rn() intrinsic is the native approximate square root while __fmul_rn() / __fadd_rn() / __fdiv_rn() are the plain operators, so they contract too (__clang_hip_math.h, ROCm 7.15).
Decision¶
- The ADR-1358 chain, one pipeline for both twins.
core/src/feature/hip/speed_hip_pipeline.{h,c}owns one device arena, pinned staging and result blocks, and a stream (VmafHipKernelLifecycle). Per frame the extractor uploads its raw planes throughvmaf_hip_picture_upload_staged()(ADR-1377) and submits eight kernels (speed/speed_pipeline.hip): optional prescale, the anti-alias filter at the 16x-decimated points, mean subtraction with the independent term, the 25 submatrix means, the 25x25 covariance as exact fp32 pairs, eigenvalues + regularity + QR in one work-group per channel, the per-block solve / variance / entropy, and the frame score; then oneSpeedGpuFrameResultcopy.collect()is the only wait.speed_temporal_hipkeeps the previous frame's raw planes on the device in two slots and differences them there. - Exact arithmetic per TU, not per operation. The SpEED kernel TU is built with
-ffp-contract=offand-fhip-fp32-correctly-rounded-divide-sqrt(hip_cu_extra_flagsincore/src/meson.build), and the device code writes plain*,+,/andsqrtf(). Per-operation__fmul_rn/__fadd_rn/__fdiv_rn/__fsqrt_rnwould be wrong on HIP (see Context). The device LLVM IR carries nocontractflag, no!fpmathrelaxation and nodoublein any SpEED kernel; the onlyllvm.fmuladdcalls come fromsinpif()in the lanczos4 prescale path, which ADR-1358 already leaves at the ADR-0214 tolerance because the CPU evaluates those weights in fp64. log2 isspeed_hd_log2_rn(), fp32-pair evaluated and rounded once, as on SYCL and CUDA. - One header for the kernels and the replay. Every per-work-item routine is in
speed/speed_hip_device.h, compiled by hipcc and, with contraction off, by the host compiler forcore/test/test_hip_speed_device_math.c; the replay builds its parameter block with the pipeline's ownspeed_hip_params_fill(),speed_hip_taps_fill()andspeed_hip_bindings_*(). - One init-time configure for every backend. The geometry, filter taps and scoring constants come from
speed_internal_gpu_configure()inspeed_internal.c, filling the backend-neutralSpeedGpu*types ofspeed_gpu_common.h; the SYCL host setup (speed_sycl_host.cpp) now calls it too and its types alias the shared ones. The CUDA port carries the identical routine; whichever lands second keeps one copy. - The contract. Every per-frame
speed_chroma_u/v/uvandspeed_temporalequals the CPU extractor bit for bit when the CPU'slog2fis correctly rounded (libimf, or glibc with a correctly roundedlog2f); the host replay asserts this on the Netflix pair and synthetic 10-bit, prescale, weighting-mode,speed_use_ref_diffand singular fixtures. glibc'slog2fmisrounds 0.38 % of arguments in [1, 8), so against a glibc-built CPU extractor a few chroma frames differ in the last float bits (6 of 48 on the Netflix pair, at most 4.8e-7);speed_temporalstays identical there.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
Per-TU -ffp-contract=off + correctly rounded / and sqrtf() (chosen) | Every operation rounds where speed.c rounds; plain code the host replay compiles identically; checked in the IR | Relies on a build flag the contract test pins | — |
__fmul_rn / __fadd_rn / __fdiv_rn / __fsqrt_rn per operation | Explicit in the source | On HIP these are the plain operators (contracting) and the native square root unless OCML_BASIC_ROUNDED_OPERATIONS is set; would silently not round correctly | Wrong on this toolchain |
| Keep the host linear algebra, one readback per frame | Smaller change | Still a round trip and a mid-frame wait; the brief rules it out | Brief |
| A HIP-private copy of the SYCL init-time configure | No change outside HIP | A second implementation of the same setup (HISS-19) | HISS-19 |
Reproduce glibc's log2f on the device | Bit-exact against glibc builds | glibc evaluates in double; the device contract is fp64-free, and libimf builds would then differ | The CPU reference differs by libm; the correctly rounded value is the portable target, as on SYCL and CUDA |
Consequences¶
- Positive: no host stage, one upload, one readback and one wait per frame; the scores reproduce
speed.cin fp32 (bit-identical against a correctly roundedlog2f). The replay test (fastsuite) pins the arithmetic without an AMD device and kills every planted regression of the device header and the parameter block. - Negative: not run on an AMD device in this change; the verify and timing commands, including the correctly rounded
log2fshim for glibc builds, are in theT-HIP-SPEED-HOST-RESIDUAL-2026-09-29row ofdocs/state.md. The eigenvalue sweep runs in one work-group per channel, latency-bound as on SYCL. - Neutral / follow-ups: lanczos4 prescale stays within the ADR-0214 tolerance.
test_hip_device_resident_contract.pypins the design at the source level: no host stage, staged upload, one readback, the wait incollect(), fp64-free device code without inexact intrinsics, and the kernel's exact-arithmetic flags.
References¶
- req: RC3 port brief (2026-09-30): "there shouldnt be any gpu cpu rountrips"; SpEED "fully on the device ... bit-exact"; "on HIP use __fmul_rn/__fadd_rn/__fdiv_rn/__fsqrt_rn per operation or -ffp-contract=off per TU (say which and why)".
- ADR-1358 — the SYCL design ported here; Research-1358.
- Research-1378 — HIP rounding intrinsics, the IR check, the glibc
log2fmeasurement and the host replays. - ADR-0567, ADR-1218, ADR-1377, ADR-0214, ADR-0220.