ADR-1435: vif_hip reads the CPU's log2 table instead of evaluating log2f() on the device, and returns the CPU's scores bit for bit¶
- Status: Accepted
- Date: 2026-10-01
- Deciders: lusoris
- Tags:
hip,vif,gpu-parity,numerics,testing,ci,rc3,fork-local
Context¶
The fixed-point vif extractor (integer_vif.c) is integer arithmetic up to its last step. Per scale it filters in integers, forms three variances, and adds one or two logarithms per pixel into int64 accumulators; only vif_store_residuals() leaves the integers, with two float values per scale and a single-precision ratio. A GPU twin that accumulates the same integers returns the same scores.
vif_hip did not. Measured on a gfx1036 (ROCm 7.2.4) at --precision max against the CPU extractor on origin/master 80c5a0332:
| Fixture | Frames | Identical on scale 0 / 1 / 2 / 3 | Max abs diff |
|---|---|---|---|
| Netflix 576x324, 8 bit | 48 | 4 / 0 / 1 / 1 | 5.4e-7 |
| Checkerboard 1 px, 1920x1080 | 3 | 2 / 3 / 2 / 3 | 3.0e-8 |
| Checkerboard 10 px, 1920x1080 | 3 | 3 / 3 / 3 / 3 | 0 |
| Netflix 576x324, 10 bit | 3 | 1 / 0 / 0 / 0 | 3.6e-7 |
| Sparks 480x270, 10 bit | 5 | 0 / 0 / 3 / 1 | 3.6e-7 |
| BBB 3840x2160 | 48 | 5 / 3 / 5 / 3 | 3.0e-7 |
49 of 440 scores were the CPU's. Three runs of the twin gave the same values each time, so this is arithmetic, not the device's lost stream commands (T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01).
The host tail was already the CPU's (float sums, single-precision ratio), the filters and the variances are integers, and the gain is computed in fp64 on the device as on the CPU. What differed was the logarithm. The CPU reads it from a table of 32768 uint16_t entries that init() fills with the host math library, round(log2f(32768 + i) * 2048). The kernel computed __float2int_rn(log2f((float)v) * 2048.0f) per pixel on the device. A probe on the gfx1036 compared the two for all 32768 arguments:
- the device's
log2f()differs from glibc's in 15964 of them, by one ulp; - the product is exactly
k + 0.5for 80 arguments, whereroundf()rounds away from zero and__float2int_rn()to even; - 77 table entries end up one lower on the device (36 with
roundf()on the device, so both causes contribute).
Each pixel in the logarithm branch reads three entries, so nearly every frame had a numerator or denominator off by a few 2048ths.
Decision¶
We will make vif_hip read the CPU's table. vif_log2_table_generate() in integer_vif.h becomes the one definition of the table: the CPU extractor fills its state with it, integer_vif_hip.c builds the same 32768 values at init() and uploads them to a device buffer, and the horizontal kernels of vif_statistics.hip take that buffer as an argument and look every logarithm up (log2_lookup(), masked as log2_32() / log2_64() mask). The kernel file evaluates no logarithm. vif: hip is declared an exact twin, so the parity gate compares the CPU and HIP cells with tolerance 0 at --precision max.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Upload the table the CPU builds (this ADR) | The CPU's values by construction, whatever math library the host has; one definition of the table; three 16-bit loads per pixel instead of three log2f() evaluations | 64 KB more device memory and one more kernel argument | Chosen |
Keep the device log2f() and switch to roundf() | A one-line change | Fixes the 80 ties only; the one-ulp differences of the device's log2f() still move 36 entries | Not exact |
| Evaluate the logarithm in fp64 on the device and round to float | No table. On this host it gives the CPU's 32768 entries: glibc's log2f() is not the correctly rounded value for 5 of the arguments, and none of the 5 moves an entry | Three fp64 logarithms per pixel; and the CPU's table is whatever the host's log2f() returns, so the agreement holds for this glibc and is not a property of the code | Exact by coincidence of the host library |
Port glibc's log2f() to the kernel | No table, no memory reads | Ties the twin to one libc's algorithm and table; a macOS or Windows host has another | The CPU's values are defined by the host, so they have to come from the host |
| A static table in the kernel source | No upload | A third copy of the values, generated on one machine and wrong on a host whose log2f() differs | The same defect in another place |
| Append the table to the filter-table buffer | No new kernel argument | A buffer named for the filter taps that also holds logarithms, and an offset both sides must agree on | An explicit argument is clearer and costs nothing per frame |
Consequences¶
- Positive: measured on a gfx1036 at
--precision max, every score of every frame equals--backend cpu: 440 of 440 on the six fixtures above (49 before). Withdebug=truethe frame ratio and the per-scale numerator and denominator sums are equal too, and so are the scores at 12 and 16 bits, in 4:2:2, withvif_enhn_gain_limit=1.0and withvif_skip_scale0(see the HIP backend page). - Positive: the table has one definition.
integer_vif.c::log_generate()is gone; the CPU extractor and the HIP host callvif_log2_table_generate(). - Positive: a frame takes no longer: 44.4 and 42.7 ms at 1920x1080, 188.2 and 163.4 ms at 3840x2160, before and after (medians of 21 interleaved pairs of runs under other lanes' load; the samples overlap). A lookup replaces each
log2f(), and a pixel in the low-variance branch no longer computes the fp64 gain it does not use. - Negative: stored
vif_hipscores change by up to 5.4e-7. - Neutral / follow-ups:
integer_vif_sycl.cppandinteger_vif_metal.mmbuild the table with copies of the expression, andtest_integer_vif_log2.chas a fourth. They compute on the host and are correct; they should callvif_log2_table_generate()(other lanes own those twins).vif_cudastill evaluateslog2f()on the device. It measured equal to the CPU on an RTX 4090; whether it is on every argument has not been probed.- Guards:
test_hip_vif_parityand_large(seven cases, two frames each,==on every output; the first case fails on the old twin) andtest_hip_vif_log2_table_contract.py(five planted regressions, no device).
References¶
req(maintainer brief for the second HIP lane, 2026-10-01): "Every HIP twin returns the CPU extractor's bits, or differs only by the math library with a derived bound." and "Known start:vif_hipis 1.2e-7 off the CPU on a 4K frame (seen in #1749's test). Open its state row in your PR; it has none."- ADR-1421 (RC3: twin exactness), ADR-1397 (the exact cell), ADR-0500 (the 32768-entry table), ADR-0537 (device copies of host tables), ADR-1407, ADR-0214.
docs/state.md:T-HIP-VIF-DEVICE-LOG2-2026-10-01(opened and closed by this decision).