Research-1402: Removing the int16 wrap of the integer ADM masking centre tap¶
- Status: Active
- Workstream: ADR-1402
- Last updated: 2026-10-01
Question¶
Integer ADM narrows the centre tap of its scale-0 masking threshold to int16. Does any Netflix golden assertion depend on that wrap? If not, which inputs change score when the tap is kept in int32, by how much, what must the vector and device kernels compute to stay bit-identical to the scalar code, and what does the change cost?
Sources¶
core/src/feature/integer_adm_kernels.h:adm_cm_thresh(),adm_cm_accum_round(),adm_cm_rows().core/src/feature/adm_cm_accumulator.h:adm_cm_excess_s0().core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c:cm_thresh_band_*(),cm_excess_*(),cm_accum_*(),cm_row_*().core/src/feature/cuda/integer_adm/adm_cm.cu,core/src/feature/hip/integer_adm/adm_cm.hip,core/src/feature/sycl/integer_adm_sycl.cpp,core/src/feature/metal/integer_adm.metal.- Netflix/vmaf#1602, first revision (SIMD follows the scalar wrap) and second revision (the wrap removed).
- Host: Ryzen 9 9950X3D, RTX 4090 (CUDA), gfx1036 (HIP), Arc A380 under the Linux xe driver (SYCL), gcc 16.2.1, icpx 2026.0. Before: master
2c3acf1c9. No Apple device.
Findings¶
The golden gate does not depend on the wrap¶
The gate was run first, with only the scalar change applied and the assembly paths disabled (-Denable_asm=false), so that every test took the changed code. Then it was run on the finished change with the default dispatch.
Build (golden profile of scripts/ci/setup-golden-build.sh, gcc) | Result of the five golden test files |
|---|---|
| master, default dispatch | 271 passed, 12 skipped |
| master, scalar only | 271 passed, 12 skipped |
| scalar prototype, scalar only | 271 passed, 12 skipped, same outcome per test |
| finished change, default dispatch | 271 passed, 12 skipped |
The 13 fixture pairs the golden tests read (the src01 pair at 8, 10, 12 and 16 bits and 4:2:2, both checkerboards, flat, sparks, two akiyo crops, the 160x90 src01 crop, KristenAndSara) give JSON identical to master at --precision max: with --feature adm --feature float_adm at --cpumask 0, 48 and 4294967295, and with the default model on the eleven pairs large enough for it.
Which inputs change, and by how much¶
CPU, default options, pooled mean, before to after. The values are the same at --cpumask 0, 48 and 4294967295.
| Input | Metric | Before | After |
|---|---|---|---|
flat grey 64x64 against 4x2 patches 255 0 0 0 every 16 pixels | integer_adm_scale0 | 1.0829225419556654 | 1 |
integer_adm2 | 1.035481944668303 | 1 | |
| flat grey 24x24 against one such patch at (3, 3) | integer_adm_scale0 | 1.0701309766616138 | 1 |
integer_adm2 | 1.0301034698295983 | 1 | |
| independent uniform 8-bit noise, 576x324, 3 frames | integer_adm2 | 0.38954843646923215 | 0.38950305912838107 |
integer_adm_scale0 | 0.4549724278265123 | 0.45480076976959144 | |
integer_adm3 | 0.36814184297691904 | 0.3681193593966255 | |
integer_aim | 0.653264750515394 | 0.65326434033513 | |
| the same noise against itself plus uniform [-16, 15] | integer_adm3 | 0.9771812019019719 | 0.9772170356634652 |
integer_aim | 7.166752298663627e-05 | 0 |
Unchanged: one-pixel vertical stripes against the same stripes shifted on every third row, salt-and-pepper impulses on a gradient, 6x10 blocks against a horizontal blur of them (all 576x324), and scales 1 to 3 of every input. On the two noise pairs the default model's own ADM features (computed with adm_csf_mode=2, which weights the bands differently) and its VMAF score are unchanged as well.
float_adm gives 1 on both patch pictures before and after. integer_aim on the patch pictures is 3.1756 (64x64) and 2.1640 (24x24) before and after, where float_adm's aim is 1; this change does not touch that difference and it was not investigated.
What "bit-identical to the scalar" requires of the vector kernels¶
The scalar excess is clamp(M - T * 2^s, 0, INT32_MAX) with M = |x| and T the threshold. In the kernels x is an int16 sample times a uint16 factor, so M < 2^31, and T is a sum of 24 int16 values and three taps of at most 69904, so |T| <= 996120. With s = 12 (diagonal band) the product leaves int32 from |T| = 2^19.
T >= 0andT * 2^s < 2^31:M - (T << s)is exact in int32 and only the clamp at 0 can act. This is the old 32-bit expression. Every decoded picture is in this case: its filtered neighbours are non-negative, and reaching2^19needs an average of 19418 per term.T >= 2^(31 - s): the result is 0. The 32-bit expression wraps the product and returns a positive excess.T < 0: the result ismin(M + |T| * 2^s, INT32_MAX). With|T|clamped to2^(31 - s)the sum is exact in uint32.
The vector rows therefore sum a row with the 32-bit expression and OR its thresholds together; when a bit of ~(2^19 - 1) is set in that OR (a negative threshold or one of at least 2^19), the row is summed again with the exact form: threshold clamped to +/-2^(31 - s), signed max(d, 0) for a non-negative threshold, unsigned min(d, INT32_MAX) for a negative one. A branch per block on the same test left the AVX-512 stage 7% slower than before the change; the per-row form leaves it 4% to 5% slower (see the stage times).
Upstream's second revision uses a different vector form: lanes with T > (INT32_MAX >> s) are zeroed, the rest take the 32-bit expression. That equals upstream's scalar for T >= 0 only. Planted into cm_excess_avx2() it fails test_integer_adm_simd (67x21, negative block: 0x1.211a32p+11 against the scalar's 0x1.211a3p+11).
The scalar narrows the squared excess to int32, so for an excess of 2^30 or more the cube is negative and its shift is arithmetic. AVX-512 has _mm512_sra_epi64. AVX2 has no 64-bit arithmetic shift; an arithmetic shift of a by n equals the logical shift of a + 2^63 minus 2^63 >> n, modulo 2^64, so the kernel adds 2^63 to the rounding term, accumulates the biased terms, and subtracts lanes * (2^63 >> n) from the row total. That costs no instruction per sample; upstream's variable-count sra_epi64 helper costs five per shift.
Planted defects¶
test_integer_adm_simd compares adm_cm_avx2() and adm_cm_avx512() with the scalar kernels on hand-built bands: dense fills with filtered bands of either sign and with non-negative ones, one full-range event at every position, and 3x4 blocks of -32768 or 32767 in the filtered bands at every column. Each of these defects, planted alone, fails it:
| Defect | AVX2 | AVX-512 |
|---|---|---|
| the row never takes the exact pass | fails | fails |
| the exact pass is triggered by the sign bit only | fails | fails |
| exact form without the threshold clamp | fails | fails |
| exact form without saturation of a raised excess | fails | fails |
| int16 wrap of the centre tap restored | fails | fails |
| tail block lanes off by one | fails | fails |
| tail block dropped | n/a | fails |
| tail block thresholds left out of the row's OR | fails | n/a |
| cube shift logical (no bias) | fails | n/a |
| bias count one block short | fails | n/a |
test_integer_adm_cm_threshold (flat reference against patches, every dispatch level) fails on master with 64x64 cpumask 4294967295 frame 0 integer_adm_scale0: 1.0829225419556654.
The undefined shift in the x86 tails¶
gcc 16.2.1, -Db_sanitize=address,undefined, 24x24 flat grey against one patch at (3, 3), vmaf --feature adm --no_prediction:
- master:
adm_avx512.c:2291:17: runtime error: left shift of negative value -15176(also lines 2293 and 2295); with--cpumask 48,adm_avx2.c:2687:17(also 2689 and 2691). - after: no report at
--cpumask0, 48 and 4294967295. The thirteen ADM unit tests pass under meson'sUBSAN_OPTIONS=halt_on_error=1.
Device twins¶
Before is a build of master 2c3acf1c9, after a build of the change; 20 pairs (the 13 fixture pairs and the seven synthetic ones above), --precision max.
| Twin | Before against after | After against the scalar CPU | Tests |
|---|---|---|---|
adm_cuda, RTX 4090 | 13 fixture pairs identical; the two patch pairs and the two noise pairs move by the CPU's amounts | within 2.56e-7 on 18 pairs, identical on 2 (the residual is the host-side float finalisation, unchanged) | test_cuda_adm_parity, _tiny_frames (7 cases), _small_border, _wide_rounding pass |
adm_hip, gfx1036 | 13 fixture pairs identical; patch pairs and independent noise move by the CPU's amounts (the twin emits no aim) | identical on 19 pairs; impulses on a gradient differ by up to 3.96e-7 on scales 2 and 3, before and after | test_hip_adm_parity, _tiny_frames (7 cases), _small_border, _wide_rounding pass |
adm_sycl, Arc A380 (xe) | 13 fixture pairs identical; the same four pairs move by the CPU's amounts | identical on all 20 pairs, all seven metrics | test_sycl_adm_parity, _tiny_frames (7 cases, bit-exact arm), _parity_large, test_sycl_kernel_scratch (108 kernels audited on master 67169ca4c; the 2 that use scratch memory are in the ratchet and neither is integer ADM) pass |
integer_adm_metal | not run | not run | source change only |
vmaf --backend hip --feature adm_hip --threads 4 exits with problem flushing context on this gfx1036, on master and after alike; the HIP runs above use no worker threads.
Stage times¶
Old and new objects linked into one benchmark and alternated on the same frames (gcc 16.2.1, -O3, no LTO; minimum of 4 to 6 runs; milliseconds per call at 1920x1080 8-bit unless noted). The same benchmark with the old source in both slots gives a ratio of 1.000.
| Level | Scale-0 contrast masking, before | after | Whole ADM pipeline, before | after |
|---|---|---|---|---|
| scalar | 2.240 | 2.394 (+7%) | 22.77 | 23.19 (+1.8%) |
| AVX2 | 0.565 | 0.537 (-5%) | 7.60 | 7.44 (-2.1%) |
| AVX2, 576x324 | 0.057 | 0.050 (-12%) | 0.577 | 0.553 (-4.2%) |
| AVX-512 | 0.240 | 0.253 (+5%) | 3.19 | 3.27 (+2.6%) |
| AVX-512, 576x324 | 0.024 | 0.024 | 0.304 | 0.310 (+2.0%) |
The scalar excess was first written with an early return for a masked sample. On noise the branch mispredicts: the scalar AIM contrast-masking stage took 4.40 ms against 2.30 ms. Two selects removed that. The split x86 files at first ran the leftover columns of a row through the shared scalar kernel, which cost the AVX-512 stage 33% at 576x324 (8 leftover columns per row); the leftover columns are now the top lanes of one more vector block.
The AVX-512 pipeline as a whole is 0% to 5% slower than before the split (typically 3%) at 0.2% to 0.5% more instructions (perf, cycles:u and instructions:u); the AVX2 pipeline executes 7% fewer instructions. The AVX-512 difference is tracked as T-ADM-AVX512-SPLIT-STAGE-TIME-2026-10-01 in docs/state.md. An end-to-end vmaf --feature adm run on 60 frames of BBB 3840x2160 does not resolve any of these differences on this host while other jobs run: repeated runs of the same two binaries ranged from 6% slower to 6% faster.
Found on the way, not changed here¶
- With a non-integer
adm_enhn_gain_limitthe AVX2 and AVX-512 decouple kernels roundrst * gainwhere the scalar truncates (T-ADM-DECOUPLE-X86-FRACTIONAL-GAIN-ROUNDING-2026-10-01). - Upstream's AVX2 scale 1-3 cube shift (
i4_cm_cube_avx2()) is logical where the scalar's is arithmetic. The two differ only for an excess of 2^30 or more, which the scale 1-3 extractor inputs do not reach; it is upstream's code and was left.
Reproduce¶
# golden gate
make test-netflix-golden
# the patch pictures, every dispatch level, and the vector kernels against the scalar
meson test -C build test_integer_adm_cm_threshold test_integer_adm_simd
# before / after on a patch picture (draw it as core/test/test_integer_adm_cm_threshold.c does)
vmaf -r flat_64x64.yuv -d patches_64x64.yuv -w 64 -h 64 -p 420 -b 8 \
--feature adm --feature float_adm --no_prediction --precision max --json
# device twins
meson test -C build-cuda test_cuda_adm_tiny_frames
meson test -C build-hip test_hip_adm_tiny_frames
meson test -C build-sycl test_sycl_adm_tiny_frames test_sycl_kernel_scratch