Research-2099: Non-finite score laundering in the metric engine¶
Finding¶
A sweep of 466 feature-collector append sites found twelve expressions where IEEE-754 unordered comparisons hid a failed computation behind a finite score. The defect was not that a NaN could exist; it was that MIN, MAX, a threshold comparison, or a default output converted it into a number downstream could no longer distinguish from a measurement. The most severe mappings were SSIMULACRA2 NaN -> 100.0, ADM AIM 0/0 -> 1.0, and SSIM/MS-SSIM NaN -> max_db.
The finite paths are deliberately unchanged. Netflix golden assertions remain the authority and were not edited. The implementation validates immediately before a laundering comparison, returns -EINVAL, and avoids publishing any output for the failed frame. Where an enclosing function was already at the strict size limit, a pure scoring helper became the testable seam rather than adding an inline exception.
Reproductions¶
Three tests were first run against bug-preserving helper extractions:
piecewise_linear_mapping(NAN, ...)returned success and changed42.0to the plausible prediction0.0;- ADM with finite
num/denandaim_num/aim_den == 0/0returned success and changed the AIM output to the perfect1.0; - SSIMULACRA2's final mapping accepted
NANand returned the perfect100.0; - TransNet accepted a
NANlogit and published boundary flag0.0.
Each test failed before the finite guard and passes after it. They also cover infinities, unchanged caller outputs on failure, finite ratios and thresholds, ADM's legitimate flat-frame result, SSIMULACRA2's finite sign split, and the zero-logit TransNet threshold. The aggregate-prediction regression also drives vmaf_predict_score_at_index with a non-finite feature and intercepts its direct-source-inclusion diagnostic seam, proving exactly one diagnostic carries the stage, frame and value before the path leaves the caller output unchanged and publishes no model score. Normal builds take the adjacent vmaf_log branch with the human-readable warning format.
Backend scope¶
VIF, ADM, SSIM and MS-SSIM duplicate their host-side publication and conversion logic across CPU, CUDA, HIP, SYCL and Metal. The registered twins now validate raw operands and every enabled output before a clamp, fallback, dB conversion or first collector write. nonfinite_score.h owns the common finite-ratio, SSIM conversion and validate-before-publish operations; adm_score.h owns the ADM-family ratios and blend. This closes the same laundering defect on a selected GPU backend instead of fixing only the CPU reference.
SSIMULACRA2 likewise duplicated its host-side edge split and polynomial pool across the scalar reference, AVX2, AVX-512, NEON, SVE2, CUDA, HIP, SYCL and Metal hosts. ssimulacra2_score.h is the single implementation for both operations; each extractor still performs its own named frame log and returns -EINVAL before collector publication.
ADM's helper validates raw operands, denominators and computed results before writing either output. Aggregate numerator and denominator reductions are validated before the precision floor, because -Inf < limit would otherwise turn a failed reduction into finite zero. The finite 0/0 flat-frame ratio is defined as 1.0 for ADM, AIM and per-scale scores, while nonzero-over-zero and non-finite operands fail. ADM and AIM denominators are evaluated independently, so a flat ADM aggregate does not overwrite a finite AIM ratio or vice versa. The ADM3 helper retains the defined all-zero harmonic mean while rejecting a non-finite blend before adm_min_val can hide it.
MS-SSIM previously appended the luma plane before it knew whether an enabled chroma plane was valid. Its extraction now computes and validates all enabled planes and all L/C/S atoms first, then appends them, preserving the fail-frame atomicity stated by ADR-1302. L/C/S atoms are checked even when enable_lcs is off, closing the pow(NaN, 0) == 1 escape. Float and integer VIF collectors similarly validate all four ratios plus enabled debug atoms before publishing scale 0; scale 0 remains unclamped, while scales 1-3 then apply their configured minimum.
SSIM/MS-SSIM retain one explicit exception from ADR-1221: finite raw scores at or above 1.0 report positive infinity when dB output is enabled and clipping is disabled. The shared converter tests that case directly while rejecting NaN raw scores, invalid ceilings, and non-finite conversions. This distinguishes a documented output sentinel from the failed computations Issue #1526 targets.
Reproduce¶
meson setup core/build-nonfinite core -Db_lto=false \
-Denable_cuda=false -Denable_sycl=false -Denable_hip=false \
-Denable_metal=disabled -Denable_dnn=disabled
ninja -C core/build-nonfinite -j4
meson test -C core/build-nonfinite --print-errorlogs \
test_adm_nonfinite_score test_nonfinite_collector_wiring \
test_ssimulacra2_nonfinite \
test_ssimulacra2_coverage test_predict test_transnet_v2 \
test_float_ms_ssim_coverage
The release gate remains make test-netflix-golden; a passing unit test is not a substitute for the unchanged golden scores.
On the completed worktree, the focused seven-test command passes 7/7, the full Meson fast suite passes 150/150, and the five-file Netflix Python gate reports 271 passed, 12 skipped. No Netflix-authored assertion or expected value is changed. Release builds with CUDA 13.4/NVCC, ROCm 7.2/HIPCC (gfx1100), and oneAPI 2026.0/SYCL (SPIR-V JIT) also compile the shared SSIMULACRA2 guard and their respective host twins successfully. Metal compilation remains CI-only because this Linux workstation has no Apple toolchain; the wiring regression also checks its host and shader sources for the removed perfect-score fallbacks.
The backend-twin edits are the same mechanical helper substitution in files that carry historical HISS debt. The local governance gate was therefore run with PRAETOR_TOUCHED_DEBT_DELTA_REASON set to that exact parity rationale; the stricter debt-delta audit passes at 267 active findings against the 276 finding baseline, with no new debt.
Post-rebase hosted falsification¶
The hosted matrix on the rebased heads found four contracts that the local review had not initially proved:
- both macOS Metal lanes rejected
float_ms_ssim_metal.mmbecause the shared emitter call referenceds->enable_dbands->max_db, while the Metal state intentionally exposes onlyenable_lcsper ADR-0490 and ADR-1221; - multiple Linux/macOS build-matrix lanes rejected the compatibility Cython extension after
adm.cbegan includingnonfinite_score.h: the extension's direct-source build searchedcore/srcbut notcore/include, so the transitive publiclibvmaf/model.hinclude was unreachable; - the whole-tree Clang 22 tidy ratchet rejected the new
transnet_v2_score.hhelper because its three-way probability clamp omitted braces required byreadability-braces-around-statements; - after the public-header search path let the compatibility Cython extension finish linking, its first runtime import exposed
vmaf_logas undefined: the text-includedadm.cnow reaches the logging wrappers innonfinite_score.h, but the extension had not linkedcore/src/log.c.
The Metal call now supplies the existing surface's fixed linear-score settings, false, INFINITY, and a static wiring regression locks that contract down on non-Apple hosts. The Python extension now includes both ../core/src and ../core/include; an AST regression checks the build definition, and a real Python 3.14 editable-wheel build completes successfully with the corrected path. The TransNet helper now braces all three branches, with the hosted tidy artifact providing the exact three-diagnostic reproducer. These are build and style-only corrections. The extension now links core/src/log.c, and its setup metadata regression rejects any future removal of that runtime dependency. Together these corrections neither expand the Metal option surface nor change finite metric results.