ADR-1206: Every CUDA parity test also runs against a second, larger fixture¶
- Status: Proposed
- Date: 2026-09-06
- Deciders: Lusoris
- Tags: testing, cuda, ci, correctness
Context¶
Of the 77 cross-backend parity tests in core/test/, 67 pinned a single 256x144 fixture and the rest pinned one other small size. That is a systematic blind spot, not an incidental one: several extractors change behaviour with resolution, and none of those branches were reachable at the pinned size.
Concretely, the shared SSIM / MS-SSIM auto-scale is max(1, round(min(w, h) / 256)), so it is always 1 below min(w, h) = 384; and the ADM border crop is (int)(dim * 0.1 - 0.5), which is 0 only for small bands. Two real divergences hid behind exactly this gap — the speed_chroma 4K defect (ADR-1202) and the float-ADM edge-indexing defect (ADR-1204). Having the same gap produce two independent bugs makes closing it more valuable than either individual fix.
Adding a second fixture immediately paid for itself: it is what surfaced the float_ssim_cuda scale=1-only refusal at real resolutions, and it re-confirmed the ADM defect from the width side as well as the height side.
Decision¶
We will register a _large variant of every CUDA parity test, built from the same translation unit against a 960x540 fixture. 960x540 is chosen because min(w, h) = 540 puts the auto-scale at 2 (crossing the decimation boundary) and 540 is not a multiple of the 16/32-wide kernel blocks, so tail-bound handling is exercised as well. The fixture macros become #ifndef-guarded so the second size is supplied by -D from meson.build with no duplicated test source.
Where a GPU twin legitimately refuses the larger resolution — float_ssim_cuda is a documented v1 scale=1-only extractor — the variant asserts that documented contract (clean refusal, reported as a skip) rather than being dropped. That way, if the twin ever stops refusing and starts returning a scale=1 score at a decimating resolution, the test fails instead of silently comparing two different metrics.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
Second fixture size per test via -D and a meson foreach (chosen) | No duplicated test source; one obvious knob; catches the whole class | 16 extra small test binaries to build and run | — |
| Change the existing fixture to 960x540 instead of adding a variant | No new targets | Loses coverage of the small-band paths, which is where the ADM defect actually lives | Rejected — would trade one blind spot for another |
| Randomise the fixture size per run | Broadest coverage over time | Non-deterministic CI; a failure may not reproduce | Rejected — parity gates must be reproducible |
| Add sizes only to the tests known to be resolution-sensitive | Cheapest | Requires knowing which those are, which is the thing we got wrong twice | Rejected |
Consequences¶
- Positive: resolution-dependent divergence is now a tested property rather than something found by accident. The 16 variants run in ~2-6 s each.
- Negative: 16 additional test binaries to compile and run in the GPU lane. Build and runtime cost is small but not zero.
- Neutral / follow-ups: the CUDA and SYCL families are covered. Both were verified on real hardware — CUDA on an RTX 4090, SYCL on the Arc A380 — and the SYCL sweep immediately paid for itself the same way: it surfaced the
float_ssim_syclscale=1-only refusal and amotion_add_uvtolerance gap (below). HIP and Metal are deliberately not registered yet: this workstation cannot verify Metal at all, and shipping test registrations that have never been run is how a lane goes red for reasons nobody has looked at. They remain a follow-up. - Resolved follow-up (ADR-1326):
test_sycl_motion_add_uv_paritywas originally excluded because it compared CPU float arithmetic with the SYCL fixed-point path under a fixture-calibrated 2e-4 budget. ADR-1326 replaces that comparison with an arithmetic-identical scalar fixed-point oracle and a derived2*gamma_5host-roundoff bound. The 960x540 variant is now registered under the same resolution-independent contract.