Skip to content

ADRs tagged rc3

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

77 ADR(s) carry this tag.

ID Title
ADR-1357 Run the SYCL CAMBI extractor entirely on the device
ADR-1363 The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame
ADR-1367 Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root
ADR-1369 SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence
ADR-1370 float_ssim_sycl decimates on the device, bit-identical to the CPU
ADR-1378 Run the HIP CAMBI extractor entirely on the device
ADR-1379 Run the CUDA CAMBI extractor entirely on the device
ADR-1380 The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic
ADR-1384 Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic
ADR-1386 One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks
ADR-1390 The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur
ADR-1391 The CUDA ssimulacra2 twin is device-resident
ADR-1395 SYCL kernels use no scratch memory on Intel GPUs
ADR-1397 GPU psnr_hvs twins reproduce the CPU's running float sum bit for bit
ADR-1399 float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic
ADR-1401 psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root
ADR-1403 Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs
ADR-1405 float_ssim_hip decimates on the device, bit-identical to the CPU
ADR-1407 Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root
ADR-1408 A VmafContext uploads each plane of a frame once and every HIP twin reads that copy
ADR-1409 float_motion_cuda adds its SAD in the CPU's order and returns the CPU's scores bit for bit
ADR-1410 SYCL CLI picture pool allocates pinned host USM to bypass staging upload
ADR-1411 float_motion_sycl adds its SAD in the CPU's order and returns the CPU's scores bit for bit
ADR-1412 float_vif_cuda computes the CPU's arithmetic, adds in the CPU's order and returns its scores bit for bit
ADR-1414 float_ms_ssim_sycl computes the CPU's arithmetic and returns its per-scale means bit for bit
ADR-1415 every x86 SIMD library is built without FP contraction
ADR-1416 adm_cuda takes its CSF weights, its rounding shifts and its score conclusion from the CPU's routines and folds the denominator once per row
ADR-1418 Motion parity cells compare what every twin emits; a missing metric is a cell error
ADR-1419 float_motion_hip stores its differences transposed and adds each row in the CPU's order
ADR-1420 float_adm_cuda computes the CPU's arithmetic, divides through the host's reciprocal estimate and returns the CPU's scores bit for bit
ADR-1422 float_vif_sycl computes the CPU's arithmetic without an fp64 type and returns its scores bit for bit
ADR-1423 adm_hip takes its weights, shifts and score conclusion from the CPU, folds the denominator once per row and clears its accumulators after the upload
ADR-1424 integer_ssim_cuda adds its terms in the CPU's raster order, on the host, and returns the CPU's score bit for bit
ADR-1426 ciede_cuda computes the CPU's arithmetic and adds in the CPU's order; what remains is the math library, and the gate bounds it
ADR-1427 A HIP frame queues its accumulator clears after its upload
ADR-1428 Exact GPU twins are declared by one fragment file each, not by a shared literal
ADR-1430 speed_chroma_cuda keeps its correctly rounded log2; the gate bounds what glibc's log2f adds
ADR-1432 vif_sycl computes the gain terms of the integer VIF in exact integer arithmetic and returns the CPU's scores bit for bit
ADR-1433 ssimulacra2_cuda returns the sums of the CPU's loops, formed on the device from integer increments per binade
ADR-1434 float_adm_sycl computes the CPU's arithmetic without an fp64 type, adds in the CPU's order and returns the CPU's scores bit for bit
ADR-1435 vif_hip reads the CPU's log2 table instead of evaluating log2f() on the device, and returns the CPU's scores bit for bit
ADR-1436 ciede_sycl runs the CPU's statements on fp32 pairs and adds in the CPU's order; it lands where the CUDA twin does, 1.4e-11 from the CPU
ADR-1437 motion_hip, motion_v2_hip, psnr_hip, integer_ms_ssim_hip and cambi_hip are declared exact twins; float_psnr_hip and float_moment_hip are not
ADR-1438 integer_ssim_hip adds its terms in the CPU's raster order at every frame size and returns the CPU's score bit for bit
ADR-1440 float_psnr_hip adds its squared differences as integers and returns the CPU's score bit for bit
ADR-1441 float_ssim_hip forms its window sums through the arithmetic float_ms_ssim_hip shares with the CPU and returns the CPU's score bit for bit
ADR-1442 float ADM divides; the processor's reciprocal estimate leaves the reference and the CUDA twin
ADR-1443 integer_ssim_sycl computes the CPU's fp64 term in 64-bit integers and adds in the CPU's order; it returns the CPU's ssim bit for bit
ADR-1444 float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit
ADR-1445 ssimulacra2_hip evaluates the CPU's fp64 terms and returns the sums of the CPU's loops
ADR-1446 ssimulacra2_sycl forms the CPU's fp64 terms in 64-bit integers and returns the sums of the CPU's loops; it returns the CPU's score bit for bit
ADR-1447 float_moment_hip adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1448 ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair
ADR-1449 float_moment_sycl adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1450 float_psnr_sycl adds its squared differences as integers, and is bit-identical to the CPU
ADR-1451 adm_sycl, motion_sycl, motion_v2_sycl, psnr_sycl, float_ssim_sycl and cambi_sycl are declared exact twins; speed_chroma_sycl is not
ADR-1452 the gate bounds speed_chroma_hip against the CPU by what glibc's log2f adds, as it does for the CUDA twin
ADR-1453 float_moment_cuda adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1454 A large subtree AGENTS.md is a generated index over one page per topic
ADR-1455 float_psnr_cuda adds its squared differences as integers, and is bit-identical to the CPU
ADR-1456 vif_cuda keeps its device log2f(), proven equal to the CPU's log2 table on every entry, and is declared an exact twin
ADR-1457 motion_cuda, motion_v2_cuda, psnr_cuda, float_ssim_cuda, float_ms_ssim_cuda and cambi_cuda are declared exact twins
ADR-1458 float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit
ADR-1459 SpEED's covariance kernels return the scalar kernel's bits; they vectorise across sums, not within one
ADR-1460 speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell
ADR-1461 no C or C++ translation unit is built with floating-point contraction; the strict policy is a project argument
ADR-1462 vif_cuda reads the CPU's log2 table instead of evaluating log2f() on the device
ADR-1463 float_ssim_sycl forms the CPU's fp64 terms in 64-bit integers and adds them on the host in the CPU's raster order
ADR-1464 float_ssim_cuda adds its frame sums in the CPU's raster order, on the host
ADR-1465 float_ms_ssim_cuda adds the terms of every scale in the CPU's raster order, on the host
ADR-1466 float_ms_ssim_sycl stores every window's l, c and s of every scale and adds them on the host in the CPU's raster order
ADR-1467 ciede.c writes its squares as products, so ciede2000 no longer depends on the compiler or on the C library's powf for them; a GCC build moves by up to 2e-11
ADR-1468 A SYCL kernel requires only a sub-group size every default AOT target accepts (16 or 32); the six kernels that required 8 move to 16
ADR-1469 the psnr_hvs SIMD butterfly is two functions, cut at the same statement in the AVX2 and the NEON twin
ADR-1470 An enum that C and C++ translation units both see has one size: no C++-only underlying type other than int's width
ADR-1472 The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale
ADR-1473 The x86 float ADM wavelet and CSF kernels return the scalar bits and are dispatched; the two reduction kernels are removed