ADRs tagged rc3¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
77 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-1357 | Run the SYCL CAMBI extractor entirely on the device |
| ADR-1363 | The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame |
| ADR-1367 | Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1369 | SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence |
| ADR-1370 | float_ssim_sycl decimates on the device, bit-identical to the CPU |
| ADR-1378 | Run the HIP CAMBI extractor entirely on the device |
| ADR-1379 | Run the CUDA CAMBI extractor entirely on the device |
| ADR-1380 | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| ADR-1384 | Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic |
| ADR-1386 | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
| ADR-1390 | The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur |
| ADR-1391 | The CUDA ssimulacra2 twin is device-resident |
| ADR-1395 | SYCL kernels use no scratch memory on Intel GPUs |
| ADR-1397 | GPU psnr_hvs twins reproduce the CPU's running float sum bit for bit |
| ADR-1399 | float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic |
| ADR-1401 | psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root |
| ADR-1403 | Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs |
| ADR-1405 | float_ssim_hip decimates on the device, bit-identical to the CPU |
| ADR-1407 | Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1408 | A VmafContext uploads each plane of a frame once and every HIP twin reads that copy |
| ADR-1409 | float_motion_cuda adds its SAD in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1410 | SYCL CLI picture pool allocates pinned host USM to bypass staging upload |
| ADR-1411 | float_motion_sycl adds its SAD in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1412 | float_vif_cuda computes the CPU's arithmetic, adds in the CPU's order and returns its scores bit for bit |
| ADR-1414 | float_ms_ssim_sycl computes the CPU's arithmetic and returns its per-scale means bit for bit |
| ADR-1415 | every x86 SIMD library is built without FP contraction |
| ADR-1416 | adm_cuda takes its CSF weights, its rounding shifts and its score conclusion from the CPU's routines and folds the denominator once per row |
| ADR-1418 | Motion parity cells compare what every twin emits; a missing metric is a cell error |
| ADR-1419 | float_motion_hip stores its differences transposed and adds each row in the CPU's order |
| ADR-1420 | float_adm_cuda computes the CPU's arithmetic, divides through the host's reciprocal estimate and returns the CPU's scores bit for bit |
| ADR-1422 | float_vif_sycl computes the CPU's arithmetic without an fp64 type and returns its scores bit for bit |
| ADR-1423 | adm_hip takes its weights, shifts and score conclusion from the CPU, folds the denominator once per row and clears its accumulators after the upload |
| ADR-1424 | integer_ssim_cuda adds its terms in the CPU's raster order, on the host, and returns the CPU's score bit for bit |
| ADR-1426 | ciede_cuda computes the CPU's arithmetic and adds in the CPU's order; what remains is the math library, and the gate bounds it |
| ADR-1427 | A HIP frame queues its accumulator clears after its upload |
| ADR-1428 | Exact GPU twins are declared by one fragment file each, not by a shared literal |
| ADR-1430 | speed_chroma_cuda keeps its correctly rounded log2; the gate bounds what glibc's log2f adds |
| ADR-1432 | vif_sycl computes the gain terms of the integer VIF in exact integer arithmetic and returns the CPU's scores bit for bit |
| ADR-1433 | ssimulacra2_cuda returns the sums of the CPU's loops, formed on the device from integer increments per binade |
| ADR-1434 | float_adm_sycl computes the CPU's arithmetic without an fp64 type, adds in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1435 | vif_hip reads the CPU's log2 table instead of evaluating log2f() on the device, and returns the CPU's scores bit for bit |
| ADR-1436 | ciede_sycl runs the CPU's statements on fp32 pairs and adds in the CPU's order; it lands where the CUDA twin does, 1.4e-11 from the CPU |
| ADR-1437 | motion_hip, motion_v2_hip, psnr_hip, integer_ms_ssim_hip and cambi_hip are declared exact twins; float_psnr_hip and float_moment_hip are not |
| ADR-1438 | integer_ssim_hip adds its terms in the CPU's raster order at every frame size and returns the CPU's score bit for bit |
| ADR-1440 | float_psnr_hip adds its squared differences as integers and returns the CPU's score bit for bit |
| ADR-1441 | float_ssim_hip forms its window sums through the arithmetic float_ms_ssim_hip shares with the CPU and returns the CPU's score bit for bit |
| ADR-1442 | float ADM divides; the processor's reciprocal estimate leaves the reference and the CUDA twin |
| ADR-1443 | integer_ssim_sycl computes the CPU's fp64 term in 64-bit integers and adds in the CPU's order; it returns the CPU's ssim bit for bit |
| ADR-1444 | float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit |
| ADR-1445 | ssimulacra2_hip evaluates the CPU's fp64 terms and returns the sums of the CPU's loops |
| ADR-1446 | ssimulacra2_sycl forms the CPU's fp64 terms in 64-bit integers and returns the sums of the CPU's loops; it returns the CPU's score bit for bit |
| ADR-1447 | float_moment_hip adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1448 | ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair |
| ADR-1449 | float_moment_sycl adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1450 | float_psnr_sycl adds its squared differences as integers, and is bit-identical to the CPU |
| ADR-1451 | adm_sycl, motion_sycl, motion_v2_sycl, psnr_sycl, float_ssim_sycl and cambi_sycl are declared exact twins; speed_chroma_sycl is not |
| ADR-1452 | the gate bounds speed_chroma_hip against the CPU by what glibc's log2f adds, as it does for the CUDA twin |
| ADR-1453 | float_moment_cuda adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1454 | A large subtree AGENTS.md is a generated index over one page per topic |
| ADR-1455 | float_psnr_cuda adds its squared differences as integers, and is bit-identical to the CPU |
| ADR-1456 | vif_cuda keeps its device log2f(), proven equal to the CPU's log2 table on every entry, and is declared an exact twin |
| ADR-1457 | motion_cuda, motion_v2_cuda, psnr_cuda, float_ssim_cuda, float_ms_ssim_cuda and cambi_cuda are declared exact twins |
| ADR-1458 | float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit |
| ADR-1459 | SpEED's covariance kernels return the scalar kernel's bits; they vectorise across sums, not within one |
| ADR-1460 | speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell |
| ADR-1461 | no C or C++ translation unit is built with floating-point contraction; the strict policy is a project argument |
| ADR-1462 | vif_cuda reads the CPU's log2 table instead of evaluating log2f() on the device |
| ADR-1463 | float_ssim_sycl forms the CPU's fp64 terms in 64-bit integers and adds them on the host in the CPU's raster order |
| ADR-1464 | float_ssim_cuda adds its frame sums in the CPU's raster order, on the host |
| ADR-1465 | float_ms_ssim_cuda adds the terms of every scale in the CPU's raster order, on the host |
| ADR-1466 | float_ms_ssim_sycl stores every window's l, c and s of every scale and adds them on the host in the CPU's raster order |
| ADR-1467 | ciede.c writes its squares as products, so ciede2000 no longer depends on the compiler or on the C library's powf for them; a GCC build moves by up to 2e-11 |
| ADR-1468 | A SYCL kernel requires only a sub-group size every default AOT target accepts (16 or 32); the six kernels that required 8 move to 16 |
| ADR-1469 | the psnr_hvs SIMD butterfly is two functions, cut at the same statement in the AVX2 and the NEON twin |
| ADR-1470 | An enum that C and C++ translation units both see has one size: no C++-only underlying type other than int's width |
| ADR-1472 | The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale |
| ADR-1473 | The x86 float ADM wavelet and CSF kernels return the scalar bits and are dispatched; the two reduction kernels are removed |