vmafx-tune-go — Go port of vmaf-tune¶
vmafx-tune-go is the Go port of the vmaf-tune rate-quality tuning CLI. Every vmaf-tune subcommand is now ported: compare, ladder, report, recommend, predict, recommend-saliency, prefilter, tune-per-shot, fast, corpus, sidecar, benchmark, encode-profile and auto. It is the active tuning binary after the retirement of the Python CLI shadow.
A few individual flags still require Python — see Python-only flags. Each fails with a message naming the fallback rather than degrading quietly.
This page documents the Go binary. For the full Python vmaf-tune reference, see vmaf-tune.md.
Build¶
go build -o vmafx-tune-go ./cmd/vmafx-tune
# or with version injection:
go build -ldflags "-X main.version=$(cat VERSION)" -o vmafx-tune-go ./cmd/vmafx-tune
Ported subcommands (Stage 1)¶
compare — Rate-quality sweep¶
Runs a VMAF-target bisect for each (codec, target) pair and emits a ranked report.
Required flags:
| Flag | Description |
|---|---|
--reference, -r | Path to the reference video file |
Optional flags:
| Flag | Default | Description |
|---|---|---|
--codecs, -c | libx264,libx265 | Comma-separated encoder names. Stage-1 supports software encoders only (libx264, libx265). |
--targets, -t | 85 | Comma-separated VMAF target(s). Multiple targets produce a schema-v2 sweep. |
--output, -o | stdout | Output file path. |
--format | json | Output format: json or markdown. |
--ffmpeg | ffmpeg | Path to the ffmpeg binary. |
--vmaf | vmaf | Path to the vmaf binary for scoring. |
--work-dir | OS temp dir | Directory for temporary encode outputs. |
--crf-lo | 0 | Lower bound of CRF search window (best quality). |
--crf-hi | 0 | Upper bound (0 = encoder default: 51 for x264/x265). |
--max-iter | 12 | Maximum bisect iterations per (codec, target) pair. |
Example — single target, JSON output:
vmafx-tune-go compare \
--reference src.mp4 \
--codecs libx264,libx265 \
--targets 85 \
--output results.json
Example — multi-target sweep, Markdown:
vmafx-tune-go compare \
--reference src.mp4 \
--codecs libx264 \
--targets 80,85,90,95 \
--format markdown
Output schema¶
Single-target output (--targets has one value) uses schema-v1, identical to the Python vmaf-tune compare JSON output:
{ "src": "src.mp4", "target_vmaf": 85.0, "tool_version": "dev", "wall_time_ms": 4200, "rows": [
{
"codec": "libx264",
"best_crf": 23,
"bitrate_kbps": 1234.5,
"encode_time_ms": 2100,
"vmaf_score": 85.34,
"target_vmaf": 85.0,
"ok": true,
"error": "",
"bisect_samples": [...]
}
]
}
Multi-target output uses schema-v2 (adds schema_version: 2 and target_vmafs: [...]), also compatible with the Python report.py renderer.
Non-finite float values (NaN, Inf) are serialized as null — the output is RFC 8259 strict JSON, parseable by Go, Rust, and jq --strict.
Ported subcommands (Stage 4)¶
report — Render Markdown or HTML from prior runs¶
Reads one or more JSON files produced by compare or ladder and renders a human-readable report.
The subcommand auto-detects whether each input is a compare or ladder JSON file and renders a unified report.
Optional flags:
| Flag | Default | Description |
|---|---|---|
--output, -o | stdout | Output file path. |
--format | markdown | Output format: markdown or html. |
Example — Markdown to stdout:
Example — merge compare + ladder results into one HTML report:
The HTML output is self-contained (inlined CSS, no external dependencies) and renders correctly without a network connection.
Ported subcommands (ML-driven group)¶
These four subcommands cover the predictor-, saliency- and search-driven half of the CLI. Their JSON payloads are byte-identical to the Python originals — key order, ", " / ": " separators, NaN / Infinity tokens and float rendering all match json.dumps — so a downstream consumer cannot tell which binary produced a file.
recommend — pick the CRF meeting a target¶
Two modes. --from-corpus picks from an existing corpus JSONL with no new encodes; without it, a coarse-to-fine CRF search runs the encodes first, writes every visited point to --output, and picks from those rows.
Predicates (mutually exclusive):
| Flag | Description |
|---|---|
--target-vmaf | Smallest CRF whose VMAF meets the target. A smaller CRF is higher quality, so this is the best quality that clears the gate. Falls back to the closest miss, tagged (UNMET), when nothing clears it. |
--target-bitrate | Row whose bitrate_kbps is closest to the target; ties go to the lower CRF. --from-corpus only. |
Source and encode flags (required unless --from-corpus is used):
| Flag | Default | Description |
|---|---|---|
--source | — | Reference video; repeatable to sweep several sources. |
--width / --height | — | Raw-YUV reference geometry. |
--preset | — | Encoder preset; repeatable to sweep several presets. |
--encoder | libx264 | Codec adapter. |
--pix-fmt | yuv420p | ffmpeg pix_fmt. |
--framerate | 24 | Reference framerate. |
--duration | 0 | Clip duration in seconds; bounds the encode and derives achieved kbps. |
--output | corpus.jsonl | JSONL destination for the visited points. |
--encode-dir | .workingdir/cache/vmafx-tune/encodes | Bounded scratch directory for the probe encodes. |
--keep-encodes | off | Keep the encoded artefacts instead of deleting them after scoring. |
--no-source-hash | off | Skip the source SHA-256 (faster on very large sources). |
--vmaf-model | vmaf_v1.0.16_3d0h | libvmaf model version, or a path=... string. |
--score-backend | auto | libvmaf backend: auto, cpu, cuda, sycl, hip. |
--ffmpeg-bin / --vmaf-bin | ffmpeg / vmaf | Binary paths. |
Search flags:
| Flag | Default | Description |
|---|---|---|
--coarse-step | 10 | CRF step for the coarse pass (10 → 10, 20, 30, 40, 50). |
--fine-radius | 5 | Radius around the best-coarse CRF for the fine pass. |
--fine-step | 1 | CRF step for the fine pass. |
With the defaults that is 5 coarse plus up to 10 fine encodes, against 42 for a full 10–50 sweep (ADR-0296). The CRF window is fixed at 10–50, matching the Python defaults, so a Go and a Python run over the same source visit the same cells and their corpora stay comparable.
Uncertainty flags (ADR-0279):
| Flag | Default | Description |
|---|---|---|
--with-uncertainty | off | Consume the conformal prediction intervals carried in each row's vmaf_interval block. |
--uncertainty-sidecar | — | Calibration sidecar JSON. Without one the documented Research-0067 floor (tight 2.0, wide 5.0 VMAF) applies. |
A tight interval whose lower bound already clears the target short-circuits the search at that row; a wide interval refuses to short-circuit and tags the result (UNCERTAIN). This changes which encodes get probed, never which get shipped — the production-flip gate stays in the predictor's validation harness.
Examples:
# Pick from an existing corpus, machine-readable.
vmafx-tune-go recommend --from-corpus corpus.jsonl --target-vmaf 93 --json
# Run the search, then pick.
vmafx-tune-go recommend \
--source src.yuv --width 1920 --height 1080 --duration 10 \
--preset medium --target-vmaf 93 --output corpus.jsonl
predict — predict per-shot VMAF, then verify¶
Predicts VMAF per shot from cheap signals, then verifies the prediction against real libvmaf on K stratified shots and emits a verdict.
| Flag | Default | Description |
|---|---|---|
--source | — | Reference video, any FFmpeg-readable container (required). |
--codec | libx264 | Codec adapter. |
--target-vmaf | 93 | Target pooled-mean VMAF. |
--validate-k | 8 | Shots to verify against real libvmaf. |
--residual-threshold | 1.5 | Max abs(predicted - measured) before the verdict leaves gospel. |
--model | — | predictor_<codec>.onnx; without it the per-codec analytical curve runs. |
--use-saliency | off | Include saliency mean/variance in the feature vector. |
--saliency-model | shipped model | saliency_student_v1.onnx path. |
--per-shot-bin | vmaf-perShot | Shot detector binary. |
--bitdepth | 8 | Source bit depth (8, 10 or 12), forwarded to the detector. |
--total-frames | 0 | Frame count for the single-shot fallback. |
--report-out | stdout | Validation report destination. |
--with-uncertainty | off | Emit conformal intervals beside each predicted VMAF. |
--calibration-sidecar | — | Split-conformal calibration JSON. |
--alpha | sidecar's | Override the miscoverage level (0.05 = 95 % coverage). |
Verdicts and exit codes:
| Verdict | Meaning | Exit |
|---|---|---|
gospel | Every residual within the threshold; trust the predictor on the remaining shots. | 0 |
recalibrate | Residuals biased but tight; add the reported bias_correction and redo the picks. No retraining needed. | 0 |
fall_back | Residuals too wide; degrade to the full encode-and-score loop. | 2 |
Shot detection degrades to a single shot spanning the clip when vmaf-perShot is unavailable, with a WARN log line. Without --calibration-sidecar, --with-uncertainty emits degenerate low == high == point intervals and flags the report "calibrated": false — do not read a coverage guarantee into a zero-width interval.
vmafx-tune-go predict --source movie.mkv --codec libx264 \
--target-vmaf 93 --validate-k 8 --report-out predict.json
recommend-saliency — saliency-aware ROI encode¶
Scores the fork's saliency_student_v1 model over sampled frames, reduces the result to a per-block QP-offset map, and hands it to the encoder through its native ROI channel.
vmafx-tune-go recommend-saliency --src <yuv> --width W --height H \
--duration-frames N --output <path> [flags]
| Flag | Default | Description |
|---|---|---|
--saliency-aware | off | Enable the biasing; without it this is a plain encode. |
--saliency-offset | -4 | QP delta at peak saliency, clamped to ±12. Negative spends more bits on salient regions. |
--saliency-aggregator | mean | Temporal reducer: mean, ema, max or motion-weighted. |
--saliency-ema-alpha | 0.6 | Current-frame weight for ema. |
--saliency-model | shipped model | ONNX path. |
--saliency-fallback-plain | off | Accept a plain encode on an encoder with no ROI dispatch instead of exiting 2 (ADR-0546). |
--encoder | libx264 | Codec adapter. |
--crf | adapter default | Explicit CRF. |
--preset | medium | Encoder preset. |
ROI channel per encoder:
| Encoder | Channel | Granularity |
|---|---|---|
libx264 | -x264-params qpfile=… | 16×16 macroblocks |
libaom-av1 | -qpfile … (patched FFmpeg bridge) | 16×16 macroblocks |
libx265 | -x265-params zones=… | per-clip spatial mean |
libsvtav1 | -svtav1-params qp-file=… | 64×64 super-blocks |
libvvenc | -vvenc-params ROIFile=… | 64×64 CTUs |
Any other encoder exits 2 unless --saliency-fallback-plain is set or VMAFTUNE_SALIENCY_FALLBACK_OK=1 is exported.
An --output ending in .json is treated as a report destination: the encode goes to a sibling <stem>_encoded.mp4 and the path is echoed on stdout.
Inference availability. The Go port ships the entire numeric pipeline — YUV to ImageNet tensor, all four temporal aggregators, the QP mapping, the per-block reduce and every sidecar format — but has no in-process ONNX Runtime. With
--saliency-awareit therefore logs a warning and falls back to a plain encode, the same degradation the Python takes whenonnxruntimeis not installed. See Known gaps below.
prefilter — joint deband + CRF autotune¶
Optimises the ten frozen Pelorus deband knobs (the ADR-0110 control-plane contract) and the CRF axis together in one TPE study, with VMAF as the oracle.
| Flag | Default | Description |
|---|---|---|
--target-vmaf | — | Quality target on the VMAF [0, 100] scale (required). |
--smoke | off | Use the synthetic surface; no ffmpeg, Vulkan or GPU. |
--src | — | Source video; required for the live loop. |
--width / --height | — | Raw-YUV geometry; required for the live loop. |
--sweep-knob | all ten | Restrict the search to this knob; repeatable. |
--crf-min / --crf-max | 18 / 40 | Joint search range. |
--n-trials | 60 live, 40 smoke | TPE trial budget. |
--time-budget-s | 600 | Soft wall-clock cap. |
--seed | 0 | Sampler seed; the same seed reproduces the same recommendation. |
--encoder | libx264 | Codec performing the post-deband encode. |
--filter | pelorus_deband | Filter adapter to autotune. |
--encode-dir | .workingdir/cache/vmafx-tune/prefilter | Probe scratch directory. |
--output | stdout | JSON destination. |
The objective is |achieved - target| + λ·kbps, so the search converges on the lowest-bitrate combination that hits the target. The bitrate weight is small enough that it only breaks ties between equally-good quality points.
vmafx never runs the deband filter itself — it only emits the -vf string — so the live loop requires pelorus_deband_vulkan in the ffmpeg build and refuses to start (exit 2) with an actionable message when it is absent. --smoke exercises the whole search without it.
The ten swept knobs are range, thry, thrc, grainy, grainc, softness, detail, dither, dynamic and protect. The deliberately out-of-contract options (sample, blur, planes, meta) are pipeline switches set once per run and are rejected if passed to --sweep-knob.
# CI-friendly: no ffmpeg, no Vulkan, no GPU.
vmafx-tune-go prefilter --smoke --target-vmaf 93 --output rec.json
# Live loop.
vmafx-tune-go prefilter \
--src ref.yuv --width 1920 --height 1080 --duration 10 \
--target-vmaf 93 --encoder libx264 --output rec.json
Unlike the Python subcommand, this needs no optional [fast] extra: the TPE sampler is implemented natively (see Known gaps).
Known gaps¶
Two behaviours in this group are not at full parity. Both are documented here rather than hidden behind a silent degradation.
Saliency ONNX inference¶
recommend-saliency --saliency-aware falls back to a plain encode. Everything around the model is ported and tested against the Python — only the single forward pass is missing, because:
- Go has no in-process ONNX Runtime in this module. The fork's bridge (
pkg/ai, ADR-0713) shells out tovmafx-ort-runnerand passes the input tensor as a JSON array in argv. That works for the per-shot predictor's 14 floats; it cannot carry saliency's 3×H×W input, which is 6.2 million floats (about 75 MB of JSON) for a 1080p frame. - The runner itself is built here (
cmd/vmafx-ort-runner, ADR-1134 — see vmafx-ort-runner.md) and serves the predictor path; it has no transport for tensors that do not fit in argv.
Either of two changes unblocks it: a cgo ONNX Runtime binding (an ADR-level decision, since the binary currently builds without cgo), or a runner protocol that streams tensors over stdin — a protocol extension of the in-tree runner and pkg/ai, not a new dependency.
The per-shot predictor ONNX (predict --model) does route through pkg/ai, because its 14-float input fits argv comfortably; it degrades to the analytical curve when the runner is absent from PATH or linked against a libvmaf built without ONNX Runtime (exit 3), and the log says which.
TPE sampler trajectory¶
prefilter implements TPE natively (Bergstra et al. 2011 §4, the construction Optuna follows) rather than depending on a Go Optuna port that would pull gorm plus the MySQL, Postgres and cgo-SQLite drivers into a one-shot CLI. The search space, the objective, the emitted JSON and per-seed reproducibility are identical to the Python; the trial-by-trial trajectory for a given seed is not, and cannot be, because the two use different RNG streams.
fast — Proxy + TPE + GPU-verify recommend¶
Recommends a CRF for a VMAF target without running the full Phase A grid. Instead of a sweep, a TPE (Tree-structured Parzen Estimator) search walks the integer CRF axis: each trial encodes a short probe slice, extracts the canonical-6 libvmaf features, and predicts VMAF with the fr_regressor_v2 proxy. A single real encode + libvmaf score at the chosen CRF then verifies the recommendation — the proxy alone never wins (ADR-0304).
Port status.
--smokeruns end to end today. Production mode runs the backend selection, the probe encodes, the canonical-6 extraction and the verify pass, but stops at the proxy-inference step:fr_regressor_v2is a two-named-input ONNX graph and the Go inference seam drives a single flat input vector only. See Production-mode blocker below and usevmaf-tune fastfor a production run in the meantime.
Flags:
| Flag | Default | Description |
|---|---|---|
--src | — | Source video (raw YUV or any ffmpeg-readable container). Required unless --smoke. |
--width / --height | 0 | Raw-YUV reference geometry. Required in production mode. |
--pix-fmt | yuv420p | ffmpeg pixel format. |
--framerate | 24 | Reference framerate. |
--encoder | libx264 | Codec adapter. Must be in the proxy model's encoder vocabulary in production mode. |
--preset | medium | Encoder preset for the probe + verify encodes. |
--crf-min / --crf-max | 10 / 51 | Inclusive CRF search range. |
--n-trials | 30 prod / 50 smoke | TPE trial budget. |
--time-budget-s | 300 | Soft wall-clock cap on the TPE loop. In-flight trials are allowed to finish. |
--proxy-tolerance | 1.5 | Max absolute proxy/verify VMAF gap before the result is flagged out-of-distribution. |
--sample-chunk-seconds | 5.0 | Probe-encode slice length per trial. Shorter = faster trials, longer = more stable features. |
--smoke | false | Deterministic synthetic CRF→VMAF curve. No ffmpeg, no ONNX, no GPU verify. |
--score-backend | auto | libvmaf backend for the verify pass: auto, cpu, cuda, sycl, hip. auto walks cuda → sycl → hip → cpu; an explicit value is honoured strictly and errors rather than downgrading. |
--ffmpeg-bin | ffmpeg | Path to the ffmpeg binary. |
--vmaf-bin | vmaf | Path to the libvmaf CLI binary. |
--vmaf-model | vmaf_v1.0.16_3d0h | vmaf model version string. |
--encode-dir | .workingdir/cache/vmafx-tune/fast | Scratch dir for probe + verify encodes. |
--output, -o | stdout | JSON destination for the recommendation payload. |
--crf-max | — | See vmafx-tune-go fast --help. |
--height | — | See vmafx-tune-go fast --help. |
--target-vmaf | — | See vmafx-tune-go fast --help. |
Production-mode blocker: ONNX named inputs¶
The shipped proxy model/tiny/fr_regressor_v2.onnx declares two named input ports — features (shape [N, 6]) and codec (shape [N, 14]) — as recorded in its sidecar's "input_names". The only ONNX inference seam in the Go tree, pkg/ai.Registry.Infer, serialises a single flat []float64 to the vmafx-ort-runner subprocess and has no wire format for a second port; Registry.InferDirect, the CGO path, is an explicit Stage-2 stub.
Flattening the two ports into one 20-D vector is not a workaround — vmaftune/proxy.py documents that exact mistake: the graph's first dense layer reads the 6-D features port only, so the 14 codec dimensions are silently interpreted as batch padding and codec receives nothing. Rather than return a quietly-wrong score, pkg/fast fails with ErrProxyPortsUnsupported and a diagnostic naming both ports.
Any one of these unblocks it:
- A
vmafx-ort-runnerprotocol that accepts named input tensors, plus a matchingpkg/ai.Registry.InferNamed. The runner is in-tree since ADR-1134 (vmafx-ort-runner.md), so this is a protocol extension rather than an external dependency. - Promoting
pkg/ai.Registry.InferDirectonto a CGO ONNX Runtime binding (e.g.github.com/yalue/onnxruntime_go), whichpkg/aidefers to Stage 2 precisely because it couples the build tolibonnxruntime. - A single-port re-export of
fr_regressor_v2that concatenates the two inputs inside the graph, shipped alongside the current model.
Divergences from the Python fast implementation¶
The Go port fixes four defects found in vmaf-tune fast while reading it. The CLI surface and the JSON schema are unchanged; only the numbers the proxy would see differ.
| # | Python behaviour | Go behaviour |
|---|---|---|
| 1 | cli._build_fast_sample_extractor hands the probe .mp4 straight to the libvmaf CLI, which reads raw YUV only — every probe score fails and the feature vector degrades to six zeros. | Container-shaped encodes are decoded to raw YUV first (the score.maybe_decode_distorted step the Python probe leg skips), on both the probe and verify legs. |
| 2 | cli._parse_canonical6_means looks up bare adm2 / vif_scale0 keys in pooled_metrics; modern libvmaf emits integer_adm2 / integer_vif_scale0. score.py knows this and carries a mapping, but the fast path does not use it. | The integer_-prefixed key is tried first, then the bare key, then a per-frame average of either. |
| 3 | proxy.py's hardcoded ENCODER_VOCAB_V2 disagrees with the trainer and the shipped sidecar from index 3 on (libaom-av1 vs libvvenc), so the codec one-hot lands in the wrong slot for every codec past libsvtav1, and the model's own unknown catch-all is unreachable. | The vocabulary is read from the model sidecar's encoder_vocab, so it cannot drift from the installed checkpoint. Out-of-vocabulary codecs map to unknown when the model has that slot, and are a hard error otherwise. |
| 4 | The sidecar ships feature_mean / feature_std, and run_proxy documents that the caller must apply them — but no caller on the fast path does, so raw libvmaf means reach a model trained on standardised features. | The sidecar's StandardScaler is applied before inference. |
One behaviour is weaker in Go than in Python:
- TPE reproducibility. Optuna's
TPESampler(seed=0)makes a run bit-reproducible. The Go port usesgithub.com/c-bata/goptuna, whose TPE sampler honours its seed only partially:tpe.SamplerOptionSeedseeds the sampler's own RNG and its startup random sampler, butgoptuna/internal/random.ArgMaxMultinomialdraws from the process-globalmath/randsource, which Go seeds randomly at startup and whichrand.Seedcan no longer override. Repeat runs therefore explore slightly different trial sequences and may return a neighbouring CRF when two candidates score within about a VMAF point of each other. Measured on the ADR-0276 smoke curve: at the shipped budgets the recommendation stays within ±1 of the brute-force optimum in ~99–100 % of runs, and at 150 trials it hit the exact optimum in 150 of 150 runs for every target tested. Closing the gap needs an upstream goptuna change threading the sampler RNG intointernal/random.
Ported subcommands (Stage 5 — corpus + sidecar)¶
corpus — Phase A grid sweep¶
Sweeps a (preset, crf) grid against one or more references, encodes each cell, scores it against the reference with the libvmaf CLI, and writes one JSONL row per (source, preset, crf) combination.
The JSONL schema (v3) is the API contract the Phase B target-VMAF bisect and the Phase C per-title CRF predictor consume — see vmaf-tune.md for the column reference. When scoring with models that omit VIF (such as the default vmaf_v1.0.16_3d0h model, ADR-1168/1169), every Go libvmaf driver (corpus, fast, scorecli, tune/executor) automatically passes --feature vif to libvmaf and parses options-suffixed keys, ensuring canonical-6 columns (adm2, vif_scale0..3, motion2) are populated. The Go writer emits the same bytes the Python writer does, including the bare NaN tokens CPython's json module produces for columns libvmaf did not populate, so a corpus written by either binary is readable by the same trainers.
Source and encode flags:
| Flag | Default | Description |
|---|---|---|
--pix-fmt | yuv420p | ffmpeg pix_fmt of the reference. |
--framerate | 24 | Reference framerate. |
--duration | 0 | Reference duration in seconds. Bounds the encode and the bitrate calculation; 0 means the full source. |
--encoder | libx264 | Codec adapter. Any registered adapter is accepted — see vmaf-tune-codec-adapters.md. |
--crf | — | Quality value. Repeat for multiple cells. Required unless --coarse-to-fine derives the axis. |
--two-pass | off | Run a 2-pass encode for codecs that support it (libx264 / libx265). Adapters without true 2-pass emit a one-line stderr warning and run single-pass. |
--sample-clip-seconds | 0 | Encode and score only the centre N-second slice of each source. Encode time scales linearly with the slice; expect a 1–2 VMAF-point delta versus full-clip on diverse content. |
--encode-dir | .workingdir/cache/vmafx-tune/encodes | Bounded scratch directory for encodes. |
--keep-encodes | off | Retain encoded outputs after scoring and record their paths in encode_path. |
--no-source-hash | off | Skip src_sha256. Faster on huge YUVs; loses provenance. |
--source | — | Reference video. Repeat for multiple sources. |
--width | — | Rung target width in pixels. |
--height | — | Rung target height in pixels. |
--preset | — | Encoder preset. Repeat for multiple presets. |
Scoring flags:
| Flag | Default | Description |
|---|---|---|
--vmaf-model | vmaf_v1.0.16_3d0h | libvmaf model version string. |
--neg | off | Use the VMAF NEG (No Enhancement Gain) variant. Use for codec A-vs-B comparisons; not for production monitoring — see vmaf-neg.md. |
--score-backend | auto | libvmaf backend: auto, cpu, cuda, sycl, hip. auto picks the fastest available (cuda > sycl > hip > cpu); a specific name is honoured strictly and errors out when unavailable. |
--ffmpeg-bin | ffmpeg | Path to the ffmpeg binary. |
--vmaf-bin | vmaf | Path to the vmaf binary. |
--ffprobe-bin | ffprobe | Path to the ffprobe binary (used for HDR detection). |
--output | corpus.jsonl | JSONL output path. |
Search-mode flags:
| Flag | Default | Description |
|---|---|---|
--coarse-to-fine | off | Run a 2-pass coarse-then-fine CRF search instead of the full grid. With the defaults that is 15 encodes rather than 52 — see vmaf-tune-coarse-to-fine.md. |
--coarse-step | 10 | CRF step for the coarse pass. |
--fine-radius | 5 | ± radius around the best-coarse CRF for the fine pass. |
--fine-step | 1 | CRF step for the fine pass. |
--target-vmaf | unset | Target VMAF. The search refines around the smallest CRF whose score meets it; without a target it refines around the highest-VMAF coarse point. |
HDR flags (mutually exclusive — see vmaf-tune-hdr-and-sampling.md):
| Flag | Description |
|---|---|
--auto-hdr | (default) Probe each source with ffprobe and inject HDR codec args when PQ / HLG signalling is detected. |
--force-sdr | Treat every source as SDR; skip detection and flag injection. |
--force-hdr-pq | Treat every source as HDR PQ (SMPTE-2084) regardless of the probe. |
--force-hdr-hlg | Treat every source as HDR HLG (ARIB STD-B67) regardless of the probe. |
Example — full grid:
vmafx-tune-go corpus \
--source ref.yuv \
--width 1920 --height 1080 \
--framerate 24 --duration 10 \
--preset medium \
--crf 20 --crf 26 --crf 32 \
--output corpus.jsonl
Example — coarse-to-fine against a VMAF target:
vmafx-tune-go corpus \
--source ref.yuv \
--width 1920 --height 1080 \
--preset medium \
--coarse-to-fine --target-vmaf 93 \
--output corpus.jsonl
Rows stream to the output file as each cell completes, so an interrupted sweep leaves a usable partial corpus. The selected scoring backend is echoed on stderr (vmafx-tune: scoring backend = cpu) before the first encode, and the row count is echoed when the sweep finishes.
When --keep-encodes is off, cleanup is part of a successful cell. A missing temporary encode is harmless (an injected runner may already have removed it), but any other removal failure stops the sweep instead of silently leaking data. Likewise, when source hashing is enabled, an unreadable source fails before the first encode; use --no-source-hash only when provenance is deliberately not required.
Ported subcommands (encoder introspection)¶
benchmark — Rank encoders from an existing corpus¶
Answers the standard post-sweep question: which encoder hit the target quality at the lowest bitrate? It reads a Phase-A corpus JSONL written by vmaf-tune corpus and launches no ffmpeg and no libvmaf — the corpus stays the source of truth.
For every encoder in the corpus the report picks the lowest-bitrate row whose measured VMAF clears --target-vmaf. An encoder that never clears is reported with status unmet and its closest miss, so a missing encoder build is never mistaken for a quality result.
Required flags:
| Flag | Description |
|---|---|
--from-corpus | Phase-A corpus JSONL to benchmark |
Optional flags:
| Flag | Default | Description |
|---|---|---|
--target-vmaf | 92 | Matched-quality threshold each encoder must clear. |
--baseline-encoder | lowest-bitrate encoder that clears | Encoder used for the bitrate_delta_pct column. |
--format | markdown | Report format: markdown, json or csv. |
--output, -o | stdout | Report destination. Parent directories are created. |
Example — Markdown to stdout:
Example — CSV at a stricter target, with a pinned baseline:
vmafx-tune-go benchmark \
--from-corpus corpus.jsonl \
--target-vmaf 95 \
--baseline-encoder libx264 \
--format csv \
--output benchmark.csv
Rows are ranked with cleared encoders first (ascending bitrate), then the unmet ones. A row reports:
| Field | Meaning |
|---|---|
encoder | Encoder token from the corpus rows. |
status | ok when the encoder cleared the target, unmet otherwise. |
target_vmaf / margin | The requested threshold, and the selected row's VMAF minus it (negative when unmet). |
bitrate_kbps | Bitrate of the selected row. |
bitrate_delta_pct | Percentage difference against the baseline encoder; null / blank when no encoder cleared. |
rows / source_count / preset_count | How many eligible corpus rows the encoder contributed, and how many distinct sources / presets they span. |
encode_fps / score_fps | Means over the encoder's rows, counting positive finite samples only; null / blank when no row supplied timings. |
best | The selected corpus row (src, preset, crf, vmaf_score, bitrate_kbps, vmaf_model). |
Rows are excluded from the report when exit_status is non-zero, when vmaf_score or bitrate_kbps is missing or non-finite, or when the row carries no encoder name.
The CSV output uses CRLF line endings (Python's csv "excel" dialect), and the JSON output is stable pretty RFC 8259 with sorted keys — both byte-identical to vmaf-tune benchmark.
Ported subcommands (Stage 5)¶
auto — Phase F adaptive recipe-aware planner¶
Composes the per-phase tuning stages into one deterministic decision tree and emits a JSON plan. Optionally realises the winning cell as a real encode plus a libvmaf score.
Required flags:
| Flag | Description |
|---|---|
--target-vmaf | Quality target on the standard VMAF [0, 100] scale |
--src | Source video. Required unless --smoke. |
Optional flags:
| Flag | Default | Description |
|---|---|---|
--target-vmaf | 93 | Target pooled-mean VMAF. |
--max-budget-bitrate | 8000 | Upper bound on the picked rendition's bitrate, in kbps. |
--allow-codecs | libx264 | Comma-separated codec list the tree may pick from. A single entry short-circuits the compare-shortlist stage. |
--codec | (unset) | Pin the codec choice, overriding the --allow-codecs ranking. Also short-circuits the shortlist stage. |
--sample-clip-seconds | 0 | Propagate this clip length to internal sweeps rather than re-deciding per stage. 0 = full source. |
--smoke | false | Exercise the composition with synthetic metadata — no ffprobe, no ffmpeg, no ONNX. |
--output | stdout | Write the JSON plan here. |
--execute | false | After planning, run real FFmpeg encodes and libvmaf scores for the selected cell(s). |
--runs-dir | runs | Output directory for encoded files and tune_results.jsonl (used with --execute). |
--execute-all | false | With --execute: run every plan cell, not just the winner. Useful for post-hoc A/B comparison. |
--model | (unset) | Optional predictor_<codec>.onnx path. Default uses the analytical fallback curve. |
Exit codes (identical to vmaf-tune fast):
| Code | Meaning |
|---|---|
0 | Recommendation emitted; proxy and verify agree within --proxy-tolerance. |
2 | Usage or environment error (bad CRF range, missing --src, unavailable backend, proxy unavailable). |
3 | Recommendation emitted, but the proxy/verify gap exceeds tolerance. Fall back to the slow Phase A grid (ADR-0276). The payload is still written. |
Example — smoke run (works on any host):
vmafx-tune-go fast --smoke --target-vmaf 90
### `encode-profile` — Reproduce one recommendation from a report
Every vmaf-tune report embeds a machine-readable `encoder_profile` payload.
This subcommand reads that payload, selects one recommendation, and runs the
matching FFmpeg encode.
```text
vmafx-tune-go encode-profile --profile <FILE> --output <FILE> [flags]
The profile is accepted in any of the three shapes a report ships in:
| Input | How the payload is found |
|---|---|
Report JSON (.json) | Parsed directly; the encoder_profile block is unwrapped if present. |
Report HTML (.html / .htm) | Extracted from the raw-JSON <pre> block and HTML-unescaped. |
| Report Markdown (anything else) | Extracted from the fenced JSON payload. |
Selection defaults to the first Pareto-selected row with the lowest bitrate. --codec and --target-vmaf narrow the candidate set; --recommendation-index then picks the Nth survivor (zero-based, applied after filtering).
Required flags:
| Flag | Description |
|---|---|
--profile | Report JSON / HTML / Markdown containing encoder_profile |
--output, -o | Encoded output path |
Selection flags:
| Flag | Default | Description |
|---|---|---|
--codec | all | Restrict selection to one codec. |
--target-vmaf | all | Restrict selection to one target VMAF (matched with a 1e-6 absolute tolerance). |
--recommendation-index | 0 (first) | Zero-based index after the filters are applied. |
Override flags (each falls back to the profile value when omitted):
| Flag | Description |
|---|---|
--src | Override the source path stored in the profile. |
--preset | Override the stored / adapter-default preset. |
--pix-fmt | Override the raw-source pixel format (default yuv420p). |
--framerate, --width, --height | Override the raw-source geometry. |
--duration | Override the encode duration in seconds. Passing 0 explicitly suppresses the profile's own duration bound. |
--source-kind | auto (default), container or raw. Under auto, .yuv / .raw / .rgb / .gray are raw and everything else is a container. |
--sample-clip-seconds, --sample-clip-start-s | Input-side clip length / offset forwarded to FFmpeg. |
--extra-ffmpeg-arg | Append one raw FFmpeg argv token after the codec args; repeat as needed. Use --extra-ffmpeg-arg=-movflags for tokens starting with -. |
--ffmpeg-bin | Override the profile's ffmpeg_bin (default: profile value, then ffmpeg). |
--dry-run | Print the selected recommendation and the exact FFmpeg argv without encoding. |
Example — inspect the selection without encoding:
Example — reproduce the best x265 row at target 95:
vmafx-tune-go encode-profile \
--profile report.html \
--codec libx265 \
--target-vmaf 95 \
--output encoded.mkv
Output schema. The result is a single JSON document on stdout (sorted keys, two-space indent). A --dry-run emits ok, dry_run, profile, selected, ffmpeg_argv and output; a real run replaces dry_run with the encode outcome:
```json
{
"ok": true,
"profile": "report.json",
"selected": { "codec": "libx265", "crf": 28 },
"ffmpeg_argv": ["ffmpeg", "-y", "-hide_banner"],
"output": "encoded.mkv",
"exit_status": 0,
"encode_size_bytes": 148213,
"encode_time_ms": 1843.2,
"encoder_version": "libx265-3.5",
"ffmpeg_version": "n8.1",
"stderr_tail": "..."
}
Exit status. On a real run the process exit status is FFmpeg's own, so a failed encode surfaces the encoder's code (for example 254) rather than a generic 1. The JSON payload is still written first, so a wrapper script can read exit_status and stderr_tail regardless.
Exit codes¶
benchmark, encode-profile and the sidecar group use the same exit-status convention as the Python CLI they replace:
| Status | Meaning |
|---|---|
0 | Success. |
2 | A usage or validation failure — a missing/unknown flag, an unparseable flag value, a missing input file, a filter that matches no recommendation, a baseline encoder absent from the corpus, an unknown --codec or unreadable feature / capture file on sidecar. |
1 | sidecar only: the cache directory, host UUID or state.json could not be written (an uncaught OSError in Python). See Exit codes and diagnostics. |
| other | encode-profile only: FFmpeg's own exit status from a failed encode. |
Note that the earlier ports (compare, ladder, report) still report every failure as 1; that pre-existing inconsistency is tracked separately.
Hardware-encoder caveats.
- The emitted argv contains no
-init_hw_devicechain. FFmpeg's QSV bridge on Linux needs-init_hw_device vaapi=va:<node> -init_hw_device qsv=qsv_dev@va -filter_hw_device vabefore the first-i, plus aformat=nv12,hwuploadfilter, or the encode fails with-22(see ADR-0601). This matches the Python implementation exactly:vmaf-tuneinjects that chain in itscomparesweep, never inencode-profile. Supply the flags yourself with repeated--extra-ffmpeg-arg, or drive QSV throughcompare. - Hardware encoders (NVENC / QSV / AMF) reject sources below roughly 320x240. A profile built from a smaller clip will fail at the encoder.
av1_videotoolboxis a placeholder: upstream FFmpeg ships no such encoder, so the adapter refuses to emit an argv shape it cannot verify (ADR-0339).
--model is a Go-side addition: the Python auto driver always constructs its predictor without a model path, which is the analytical fallback this flag defaults to. Supplying a model routes inference through the ONNX bridge in pkg/ai, which degrades back to the analytical curve when the ORT runner is not on PATH.
Example — plan only:
Example — plan and realise the winner:
vmafx-tune-go auto \
--src src.mp4 \
--target-vmaf 95 \
--max-budget-bitrate 6000 \
--execute --runs-dir runs/
How the plan is built¶
- Probe the source: geometry and duration via
ffprobe, HDR signalling via the colour-metadata classifier. Every probe failure degrades to conservative defaults (1920x1080, duration 0, SDR) rather than aborting the run. - Apply the content recipe. The content class selects a small override set — a narrower or wider conformal-confidence gate, a forced single-rung ladder, a saliency intensity, and a target-VMAF offset. An HDR source carrying only a generic
live_actionlabel is promoted to the HDR recipe. The overrides load fromai/data/phase_f_recipes_calibrated.jsonwhen that file is reachable, and fall back to the documented placeholders with a one-line warning otherwise. - Walk the ten short-circuits, recording each one that fires under
metadata.short_circuits. - Estimate each
(rung, codec)cell: invert the predictor for a CRF, then estimate the VMAF and bitrate that CRF would produce. - Pick a winner against the VMAF target and the bitrate budget.
Short-circuits¶
Each predicate names a stage the tree can skip. They are evaluated in this order, and the order is part of the output contract.
| Name | Fires when |
|---|---|
ladder-single-rung | Source height is below 2160, or the recipe forces a single rung. |
codec-pinned | --codec is set, or --allow-codecs resolves to one entry. |
predictor-gospel | The predictor verdict is GOSPEL; trust its CRF and skip the coarse-to-fine fallback. |
skip-saliency | Content class is neither animation nor screen_content. |
sdr-skip | The source carries no HDR signalling. |
sample-clip-propagate | --sample-clip-seconds is positive; propagate it verbatim to internal sweeps. |
skip-per-shot | The source is both shorter than 5 minutes and below 0.15 shot variance. |
low-complexity | The probe-encode bitrate is under 200 kbps. Dormant when no probe has run. |
baseline-meets-target | A default-CRF encode already meets the target. Dormant when no baseline was scored. |
no-two-pass | The resolved codec adapter does not support two-pass encoding. |
Confidence-aware escalation¶
Each cell carries a conformal interval width, and that width decides whether the predictor's own verdict is overridden:
| Interval width | Decision |
|---|---|
<= tight (default 2.0) | skip-escalation — trust the point estimate even on a FALL_BACK verdict. |
>= wide (default 5.0) | force-escalation — escalate even on a GOSPEL verdict. |
| between | Defer to the native verdict. |
NaN (uncalibrated) | Defer to the native verdict. |
Without a calibration sidecar the interval is uncalibrated, so cells carry NaN and no override happens. The 2.0 / 5.0 defaults are an emergency floor, not a corpus fit.
Plan JSON¶
The plan is a {"cells": [...], "metadata": {...}} object with sorted keys, byte-compatible with the Python vmaf-tune auto output.
Each cell carries: rung, codec, verdict, crf, estimated_vmaf, estimated_bitrate_kbps, hdr_args, sample_clip_seconds, confidence_decision, interval_width, effective_predictor_target_vmaf, prediction_source, saliency_intensity, and selected.
metadata.winner.status is one of:
| Status | Meaning |
|---|---|
budget_and_quality_met | A cell satisfies both the target and the budget. |
quality_met_budget_exceeded | Quality is reachable, but every such cell is over budget; the smallest overage wins. |
target_unmet | No cell reaches the target; the closest miss is returned so you get a concrete next encode. |
no_eligible_cells | No cell carried finite estimates. |
The plan JSON is not strict RFC 8259. An uncalibrated
interval_widthis emitted as the bare tokenNaN, exactly as CPython'sjson.dumpsdoes with its defaultallow_nan=True. This is deliberate byte-compatibility with the Python emitter. Parse it with Python'sjsonmodule or another permissive parser;jq --strictand Go'sencoding/jsonwill reject it. The--executeresults log (tune_results.jsonl) is strict — non-finite values there are rendered asnull.
Execute mode¶
With --execute the selected cell is encoded and scored, and one row per executed cell is appended to <runs-dir>/tune_results.jsonl. Encoded files land beside it as encode_<index>_<codec>_<preset>_crf<n>.mkv.
A failed encode is recorded in its row with a non-zero encode_exit_status, and scoring is skipped for that cell. The command exits non-zero only when cells were executed and none scored successfully.
sidecar — Local on-host predictor sidecar¶
Trains and inspects a bias-correction term on top of the shipped predictor:
The shipped predictor is never mutated, so model upgrades stay deterministic and reproducible across hosts.
Flags shared by every nested subcommand:
| Flag | Default | Description |
|---|---|---|
--codec | libx264 | Codec bucket for the sidecar state. Must be one of the 19 registered codec names; an unknown name is a usage error whose message lists them. |
--cache-dir | ${XDG_CACHE_HOME:-~/.cache}/vmaf-tune/sidecar | Sidecar cache root. A relative path stays relative in state_path. |
--predictor-version | predictor_v1 | Predictor-version namespace. |
--model | (unset) | Optional ONNX predictor — see the ONNX note. Unset uses the analytical fallback. |
--json | false | Emit machine-readable JSON instead of the one-line text form. |
Nested subcommands:
| Subcommand | Extra flags | Purpose |
|---|---|---|
status | — | Print state metadata: codec, host UUID, state path, predictor version, update count, residual RMS. |
predict | --features-json, --crf | Predict VMAF with the correction folded in. Reports the base score, the correction, and the sum. |
record | --features-json, --crf, --observed-vmaf, --no-persist | Fold one observed encode result into the fit. |
batch-record | --captures-jsonl | Fold a JSONL capture file, one observation per row, persisting once at the end. |
Example — inspect, train from a capture log, then predict:
vmafx-tune-go sidecar status --json
vmafx-tune-go sidecar batch-record --captures-jsonl captures.jsonl --json
vmafx-tune-go sidecar predict --features-json shot.json --crf 26 --json
sidecar status¶
Prints the state metadata for one codec bucket: the anonymous host UUID, the on-disk state path, the predictor-version namespace, the number of folded-in captures, and the RMS of the buffered residuals (the drift signal).
{
"codec": "libx264",
"host_uuid": "0123456789abcdef0123456789abcdef",
"n_updates": 0,
"predictor_version": "predictor_v1",
"recent_residual_rms": 0.0,
"schema": "vmaf-tune-sidecar-status/v1",
"schema_version": 1,
"state_path": "/home/u/.cache/vmaf-tune/sidecar/predictor_v1/libx264/state.json"
}
Without --json the same fields print as one line:
codec=libx264 predictor_version=predictor_v1 updates=0 residual_rms=0.000000 state=/home/u/.cache/vmaf-tune/sidecar/predictor_v1/libx264/state.json
sidecar predict¶
Predicts VMAF for one shot at one CRF with the correction applied.
| Flag | Description |
|---|---|
--features-json | (required) Path to a JSON object carrying the shot's feature values. |
--crf | (required) CRF to predict at. |
The payload (schema vmaf-tune-sidecar-predict/v1) carries base_vmaf (the bare predictor), correction, and sidecar_vmaf (the sum, clamped to [0, 100]), plus codec, crf and n_updates. The text form is base=<b> correction=<c> sidecar=<s> updates=<n> with six decimals.
sidecar record¶
Folds one observed VMAF measurement into the ridge fit.
| Flag | Description |
|---|---|
--features-json | (required) Path to the shot's feature JSON. |
--crf | (required) CRF the observation was measured at. |
--observed-vmaf | (required) Observed libvmaf score for the encode. |
--no-persist | Update in memory only; mainly useful for tests. |
vmafx-tune-go sidecar record \
--features-json shot.json \
--crf 26 \
--observed-vmaf 91.75 \
--json
The residual is computed against the bare predictor, never against the sidecar-corrected value, so repeated captures converge rather than compounding. The payload (schema vmaf-tune-sidecar-record/v1) is the status payload plus crf, observed_vmaf, base_vmaf and residual (observed_vmaf - base_vmaf); the text form is recorded updates=<n> residual=<r> state=<path>.
sidecar batch-record¶
Folds a whole JSONL capture file into the fit, one observation per line.
| Flag | Description |
|---|---|
--captures-jsonl | (required) Path to the JSONL capture file. |
Each line is a JSON object carrying the feature fields (either at the top level or nested under a features key) plus crf and observed_vmaf. Lines are split on \n, \r\n or a lone \r (CPython's universal-newline rule) with no length limit; blank lines are ignored but still numbered. A malformed row — not a JSON object, a missing or non-numeric required field, a crf string that is not an integer literal — is reported as vmafx-tune sidecar batch-record: skip line <n>: <reason> on stderr, counted in rows_skipped, and the run still succeeds. That is deliberate: a capture log is often partially corrupt after an interrupted run, and losing the good rows to one bad line would be worse. State is written once at the end, and only when at least one row was recorded. The payload (schema vmaf-tune-sidecar-batch-record/v1) is the status payload plus rows_recorded and rows_skipped; the text form is recorded=<rows> skipped=<rows> updates=<n> state=<path>.
Feature JSON¶
--features-json takes an object of shot features, or a {"features": {...}} wrapper so a capture row can carry crf and observed_vmaf alongside. The four probe_* fields are required — a zero probe bitrate would train the fit on a fabricated complexity barometer — and everything else defaults to 0. Values may be JSON numbers or numeric strings ("2400"), as CPython's float() accepts.
| Key | Required | Meaning |
|---|---|---|
probe_bitrate_kbps | yes | Average bitrate over the probe encode. |
probe_i_frame_avg_bytes | yes | Mean I-frame size. |
probe_p_frame_avg_bytes | yes | Mean P-frame size. |
probe_b_frame_avg_bytes | yes | Mean B-frame size (0 for codecs without B-frames). |
saliency_mean, saliency_var | no | Saliency signals; 0 when unavailable. |
frame_diff_mean, y_avg, y_var | no | FFmpeg signalstats aggregates. |
shot_length_frames, fps, width, height | no | Structural metadata. |
{
"probe_bitrate_kbps": 4200.5,
"probe_i_frame_avg_bytes": 51234.0,
"probe_p_frame_avg_bytes": 8123.25,
"probe_b_frame_avg_bytes": 2011.75,
"saliency_mean": 0.42,
"saliency_var": 0.031,
"frame_diff_mean": 7.5,
"y_avg": 112.25,
"y_var": 1830.5,
"shot_length_frames": 240,
"fps": 24.0,
"width": 1920,
"height": 1080
}
State and privacy¶
State lives at:
${XDG_CACHE_HOME:-~/.cache}/vmaf-tune/sidecar/
host-uuid # random 128-bit token
<predictor-version>/<codec>/state.json # ridge weights + inverse Gram
The host UUID is drawn from a CSPRNG on first use. It is never derived from a MAC address, hostname, /etc/machine-id, CPUID, or any other machine-identifying signal.
A predictor-version or schema mismatch on load discards the fit and resets to cold start, keeping only the host UUID. That is what makes a shipped-model upgrade safe: a stale correction can never be replayed against a refreshed predictor. At cold start the weights are zero, so the correction is exactly 0.0 and the sidecar returns the bare predictor's value untouched.
A corrupt state.json also cold-starts — including one whose weights or a_inv carry null, the residue of a NaN capture — and the corrupt file is left in place so you can inspect it.
Exit codes and diagnostics¶
The four subcommands follow the Python vmaf-tune sidecar exit contract:
| Status | Meaning |
|---|---|
0 | Success — including a batch-record run in which every row was skipped. |
1 | An I/O failure the Python CLI does not catch either: the cache directory cannot be created, the host UUID or state.json cannot be written, or stdout cannot be written. |
2 | A usage or validation failure: an unknown or unparseable flag, a missing required flag, an unknown --codec, an unresolvable --model, a --features-json that cannot be read / is not a JSON object / lacks a required key, or a --captures-jsonl that cannot be read. |
Diagnostics go to stderr and are not byte-identical to Python's: cobra prefixes them with Error: where Python prints vmaf-tune sidecar <cmd>:, the batch-record skip lines carry the vmafx-tune prefix, and the reason text is Go's rather than CPython's exception message. Nothing is written to stdout on failure.
Byte compatibility with vmaf-tune sidecar¶
For the same inputs, every stdout payload (JSON and text form) and every byte of state.json written by the Go binary is identical to the Python CLI's: the analytical predictor, the Sherman–Morrison update, and the json.dumps(..., indent=2, sort_keys=True) rendering are reproduced to the last bit, and the atomic write leaves the same single state.json behind (no .tmp residue). cmd/vmafx-tune/cmd/testdata/sidecar/ holds fixtures dumped from the Python CLI by regen.sh (pinned host UUID, relative --cache-dir), and TestSidecarPythonParity replays the same 23-step operator sequence — cold status, three records, two batch-record loads, warm status and predict, a --no-persist record, a second codec bucket, and nine error paths — requiring identical stdout, identical state.json snapshots and identical exit statuses. Known, deliberate differences:
- A capture row containing the non-standard JSON tokens
NaN/Infinityis skipped by the Go binary and counted inrows_skipped. CPython'sjson.loadsaccepts the tokens: aNaNfeature then silently skips the update while still counting the row as recorded, and aNaNobserved_vmafpoisons the ridge weights (they persist asnulland the next load cold-starts). - Hand-edited state files outside the shape either writer produces are not guaranteed to load identically (for example a numeric string where CPython's
float()would coerce and Go's decoder will not). --modelis resolved differently — see the ONNX note.
ONNX predictor models¶
The two binaries interpret --model differently, and the Python behaviour itself depends on the host:
- Python takes a filesystem path to
predictor_<codec>.onnx. Withonnxruntimeimportable, a missing file raises and the CLI exits2; without it, the flag is silently ignored and the analytical curve is used (exit0). - Go resolves the value through the model registry (
pkg/ai): an absolute path that exists, else<model-dir>/<name>.onnx, else<model-dir>/<name>, where the model dir is$VMAFX_MODEL_DIRor/usr/local/share/vmafx/model. An unresolvable name exits2. Inference then goes through thevmafx-ort-runnersubprocess, which this repository does not build; when the runner is absent fromPATHthe predictor falls back to the analytical curve silently — a warning is logged only when the runner is present and inference fails.
Omit --model to use the analytical fallback — the default, and the only path the parity fixtures exercise — and the two binaries agree byte for byte.
Ported subcommands (per-shot tuning)¶
tune-per-shot — Per-shot CRF tuning¶
Cuts the source into shots, bisects a CRF against the VMAF target inside each shot, and emits an FFmpeg encoding plan: one command per segment plus the concat-demuxer command that stitches them into the final file.
The plan is emitted, not executed. Segment encodes are independent, so you can run them sequentially or in parallel, then run the concat command.
Pipeline:
- Shot detection — shells out to the fork's
vmaf-perShotbinary (vmaf-perShot.md), which wraps TransNet V2. A missing or failing binary degrades to one shot spanning the clip. - Uniform-window splitter — any shot longer than
--max-shot-durationis sliced into equal sub-shots, so an under-cutting detector (fades, short clips) still yields a usable timeline. - Per-shot bisect — each shot is extracted to raw YUV and run through the CRF bisect against
--target-vmaf. - Plan emission — the recommendations become segment + concat commands.
Required flags:
| Flag | Description |
|---|---|
--src | Reference video: raw YUV, or any FFmpeg-readable container |
Source geometry:
| Flag | Default | Description |
|---|---|---|
--width | auto-probed | Source width. Required for raw YUV (.yuv / .raw); auto-probed via ffprobe for containers. |
--height | auto-probed | Source height. Same rule as --width. |
--pix-fmt | yuv420p | Source pixel format. |
--framerate | auto-probed | Source framerate. Falls back to 24.0 when the probe yields nothing. |
--bitdepth | 8 | Source YUV bit depth: 8, 10 or 12. |
--total-frames | 0 | Frame count for the single-shot fallback when vmaf-perShot is unavailable. |
Shot detection:
| Flag | Default | Description |
|---|---|---|
--per-shot-bin | vmaf-perShot | Path to the shot-detector binary. |
--scene-threshold | detector default (12.0) | Override the detector's mean-absolute-luma-delta cut threshold. Lower yields more shots. |
--max-shot-duration | 2.0 | Uniform-window splitter, in seconds. 0 disables it. |
Tuning:
| Flag | Default | Description |
|---|---|---|
--target-vmaf | 92 | Target pooled-mean VMAF per shot. |
--encoder | libx264 | Codec adapter (see the codec table below). |
--preset | codec default (medium) | Preset for the bisect encodes. |
--crf-min / --crf-max | codec absolute window | Inclusive bisect search bounds. Pass both or neither. |
--max-iterations | 8 | Maximum encode+score rounds per shot. |
--vmaf-model | vmaf_v1.0.16_3d0h | Model passed to the vmaf binary. |
--neg | off | Route the model to its NEG variant. There is no NEG counterpart to any vmaf_v1.0.16_* model, so --neg also selects the v0.6.1 generation (vmaf_v0.6.1neg). See vmaf-neg.md. |
--score-backend | auto | libvmaf backend: auto, cpu, cuda, sycl, hip. An explicit backend that the host cannot provide fails fast rather than silently downgrading. |
--vmaf-bin / --ffmpeg-bin | vmaf / ffmpeg | Binary paths. |
--workdir | $VMAFTUNE_WORKDIR or OS temp | Scratch space for encode / decode artefacts. Raw YUV decodes are large — point this at a volume with room. |
--max-concurrent-decodes | 1 | Concurrent reference-YUV decodes. 1 is safest on space-constrained volumes. |
Output:
| Flag | Default | Description |
|---|---|---|
--plan-out | stdout | Destination for the JSON plan. |
--output | per_shot_encode.mp4 | Final concatenated encode path named inside the plan. |
--segment-dir | <output dir>/segments | Directory the segment commands write into. |
--script-out | — | Also write the plan as a copy-paste shell script. |
Example:
vmafx-tune-go tune-per-shot \
--src src.mp4 \
--target-vmaf 92 \
--encoder libx264 \
--plan-out plan.json \
--segment-dir ./segments \
--output final.mp4
Plan JSON schema¶
Byte-compatible with vmaf-tune tune-per-shot: keys are sorted, floats keep Python's repr() form, and a shot whose predicate never measured a bitrate carries null rather than NaN.
{
"concat_command": ["ffmpeg", "-y", "-hide_banner", "-f", "concat", "..."],
"encoder": "libx264",
"framerate": 24.0,
"predicate": "bisect",
"segment_commands": [["ffmpeg", "-y", "-hide_banner", "-ss", "0.000000", "..."]],
"shots": [
{
"bitrate_kbps": 1234.57,
"crf": 22,
"end_frame": 48,
"predicted_vmaf": 92.5,
"start_frame": 0
}
],
"target_vmaf": 92.0
}
start_frame is inclusive and end_frame exclusive (half-open). The vmaf-perShot sidecar uses an inclusive end frame; the tuner normalises it.
The concat listing (concat.txt) is written next to --plan-out when --segment-dir is not given, matching the Python. Pass --segment-dir explicitly to pin the listing and the segment commands to the same directory; the command logs a WARN when the two diverge.
Supported codecs¶
tune-per-shot accepts the ten codecs the Go encoder registry can construct. Each emits its own quality knob in the plan:
| Codec | Plan argv shape |
|---|---|
libx264, libx265 | -c:v NAME -preset medium -crf N |
libsvtav1 | -c:v libsvtav1 -preset 7 -crf N (integer preset) |
libaom-av1 | -c:v libaom-av1 -cpu-used 4 -crf N |
h264_nvenc, hevc_nvenc | -c:v NAME -preset p4 -cq N |
h264_qsv, hevc_qsv | -c:v NAME -preset medium -global_quality N |
h264_amf, hevc_amf | -c:v NAME -quality balanced -rc cqp -qp_i N -qp_p N |
The Python registry carries seven more adapters — av1_nvenc, av1_qsv, av1_amf, the four VideoToolbox encoders, libvvenc and libvpx-vp9. They have no Go encoder implementation yet, so --encoder rejects them by name and points at the Python binary.
Flags with no Go implementation¶
Both fail fast with an actionable message rather than being accepted and silently ignored:
| Flag | Why | Use instead |
|---|---|---|
--predicate-module MODULE:CALLABLE | Loads a Python callable at runtime; Go has no runtime import. The Go equivalent is the pershot.PredicateFn seam, available to library callers. | vmaf-tune tune-per-shot --predicate-module ... |
--fast-nr | NR early-elimination runs the nr_metric_v1 ONNX model through onnxruntime; the Go binary has no ONNX runtime binding. | vmaf-tune tune-per-shot --fast-nr |
Python-only flags¶
Every vmaf-tune subcommand is ported. A few individual flags still need the Python implementation, because they depend on in-process ONNX inference or on importing a Python callable at runtime:
| Flag | Subcommand | Why | Use instead |
|---|---|---|---|
--fast-nr | tune-per-shot | NR early-elimination needs an ONNX forward pass per bisect midpoint | vmaf-tune tune-per-shot --fast-nr |
--predicate-module | tune-per-shot | Imports an arbitrary Python MODULE:CALLABLE at runtime | vmaf-tune tune-per-shot --predicate-module |
--saliency-aware | recommend-saliency | Requires a saliency ONNX forward pass | vmaf-tune recommend-saliency --saliency-aware |
recommend-saliency --saliency-aware and predict --use-saliency are accepted by the Go binary. When the saliency session cannot be built, recommend-saliency proceeds without an ROI map (the report's saliency_aware field then reads false), and predict logs a warning and degrades saliency moments to 0.0, matching the Python behavior.
--model (on predict, sidecar, auto) routes inference through the vmafx-ort-runner subprocess (vmafx-ort-runner.md), which the dev container and the Go CI job build from cmd/vmafx-ort-runner (ADR-1134). When the runner is absent from PATH, or present but linked against a libvmaf built without ONNX Runtime (exit 3), the predictor logs a warning carrying the runner's stderr and falls back to the analytical curve — the same fallback the Python takes without onnxruntime, but reported rather than silent. On sidecar an unresolvable model name is a usage error (exit 2); see the ONNX note for how the two binaries resolve the flag differently.
recommend's encode-driven path writes the same schema-v3 corpus JSONL the corpus subcommand does, and every key is present. Five corpus features are not carried by this group's port and their fields hold the same zero / empty values the Python emits when the feature is unavailable, so a reader filters on them exactly as it already does: the content-addressed encode cache (ADR-0298), HDR detection (ADR-0295), TransNet-V2 shot metadata (shot_count stays 0), sample-clip windowing (clip_mode stays full), and the encoder-internal pass-1 stats (the ten enc_internal_* columns stay 0.0, which is what the Python aggregator returns for an empty frame list). Those belong to the corpus port.
Configuration and logging¶
vmafx-tune-go runs each subcommand inside the golusoris clikit (cobra + fx) framework. The framework injects a structured *slog.Logger and a config tree into every subcommand, so run diagnostics (sweep start/finish, ladder build summary, report rendering) are emitted as structured log lines on stderr, while subcommand output (JSON / Markdown / HTML) goes to stdout or the --output file. This keeps machine-readable output separate from logs when you pipe stdout.
Configuration is read from environment variables under the VMAFX_ prefix. golusoris maps each underscore in the variable name to a config-path delimiter (VMAFX_LOG_LEVEL → log.level):
| Environment variable | Config key | Effect | Default |
|---|---|---|---|
VMAFX_LOG_LEVEL | log.level | Minimum log level: debug, info, warn, error | info |
VMAFX_LOG_FORMAT | log.format | Log handler: auto (tint on a TTY, JSON otherwise), tint, json | auto |
Examples:
# Quiet the per-run INFO diagnostics; keep warnings/errors.
VMAFX_LOG_LEVEL=warn vmafx-tune-go report results.json
# Force JSON logs for machine ingestion regardless of TTY.
VMAFX_LOG_FORMAT=json vmafx-tune-go compare --reference src.mp4 --targets 90
Migration roadmap¶
| Stage | Scope | ADR | Status |
|---|---|---|---|
| Stage 1 | compare subcommand, libx264/libx265, single/multi-target bisect | ADR-0705 | Merged |
| Stage 2 | ladder subcommand, hardware encoders (NVENC, QSV, AMF), convex hull + knee selection | ADR-0730 | Merged |
| Stage 3 | Downscale plumbing | — | Merged |
| Stage 4 | report subcommand, Markdown + HTML rendering | ADR-0770 | Merged |
| golusoris | Migrate the CLI root + subcommands onto the golusoris clikit (cobra + fx) framework; VMAFX_-prefixed config + injected slog | ADR-1119 | Merged |
| ML-driven | recommend, predict, recommend-saliency, prefilter; codec-adapter registry, encode/score drivers, predictor, saliency pipeline, native TPE | — | This PR |
| Encoder introspection | benchmark + encode-profile subcommands; pkg/benchmark, pkg/codecadapter, pkg/encodeprofile, pkg/pyjson | ADR-0770 | This PR | | Stage 5 | tune-per-shot subcommand, conformal CLI wiring | Planned | — |
| golusoris | Migrate the CLI root + subcommands onto the golusoris clikit (cobra + fx) framework; VMAFX_-prefixed config + injected slog | ADR-1119 | This PR | | Stage 5 (corpus/sidecar) | corpus + sidecar subcommands; pkg/codecadapter, pkg/corpus, pkg/pyjson (corpus); pkg/tune/predictor, pkg/tune/sidecar, pkg/tune/pyjson (sidecar) | ADR-1125 | Merged (#1153) | | Stage 5 (per-shot) | tune-per-shot subcommand, conformal CLI wiring | Planned | — | | Stage 6 | fast subcommand (requires ONNX Go binding) | Planned | — |
| Stage 5 | tune-per-shot subcommand: pkg/pershot, pkg/scorebackend, codec-adapter table, raw-YUV scorer | ADR-0705 | This PR | | Stage 5b | conformal CLI wiring | Planned | — | | Stage 6 | fast subcommand + tune-per-shot --fast-nr (both require an ONNX Go binding) | Planned | — |
| Stage 6 | fast subcommand + pkg/conformal + pkg/scorebackend; smoke path complete, production path blocked on ONNX named inputs | ADR-0276 / ADR-0304 | This PR | | Stage 5 | tune-per-shot subcommand, conformal CLI wiring | Planned | — | | Stage 6b | fast production mode (needs a named-input ONNX seam — see Production-mode blocker) | Planned | — |
| golusoris | Migrate the CLI root + subcommands onto the golusoris clikit (cobra + fx) framework; VMAFX_-prefixed config + injected slog | ADR-1119 | This PR | | Stage 5 | auto (Phase F planner + execute mode) and sidecar subcommands | ADR-1125 | Merged (#1153) | | Stage 6 | tune-per-shot subcommand, conformal CLI wiring | Planned | — | | Stage 7 | fast subcommand (requires ONNX Go binding) | Planned | — | | Stage N | Feature parity; rename binary to vmafx-tune | Planned | — |
Correction. An earlier revision of this table listed
pkg/conformalas merged under Stage 3. It was not: no such package existed onmaster. The package landed with Stage 6 above, and the roadmap row has been corrected.
Architecture¶
The CLI root and every subcommand are built with the golusoris clikit (cobra + fx) framework (ADR-1119): clikit.New builds the root, clikit.Command builds each subcommand, and a thin withGolusoris adapter boots a one-shot fx graph per invocation so the command receives an injected *slog.Logger and config, runs to completion, and propagates its error as the process exit code.
The Go binary uses an adapter pattern with these core packages:
pkg/encoder/—Encoderinterface + software and hardware encoder implementations. Each encoder shells out toffmpeg; nolibavcodecCGo dependency. Stage 5 adds the codec-adapter policy table (Adapter: preset vocabulary, quality windows, per-codec argv shape),AdapterEncoder, andProbeSource.
Each encoder shells out to ffmpeg; no libavcodec CGo dependency. EncodeParams.InputArgs carries ffmpeg input-side options so raw-YUV sources (-f rawvideo -pix_fmt -s -r) and sample clips (input-side -ss / -t, for fast-seek) work; EncodeParams.OutputPath pins a deterministic destination and EncodeResult.OutputSizeBytes reports the encode size for size-over-duration bitrate maths. - pkg/bisect/ — Stateless Run(src, enc, scoreFunc, params) function. The score function is injectable, enabling unit tests without a live vmaf binary. Stage 5 adds YUVScoreFunc, which decodes a containerised distorted file to raw YUV and invokes vmaf with full geometry / model / backend flags. - pkg/ladder/ — Build(src, encoder, Params) function. Convex hull (upperConvexHull), knee selection (selectRenditions), min-bitrate-gap filter. - pkg/pershot/ — Stage-5 shot detection, uniform-window splitter, per-shot tuning, and encoding-plan construction + JSON emission. - pkg/scorebackend/ — Stage-5 libvmaf backend resolution: parses the vmaf --help backend line, probes each vendor independently, and honours an explicit --score-backend strictly. - pkg/report/ — Stage-1: EmitJSON / EmitMarkdown renderers (single-run emit). Stage-4: RenderMarkdownMulti / RenderHTMLMulti (multi-file report rendering). - pkg/fast/ — the fast path: Recommend (the flow), RunTPE (the goptuna-backed search), NewSamplePredictor / NewVerifier (the probe and verify pipelines), and ORTProxy (the fr_regressor_v2 seam). - pkg/scorebackend/ — libvmaf backend detection and strict selection, ported from the selection half of vmaftune/score_backend.py. Detect intersects what the local vmaf --help advertises with what nvidia-smi / sycl-ls / rocminfo report; Select honours auto via a fallback chain and never silently downgrades an explicit request. - pkg/conformal/ — distribution-free prediction intervals for the VMAF predictor (split conformal and CV+ / jackknife+), ported from vmaftune/conformal.py. The JSON sidecar is byte-compatible with the Python writer, so a calibration produced by either implementation loads in the other. Not yet wired into a CLI flag — that is Stage 5.
The ML-driven group adds:
pkg/codecadapter/— the nineteen-codec registry (quality windows, preset mapping onto each encoder's native axis, probe argv, two-pass argv).pkg/ffencode/— the encode driver the tuning subcommands share: raw-YUV geometry, named presets, injected extra params, sample-clip windowing. Distinct frompkg/encoder, which models the narrower Stage-1 bisect abstraction.pkg/scorecli/— the libvmaf CLI driver with explicit geometry, the backend selector and the canonical-6 pooled aggregates.pkg/predictor/—ShotFeatures, the per-codec analytical curve, the binary-search CRF inversion, feature extraction and the validation harness.pkg/pershot/— shot detection viavmaf-perShotwith the single-shot fallback and the uniform long-shot splitter.pkg/saliency/— the full saliency ROI pipeline: YUV to ImageNet tensor, four temporal aggregators, QP mapping, per-block reduce, five sidecar formats.pkg/prefilter/— the frozen Pelorus knob contract plus a native TPE sampler.pkg/recommend/,pkg/conformal/,pkg/uncertainty/,pkg/corpusrow/— the predicate pickers, split-conformal intervals, confidence bands and the schema-v3 corpus row.pkg/pyjson/— renders payloads the way Python'sjson.dumpsdoes, so the emitted JSON is byte-identical to the Python binary's. It is the one CPython-JSON encoder in the tree (ADR-1137); the formerinternal/pyjson,internal/pyjsonstrictandpkg/tune/pyjsoncopies were folded into it.
The corpus and sidecar subcommands add five more:
pkg/codecadapter/— the ADR-0237 adapter registry: every codec's quality knob, preset vocabulary, validation rule, and ffmpeg argv slice. The encode driver never branches on codec identity.pkg/corpus/— the Phase A orchestrator: encode + score drivers, the pass-1 encoder-stats parser, HDR detection and codec-arg dispatch, shot metadata, scoring-backend selection, the JSONL reader / writer, and the coarse-to-fine search.pkg/predictor/—ShotFeaturesplus the per-codec analytical VMAF curve and its CRF inversion.pkg/tune/sidecar/— the online-ridge bias-correction model (Sherman-Morrison rank-1 updates) and its cache-dir persistence. The port produced two implementations of this; the other (pkg/sidecar/) was removed because nothing imported it. Thesidecarsubcommand pairs it withpkg/tune/predictor/(the analytical curve, withlog10routed throughpkg/tune/pymathfor libm parity) andpkg/tune/pyjson/— not thepkg/predictorcopy thatpredictuses or thepkg/pyjsoncopy thatcorpususes; ADR-1125 records which consumer owns which copy.pkg/pyjson/— a CPython-compatible JSON encoder. The corpus JSONL and the sidecar--jsonpayloads are cross-implementation artefacts, so the writer reproducesjson.dumpsbyte-for-byte: bareNaN/Infinitytokens,repr()-style float rendering, andensure_asciiescaping.pkg/corpusalso ports CPython's Neumaier-compensatedsum()andstatistics.pstdev()so the aggregate columns match to the last bit.
The encoder-introspection subcommands add four more:
pkg/benchmark/— corpus loading (LoadCorpusJSONL), per-encoder summarisation (Summarize) and the three renderers. No subprocess at all.pkg/codecadapter/— the argv-shaping half ofvmaftune.codec_adapters: 19 codecs, each answering "what is the-c:v ...slice for this (preset, quality)?" so nothing else branches on codec identity.pkg/encodeprofile/— profile loading from JSON / HTML / Markdown, recommendation selection,EncodeRequestconstruction, FFmpeg argv composition and the encode driver (with an injectableRunnerseam so tests never spawn ffmpeg).pkg/pyjson/— the same encoder, here rendering Go value trees byte-identically to CPython'sjson.dumps(..., indent=2, sort_keys=True)andjsonio.dumps_strict. Go'sencoding/jsondiffers on key ordering, HTML escaping, non-ASCII escaping and float formatting (float64(92)renders92in Go and92.0in CPython), so a shared encoder keeps the ported payloads diff-clean against the Python originals.
Stage 5 adds the auto / sidecar stack under pkg/tune/:
pkg/tune/auto/— the Phase F decision tree: source probing, the ten short-circuit predicates, the recipe table, the confidence policy, winner selection, and the plan emitter.pkg/tune/sidecar/— the online-ridge bias-correction model, its Sherman-Morrison rank-1 update, and the cache-dir persistence layout.pkg/tune/executor/—--executemode: the libvmaf CLI driver and the JSONL results log; its ffmpeg argv ispkg/ffencode's under the executor's name.pkg/predictor/,pkg/codecadapter/,pkg/hdr/,pkg/pyjson/,pkg/pymath/— the shared layersautoandsidecarconsume:ShotFeaturesand the per-codec analytical curve with the optional ONNX session and thePickCRFbinary-search inversion; the codec-adapter registry (quality windows, probe knobs, preset vocabularies, per-encoder ffmpeg argv); HDR detection from ffprobe colour metadata plus the per-codec HDR flag dispatch; the CPython-compatible JSON emitter; and the correctly-roundedExp2/Log10kernels that keepestimated_bitrate_kbpsandestimated_vmafon the platform-libm value CPython emits (Go'smath.Pow/math.Log10land a ULP away; the package docs record the measured residual). Each has exactly one implementation (ADR-1137); thepkg/tune/{predictor,codec,pyjson}paths survive only as thin aliases until the in-flight sidecar parity fix (#1187) lands, after which the sidecar imports move and the aliases are deleted.
See ADR-0705 for the migration rationale, ADR-0730 for Stage-2, ADR-0770 for Stage-4, and ADR-0702 for the Phase 4 umbrella.