Skip to content

vmafx-tune-go — Go port of vmaf-tune

vmafx-tune-go is the Go port of the vmaf-tune rate-quality tuning CLI. Every vmaf-tune subcommand is now ported: compare, ladder, report, recommend, predict, recommend-saliency, prefilter, tune-per-shot, fast, corpus, sidecar, benchmark, encode-profile and auto. It is the active tuning binary after the retirement of the Python CLI shadow.

A few individual flags still require Python — see Python-only flags. Each fails with a message naming the fallback rather than degrading quietly.

This page documents the Go binary. For the full Python vmaf-tune reference, see vmaf-tune.md.

Build

go build -o vmafx-tune-go ./cmd/vmafx-tune
# or with version injection:
go build -ldflags "-X main.version=$(cat VERSION)" -o vmafx-tune-go ./cmd/vmafx-tune

Ported subcommands (Stage 1)

compare — Rate-quality sweep

Runs a VMAF-target bisect for each (codec, target) pair and emits a ranked report.

vmafx-tune-go compare [flags]

Required flags:

Flag Description
--reference, -r Path to the reference video file

Optional flags:

Flag Default Description
--codecs, -c libx264,libx265 Comma-separated encoder names. Stage-1 supports software encoders only (libx264, libx265).
--targets, -t 85 Comma-separated VMAF target(s). Multiple targets produce a schema-v2 sweep.
--output, -o stdout Output file path.
--format json Output format: json or markdown.
--ffmpeg ffmpeg Path to the ffmpeg binary.
--vmaf vmaf Path to the vmaf binary for scoring.
--work-dir OS temp dir Directory for temporary encode outputs.
--crf-lo 0 Lower bound of CRF search window (best quality).
--crf-hi 0 Upper bound (0 = encoder default: 51 for x264/x265).
--max-iter 12 Maximum bisect iterations per (codec, target) pair.

Example — single target, JSON output:

vmafx-tune-go compare \
  --reference src.mp4 \
  --codecs libx264,libx265 \
  --targets 85 \
  --output results.json

Example — multi-target sweep, Markdown:

vmafx-tune-go compare \
  --reference src.mp4 \
  --codecs libx264 \
  --targets 80,85,90,95 \
  --format markdown

Output schema

Single-target output (--targets has one value) uses schema-v1, identical to the Python vmaf-tune compare JSON output:

{ "src": "src.mp4", "target_vmaf": 85.0, "tool_version": "dev", "wall_time_ms": 4200, "rows": [

    {
      "codec": "libx264",
      "best_crf": 23,
      "bitrate_kbps": 1234.5,
      "encode_time_ms": 2100,
      "vmaf_score": 85.34,
      "target_vmaf": 85.0,
      "ok": true,
      "error": "",
      "bisect_samples": [...]
    }
  ]
}

Multi-target output uses schema-v2 (adds schema_version: 2 and target_vmafs: [...]), also compatible with the Python report.py renderer.

Non-finite float values (NaN, Inf) are serialized as null — the output is RFC 8259 strict JSON, parseable by Go, Rust, and jq --strict.

Ported subcommands (Stage 4)

report — Render Markdown or HTML from prior runs

Reads one or more JSON files produced by compare or ladder and renders a human-readable report.

vmafx-tune-go report <input.json> [input2.json ...] [flags]

The subcommand auto-detects whether each input is a compare or ladder JSON file and renders a unified report.

Optional flags:

Flag Default Description
--output, -o stdout Output file path.
--format markdown Output format: markdown or html.

Example — Markdown to stdout:

vmafx-tune-go report results.json

Example — merge compare + ladder results into one HTML report:

vmafx-tune-go report compare.json ladder.json \
  --format html \
  --output report.html

The HTML output is self-contained (inlined CSS, no external dependencies) and renders correctly without a network connection.

Ported subcommands (ML-driven group)

These four subcommands cover the predictor-, saliency- and search-driven half of the CLI. Their JSON payloads are byte-identical to the Python originals — key order, ", " / ": " separators, NaN / Infinity tokens and float rendering all match json.dumps — so a downstream consumer cannot tell which binary produced a file.

recommend — pick the CRF meeting a target

Two modes. --from-corpus picks from an existing corpus JSONL with no new encodes; without it, a coarse-to-fine CRF search runs the encodes first, writes every visited point to --output, and picks from those rows.

vmafx-tune-go recommend [flags]

Predicates (mutually exclusive):

Flag Description
--target-vmaf Smallest CRF whose VMAF meets the target. A smaller CRF is higher quality, so this is the best quality that clears the gate. Falls back to the closest miss, tagged (UNMET), when nothing clears it.
--target-bitrate Row whose bitrate_kbps is closest to the target; ties go to the lower CRF. --from-corpus only.

Source and encode flags (required unless --from-corpus is used):

Flag Default Description
--source — Reference video; repeatable to sweep several sources.
--width / --height — Raw-YUV reference geometry.
--preset — Encoder preset; repeatable to sweep several presets.
--encoder libx264 Codec adapter.
--pix-fmt yuv420p ffmpeg pix_fmt.
--framerate 24 Reference framerate.
--duration 0 Clip duration in seconds; bounds the encode and derives achieved kbps.
--output corpus.jsonl JSONL destination for the visited points.
--encode-dir .workingdir/cache/vmafx-tune/encodes Bounded scratch directory for the probe encodes.
--keep-encodes off Keep the encoded artefacts instead of deleting them after scoring.
--no-source-hash off Skip the source SHA-256 (faster on very large sources).
--vmaf-model vmaf_v1.0.16_3d0h libvmaf model version, or a path=... string.
--score-backend auto libvmaf backend: auto, cpu, cuda, sycl, hip.
--ffmpeg-bin / --vmaf-bin ffmpeg / vmaf Binary paths.

Search flags:

Flag Default Description
--coarse-step 10 CRF step for the coarse pass (10 → 10, 20, 30, 40, 50).
--fine-radius 5 Radius around the best-coarse CRF for the fine pass.
--fine-step 1 CRF step for the fine pass.

With the defaults that is 5 coarse plus up to 10 fine encodes, against 42 for a full 10–50 sweep (ADR-0296). The CRF window is fixed at 10–50, matching the Python defaults, so a Go and a Python run over the same source visit the same cells and their corpora stay comparable.

Uncertainty flags (ADR-0279):

Flag Default Description
--with-uncertainty off Consume the conformal prediction intervals carried in each row's vmaf_interval block.
--uncertainty-sidecar — Calibration sidecar JSON. Without one the documented Research-0067 floor (tight 2.0, wide 5.0 VMAF) applies.

A tight interval whose lower bound already clears the target short-circuits the search at that row; a wide interval refuses to short-circuit and tags the result (UNCERTAIN). This changes which encodes get probed, never which get shipped — the production-flip gate stays in the predictor's validation harness.

Examples:

# Pick from an existing corpus, machine-readable.
vmafx-tune-go recommend --from-corpus corpus.jsonl --target-vmaf 93 --json

# Run the search, then pick.
vmafx-tune-go recommend \
  --source src.yuv --width 1920 --height 1080 --duration 10 \
  --preset medium --target-vmaf 93 --output corpus.jsonl

predict — predict per-shot VMAF, then verify

Predicts VMAF per shot from cheap signals, then verifies the prediction against real libvmaf on K stratified shots and emits a verdict.

vmafx-tune-go predict --source <video> [flags]
Flag Default Description
--source — Reference video, any FFmpeg-readable container (required).
--codec libx264 Codec adapter.
--target-vmaf 93 Target pooled-mean VMAF.
--validate-k 8 Shots to verify against real libvmaf.
--residual-threshold 1.5 Max abs(predicted - measured) before the verdict leaves gospel.
--model — predictor_<codec>.onnx; without it the per-codec analytical curve runs.
--use-saliency off Include saliency mean/variance in the feature vector.
--saliency-model shipped model saliency_student_v1.onnx path.
--per-shot-bin vmaf-perShot Shot detector binary.
--bitdepth 8 Source bit depth (8, 10 or 12), forwarded to the detector.
--total-frames 0 Frame count for the single-shot fallback.
--report-out stdout Validation report destination.
--with-uncertainty off Emit conformal intervals beside each predicted VMAF.
--calibration-sidecar — Split-conformal calibration JSON.
--alpha sidecar's Override the miscoverage level (0.05 = 95 % coverage).

Verdicts and exit codes:

Verdict Meaning Exit
gospel Every residual within the threshold; trust the predictor on the remaining shots. 0
recalibrate Residuals biased but tight; add the reported bias_correction and redo the picks. No retraining needed. 0
fall_back Residuals too wide; degrade to the full encode-and-score loop. 2

Shot detection degrades to a single shot spanning the clip when vmaf-perShot is unavailable, with a WARN log line. Without --calibration-sidecar, --with-uncertainty emits degenerate low == high == point intervals and flags the report "calibrated": false — do not read a coverage guarantee into a zero-width interval.

vmafx-tune-go predict --source movie.mkv --codec libx264 \
  --target-vmaf 93 --validate-k 8 --report-out predict.json

recommend-saliency — saliency-aware ROI encode

Scores the fork's saliency_student_v1 model over sampled frames, reduces the result to a per-block QP-offset map, and hands it to the encoder through its native ROI channel.

vmafx-tune-go recommend-saliency --src <yuv> --width W --height H \
  --duration-frames N --output <path> [flags]
Flag Default Description
--saliency-aware off Enable the biasing; without it this is a plain encode.
--saliency-offset -4 QP delta at peak saliency, clamped to ±12. Negative spends more bits on salient regions.
--saliency-aggregator mean Temporal reducer: mean, ema, max or motion-weighted.
--saliency-ema-alpha 0.6 Current-frame weight for ema.
--saliency-model shipped model ONNX path.
--saliency-fallback-plain off Accept a plain encode on an encoder with no ROI dispatch instead of exiting 2 (ADR-0546).
--encoder libx264 Codec adapter.
--crf adapter default Explicit CRF.
--preset medium Encoder preset.

ROI channel per encoder:

Encoder Channel Granularity
libx264 -x264-params qpfile=… 16×16 macroblocks
libaom-av1 -qpfile … (patched FFmpeg bridge) 16×16 macroblocks
libx265 -x265-params zones=… per-clip spatial mean
libsvtav1 -svtav1-params qp-file=… 64×64 super-blocks
libvvenc -vvenc-params ROIFile=… 64×64 CTUs

Any other encoder exits 2 unless --saliency-fallback-plain is set or VMAFTUNE_SALIENCY_FALLBACK_OK=1 is exported.

An --output ending in .json is treated as a report destination: the encode goes to a sibling <stem>_encoded.mp4 and the path is echoed on stdout.

Inference availability. The Go port ships the entire numeric pipeline — YUV to ImageNet tensor, all four temporal aggregators, the QP mapping, the per-block reduce and every sidecar format — but has no in-process ONNX Runtime. With --saliency-aware it therefore logs a warning and falls back to a plain encode, the same degradation the Python takes when onnxruntime is not installed. See Known gaps below.

prefilter — joint deband + CRF autotune

Optimises the ten frozen Pelorus deband knobs (the ADR-0110 control-plane contract) and the CRF axis together in one TPE study, with VMAF as the oracle.

vmafx-tune-go prefilter --target-vmaf T [flags]
Flag Default Description
--target-vmaf — Quality target on the VMAF [0, 100] scale (required).
--smoke off Use the synthetic surface; no ffmpeg, Vulkan or GPU.
--src — Source video; required for the live loop.
--width / --height — Raw-YUV geometry; required for the live loop.
--sweep-knob all ten Restrict the search to this knob; repeatable.
--crf-min / --crf-max 18 / 40 Joint search range.
--n-trials 60 live, 40 smoke TPE trial budget.
--time-budget-s 600 Soft wall-clock cap.
--seed 0 Sampler seed; the same seed reproduces the same recommendation.
--encoder libx264 Codec performing the post-deband encode.
--filter pelorus_deband Filter adapter to autotune.
--encode-dir .workingdir/cache/vmafx-tune/prefilter Probe scratch directory.
--output stdout JSON destination.

The objective is |achieved - target| + λ·kbps, so the search converges on the lowest-bitrate combination that hits the target. The bitrate weight is small enough that it only breaks ties between equally-good quality points.

vmafx never runs the deband filter itself — it only emits the -vf string — so the live loop requires pelorus_deband_vulkan in the ffmpeg build and refuses to start (exit 2) with an actionable message when it is absent. --smoke exercises the whole search without it.

The ten swept knobs are range, thry, thrc, grainy, grainc, softness, detail, dither, dynamic and protect. The deliberately out-of-contract options (sample, blur, planes, meta) are pipeline switches set once per run and are rejected if passed to --sweep-knob.

# CI-friendly: no ffmpeg, no Vulkan, no GPU.
vmafx-tune-go prefilter --smoke --target-vmaf 93 --output rec.json

# Live loop.
vmafx-tune-go prefilter \
  --src ref.yuv --width 1920 --height 1080 --duration 10 \
  --target-vmaf 93 --encoder libx264 --output rec.json

Unlike the Python subcommand, this needs no optional [fast] extra: the TPE sampler is implemented natively (see Known gaps).

Known gaps

Two behaviours in this group are not at full parity. Both are documented here rather than hidden behind a silent degradation.

Saliency ONNX inference

recommend-saliency --saliency-aware falls back to a plain encode. Everything around the model is ported and tested against the Python — only the single forward pass is missing, because:

  • Go has no in-process ONNX Runtime in this module. The fork's bridge (pkg/ai, ADR-0713) shells out to vmafx-ort-runner and passes the input tensor as a JSON array in argv. That works for the per-shot predictor's 14 floats; it cannot carry saliency's 3×H×W input, which is 6.2 million floats (about 75 MB of JSON) for a 1080p frame.
  • The runner itself is built here (cmd/vmafx-ort-runner, ADR-1134 — see vmafx-ort-runner.md) and serves the predictor path; it has no transport for tensors that do not fit in argv.

Either of two changes unblocks it: a cgo ONNX Runtime binding (an ADR-level decision, since the binary currently builds without cgo), or a runner protocol that streams tensors over stdin — a protocol extension of the in-tree runner and pkg/ai, not a new dependency.

The per-shot predictor ONNX (predict --model) does route through pkg/ai, because its 14-float input fits argv comfortably; it degrades to the analytical curve when the runner is absent from PATH or linked against a libvmaf built without ONNX Runtime (exit 3), and the log says which.

TPE sampler trajectory

prefilter implements TPE natively (Bergstra et al. 2011 §4, the construction Optuna follows) rather than depending on a Go Optuna port that would pull gorm plus the MySQL, Postgres and cgo-SQLite drivers into a one-shot CLI. The search space, the objective, the emitted JSON and per-seed reproducibility are identical to the Python; the trial-by-trial trajectory for a given seed is not, and cannot be, because the two use different RNG streams.

fast — Proxy + TPE + GPU-verify recommend

Recommends a CRF for a VMAF target without running the full Phase A grid. Instead of a sweep, a TPE (Tree-structured Parzen Estimator) search walks the integer CRF axis: each trial encodes a short probe slice, extracts the canonical-6 libvmaf features, and predicts VMAF with the fr_regressor_v2 proxy. A single real encode + libvmaf score at the chosen CRF then verifies the recommendation — the proxy alone never wins (ADR-0304).

vmafx-tune-go fast --target-vmaf <N> [--smoke | --src <file> --width W --height H] [flags]

Port status. --smoke runs end to end today. Production mode runs the backend selection, the probe encodes, the canonical-6 extraction and the verify pass, but stops at the proxy-inference step: fr_regressor_v2 is a two-named-input ONNX graph and the Go inference seam drives a single flat input vector only. See Production-mode blocker below and use vmaf-tune fast for a production run in the meantime.

Flags:

Flag Default Description
--src — Source video (raw YUV or any ffmpeg-readable container). Required unless --smoke.
--width / --height 0 Raw-YUV reference geometry. Required in production mode.
--pix-fmt yuv420p ffmpeg pixel format.
--framerate 24 Reference framerate.
--encoder libx264 Codec adapter. Must be in the proxy model's encoder vocabulary in production mode.
--preset medium Encoder preset for the probe + verify encodes.
--crf-min / --crf-max 10 / 51 Inclusive CRF search range.
--n-trials 30 prod / 50 smoke TPE trial budget.
--time-budget-s 300 Soft wall-clock cap on the TPE loop. In-flight trials are allowed to finish.
--proxy-tolerance 1.5 Max absolute proxy/verify VMAF gap before the result is flagged out-of-distribution.
--sample-chunk-seconds 5.0 Probe-encode slice length per trial. Shorter = faster trials, longer = more stable features.
--smoke false Deterministic synthetic CRF→VMAF curve. No ffmpeg, no ONNX, no GPU verify.
--score-backend auto libvmaf backend for the verify pass: auto, cpu, cuda, sycl, hip. auto walks cuda → sycl → hip → cpu; an explicit value is honoured strictly and errors rather than downgrading.
--ffmpeg-bin ffmpeg Path to the ffmpeg binary.
--vmaf-bin vmaf Path to the libvmaf CLI binary.
--vmaf-model vmaf_v1.0.16_3d0h vmaf model version string.
--encode-dir .workingdir/cache/vmafx-tune/fast Scratch dir for probe + verify encodes.
--output, -o stdout JSON destination for the recommendation payload.
--crf-max — See vmafx-tune-go fast --help.
--height — See vmafx-tune-go fast --help.
--target-vmaf — See vmafx-tune-go fast --help.

Production-mode blocker: ONNX named inputs

The shipped proxy model/tiny/fr_regressor_v2.onnx declares two named input ports — features (shape [N, 6]) and codec (shape [N, 14]) — as recorded in its sidecar's "input_names". The only ONNX inference seam in the Go tree, pkg/ai.Registry.Infer, serialises a single flat []float64 to the vmafx-ort-runner subprocess and has no wire format for a second port; Registry.InferDirect, the CGO path, is an explicit Stage-2 stub.

Flattening the two ports into one 20-D vector is not a workaround — vmaftune/proxy.py documents that exact mistake: the graph's first dense layer reads the 6-D features port only, so the 14 codec dimensions are silently interpreted as batch padding and codec receives nothing. Rather than return a quietly-wrong score, pkg/fast fails with ErrProxyPortsUnsupported and a diagnostic naming both ports.

Any one of these unblocks it:

  1. A vmafx-ort-runner protocol that accepts named input tensors, plus a matching pkg/ai.Registry.InferNamed. The runner is in-tree since ADR-1134 (vmafx-ort-runner.md), so this is a protocol extension rather than an external dependency.
  2. Promoting pkg/ai.Registry.InferDirect onto a CGO ONNX Runtime binding (e.g. github.com/yalue/onnxruntime_go), which pkg/ai defers to Stage 2 precisely because it couples the build to libonnxruntime.
  3. A single-port re-export of fr_regressor_v2 that concatenates the two inputs inside the graph, shipped alongside the current model.

Divergences from the Python fast implementation

The Go port fixes four defects found in vmaf-tune fast while reading it. The CLI surface and the JSON schema are unchanged; only the numbers the proxy would see differ.

# Python behaviour Go behaviour
1 cli._build_fast_sample_extractor hands the probe .mp4 straight to the libvmaf CLI, which reads raw YUV only — every probe score fails and the feature vector degrades to six zeros. Container-shaped encodes are decoded to raw YUV first (the score.maybe_decode_distorted step the Python probe leg skips), on both the probe and verify legs.
2 cli._parse_canonical6_means looks up bare adm2 / vif_scale0 keys in pooled_metrics; modern libvmaf emits integer_adm2 / integer_vif_scale0. score.py knows this and carries a mapping, but the fast path does not use it. The integer_-prefixed key is tried first, then the bare key, then a per-frame average of either.
3 proxy.py's hardcoded ENCODER_VOCAB_V2 disagrees with the trainer and the shipped sidecar from index 3 on (libaom-av1 vs libvvenc), so the codec one-hot lands in the wrong slot for every codec past libsvtav1, and the model's own unknown catch-all is unreachable. The vocabulary is read from the model sidecar's encoder_vocab, so it cannot drift from the installed checkpoint. Out-of-vocabulary codecs map to unknown when the model has that slot, and are a hard error otherwise.
4 The sidecar ships feature_mean / feature_std, and run_proxy documents that the caller must apply them — but no caller on the fast path does, so raw libvmaf means reach a model trained on standardised features. The sidecar's StandardScaler is applied before inference.

One behaviour is weaker in Go than in Python:

  • TPE reproducibility. Optuna's TPESampler(seed=0) makes a run bit-reproducible. The Go port uses github.com/c-bata/goptuna, whose TPE sampler honours its seed only partially: tpe.SamplerOptionSeed seeds the sampler's own RNG and its startup random sampler, but goptuna/internal/random.ArgMaxMultinomial draws from the process-global math/rand source, which Go seeds randomly at startup and which rand.Seed can no longer override. Repeat runs therefore explore slightly different trial sequences and may return a neighbouring CRF when two candidates score within about a VMAF point of each other. Measured on the ADR-0276 smoke curve: at the shipped budgets the recommendation stays within ±1 of the brute-force optimum in ~99–100 % of runs, and at 150 trials it hit the exact optimum in 150 of 150 runs for every target tested. Closing the gap needs an upstream goptuna change threading the sampler RNG into internal/random.

Ported subcommands (Stage 5 — corpus + sidecar)

corpus — Phase A grid sweep

Sweeps a (preset, crf) grid against one or more references, encodes each cell, scores it against the reference with the libvmaf CLI, and writes one JSONL row per (source, preset, crf) combination.

vmafx-tune-go corpus [flags]

The JSONL schema (v3) is the API contract the Phase B target-VMAF bisect and the Phase C per-title CRF predictor consume — see vmaf-tune.md for the column reference. When scoring with models that omit VIF (such as the default vmaf_v1.0.16_3d0h model, ADR-1168/1169), every Go libvmaf driver (corpus, fast, scorecli, tune/executor) automatically passes --feature vif to libvmaf and parses options-suffixed keys, ensuring canonical-6 columns (adm2, vif_scale0..3, motion2) are populated. The Go writer emits the same bytes the Python writer does, including the bare NaN tokens CPython's json module produces for columns libvmaf did not populate, so a corpus written by either binary is readable by the same trainers.

Source and encode flags:

Flag Default Description
--pix-fmt yuv420p ffmpeg pix_fmt of the reference.
--framerate 24 Reference framerate.
--duration 0 Reference duration in seconds. Bounds the encode and the bitrate calculation; 0 means the full source.
--encoder libx264 Codec adapter. Any registered adapter is accepted — see vmaf-tune-codec-adapters.md.
--crf — Quality value. Repeat for multiple cells. Required unless --coarse-to-fine derives the axis.
--two-pass off Run a 2-pass encode for codecs that support it (libx264 / libx265). Adapters without true 2-pass emit a one-line stderr warning and run single-pass.
--sample-clip-seconds 0 Encode and score only the centre N-second slice of each source. Encode time scales linearly with the slice; expect a 1–2 VMAF-point delta versus full-clip on diverse content.
--encode-dir .workingdir/cache/vmafx-tune/encodes Bounded scratch directory for encodes.
--keep-encodes off Retain encoded outputs after scoring and record their paths in encode_path.
--no-source-hash off Skip src_sha256. Faster on huge YUVs; loses provenance.
--source — Reference video. Repeat for multiple sources.
--width — Rung target width in pixels.
--height — Rung target height in pixels.
--preset — Encoder preset. Repeat for multiple presets.

Scoring flags:

Flag Default Description
--vmaf-model vmaf_v1.0.16_3d0h libvmaf model version string.
--neg off Use the VMAF NEG (No Enhancement Gain) variant. Use for codec A-vs-B comparisons; not for production monitoring — see vmaf-neg.md.
--score-backend auto libvmaf backend: auto, cpu, cuda, sycl, hip. auto picks the fastest available (cuda > sycl > hip > cpu); a specific name is honoured strictly and errors out when unavailable.
--ffmpeg-bin ffmpeg Path to the ffmpeg binary.
--vmaf-bin vmaf Path to the vmaf binary.
--ffprobe-bin ffprobe Path to the ffprobe binary (used for HDR detection).
--output corpus.jsonl JSONL output path.

Search-mode flags:

Flag Default Description
--coarse-to-fine off Run a 2-pass coarse-then-fine CRF search instead of the full grid. With the defaults that is 15 encodes rather than 52 — see vmaf-tune-coarse-to-fine.md.
--coarse-step 10 CRF step for the coarse pass.
--fine-radius 5 ± radius around the best-coarse CRF for the fine pass.
--fine-step 1 CRF step for the fine pass.
--target-vmaf unset Target VMAF. The search refines around the smallest CRF whose score meets it; without a target it refines around the highest-VMAF coarse point.

HDR flags (mutually exclusive — see vmaf-tune-hdr-and-sampling.md):

Flag Description
--auto-hdr (default) Probe each source with ffprobe and inject HDR codec args when PQ / HLG signalling is detected.
--force-sdr Treat every source as SDR; skip detection and flag injection.
--force-hdr-pq Treat every source as HDR PQ (SMPTE-2084) regardless of the probe.
--force-hdr-hlg Treat every source as HDR HLG (ARIB STD-B67) regardless of the probe.

Example — full grid:

vmafx-tune-go corpus \
  --source ref.yuv \
  --width 1920 --height 1080 \
  --framerate 24 --duration 10 \
  --preset medium \
  --crf 20 --crf 26 --crf 32 \
  --output corpus.jsonl

Example — coarse-to-fine against a VMAF target:

vmafx-tune-go corpus \
  --source ref.yuv \
  --width 1920 --height 1080 \
  --preset medium \
  --coarse-to-fine --target-vmaf 93 \
  --output corpus.jsonl

Rows stream to the output file as each cell completes, so an interrupted sweep leaves a usable partial corpus. The selected scoring backend is echoed on stderr (vmafx-tune: scoring backend = cpu) before the first encode, and the row count is echoed when the sweep finishes.

When --keep-encodes is off, cleanup is part of a successful cell. A missing temporary encode is harmless (an injected runner may already have removed it), but any other removal failure stops the sweep instead of silently leaking data. Likewise, when source hashing is enabled, an unreadable source fails before the first encode; use --no-source-hash only when provenance is deliberately not required.

Ported subcommands (encoder introspection)

benchmark — Rank encoders from an existing corpus

Answers the standard post-sweep question: which encoder hit the target quality at the lowest bitrate? It reads a Phase-A corpus JSONL written by vmaf-tune corpus and launches no ffmpeg and no libvmaf — the corpus stays the source of truth.

vmafx-tune-go benchmark --from-corpus <JSONL> [flags]

For every encoder in the corpus the report picks the lowest-bitrate row whose measured VMAF clears --target-vmaf. An encoder that never clears is reported with status unmet and its closest miss, so a missing encoder build is never mistaken for a quality result.

Required flags:

Flag Description
--from-corpus Phase-A corpus JSONL to benchmark

Optional flags:

Flag Default Description
--target-vmaf 92 Matched-quality threshold each encoder must clear.
--baseline-encoder lowest-bitrate encoder that clears Encoder used for the bitrate_delta_pct column.
--format markdown Report format: markdown, json or csv.
--output, -o stdout Report destination. Parent directories are created.

Example — Markdown to stdout:

vmafx-tune-go benchmark --from-corpus corpus.jsonl

Example — CSV at a stricter target, with a pinned baseline:

vmafx-tune-go benchmark \
  --from-corpus corpus.jsonl \
  --target-vmaf 95 \
  --baseline-encoder libx264 \
  --format csv \
  --output benchmark.csv

Rows are ranked with cleared encoders first (ascending bitrate), then the unmet ones. A row reports:

Field Meaning
encoder Encoder token from the corpus rows.
status ok when the encoder cleared the target, unmet otherwise.
target_vmaf / margin The requested threshold, and the selected row's VMAF minus it (negative when unmet).
bitrate_kbps Bitrate of the selected row.
bitrate_delta_pct Percentage difference against the baseline encoder; null / blank when no encoder cleared.
rows / source_count / preset_count How many eligible corpus rows the encoder contributed, and how many distinct sources / presets they span.
encode_fps / score_fps Means over the encoder's rows, counting positive finite samples only; null / blank when no row supplied timings.
best The selected corpus row (src, preset, crf, vmaf_score, bitrate_kbps, vmaf_model).

Rows are excluded from the report when exit_status is non-zero, when vmaf_score or bitrate_kbps is missing or non-finite, or when the row carries no encoder name.

The CSV output uses CRLF line endings (Python's csv "excel" dialect), and the JSON output is stable pretty RFC 8259 with sorted keys — both byte-identical to vmaf-tune benchmark.

Ported subcommands (Stage 5)

auto — Phase F adaptive recipe-aware planner

Composes the per-phase tuning stages into one deterministic decision tree and emits a JSON plan. Optionally realises the winning cell as a real encode plus a libvmaf score.

vmafx-tune-go auto --src <video> [flags]

Required flags:

Flag Description
--target-vmaf Quality target on the standard VMAF [0, 100] scale
--src Source video. Required unless --smoke.

Optional flags:

Flag Default Description
--target-vmaf 93 Target pooled-mean VMAF.
--max-budget-bitrate 8000 Upper bound on the picked rendition's bitrate, in kbps.
--allow-codecs libx264 Comma-separated codec list the tree may pick from. A single entry short-circuits the compare-shortlist stage.
--codec (unset) Pin the codec choice, overriding the --allow-codecs ranking. Also short-circuits the shortlist stage.
--sample-clip-seconds 0 Propagate this clip length to internal sweeps rather than re-deciding per stage. 0 = full source.
--smoke false Exercise the composition with synthetic metadata — no ffprobe, no ffmpeg, no ONNX.
--output stdout Write the JSON plan here.
--execute false After planning, run real FFmpeg encodes and libvmaf scores for the selected cell(s).
--runs-dir runs Output directory for encoded files and tune_results.jsonl (used with --execute).
--execute-all false With --execute: run every plan cell, not just the winner. Useful for post-hoc A/B comparison.
--model (unset) Optional predictor_<codec>.onnx path. Default uses the analytical fallback curve.

Exit codes (identical to vmaf-tune fast):

Code Meaning
0 Recommendation emitted; proxy and verify agree within --proxy-tolerance.
2 Usage or environment error (bad CRF range, missing --src, unavailable backend, proxy unavailable).
3 Recommendation emitted, but the proxy/verify gap exceeds tolerance. Fall back to the slow Phase A grid (ADR-0276). The payload is still written.

Example — smoke run (works on any host):

vmafx-tune-go fast --smoke --target-vmaf 90

### `encode-profile` — Reproduce one recommendation from a report

Every vmaf-tune report embeds a machine-readable `encoder_profile` payload.
This subcommand reads that payload, selects one recommendation, and runs the
matching FFmpeg encode.

```text
vmafx-tune-go encode-profile --profile <FILE> --output <FILE> [flags]

The profile is accepted in any of the three shapes a report ships in:

Input How the payload is found
Report JSON (.json) Parsed directly; the encoder_profile block is unwrapped if present.
Report HTML (.html / .htm) Extracted from the raw-JSON <pre> block and HTML-unescaped.
Report Markdown (anything else) Extracted from the fenced JSON payload.

Selection defaults to the first Pareto-selected row with the lowest bitrate. --codec and --target-vmaf narrow the candidate set; --recommendation-index then picks the Nth survivor (zero-based, applied after filtering).

Required flags:

Flag Description
--profile Report JSON / HTML / Markdown containing encoder_profile
--output, -o Encoded output path

Selection flags:

Flag Default Description
--codec all Restrict selection to one codec.
--target-vmaf all Restrict selection to one target VMAF (matched with a 1e-6 absolute tolerance).
--recommendation-index 0 (first) Zero-based index after the filters are applied.

Override flags (each falls back to the profile value when omitted):

Flag Description
--src Override the source path stored in the profile.
--preset Override the stored / adapter-default preset.
--pix-fmt Override the raw-source pixel format (default yuv420p).
--framerate, --width, --height Override the raw-source geometry.
--duration Override the encode duration in seconds. Passing 0 explicitly suppresses the profile's own duration bound.
--source-kind auto (default), container or raw. Under auto, .yuv / .raw / .rgb / .gray are raw and everything else is a container.
--sample-clip-seconds, --sample-clip-start-s Input-side clip length / offset forwarded to FFmpeg.
--extra-ffmpeg-arg Append one raw FFmpeg argv token after the codec args; repeat as needed. Use --extra-ffmpeg-arg=-movflags for tokens starting with -.
--ffmpeg-bin Override the profile's ffmpeg_bin (default: profile value, then ffmpeg).
--dry-run Print the selected recommendation and the exact FFmpeg argv without encoding.

Example — inspect the selection without encoding:

vmafx-tune-go encode-profile \
  --profile report.json \
  --output /dev/null \
  --dry-run

Example — reproduce the best x265 row at target 95:

vmafx-tune-go encode-profile \
  --profile report.html \
  --codec libx265 \
  --target-vmaf 95 \
  --output encoded.mkv

Output schema. The result is a single JSON document on stdout (sorted keys, two-space indent). A --dry-run emits ok, dry_run, profile, selected, ffmpeg_argv and output; a real run replaces dry_run with the encode outcome:

```json
{
  "ok": true,
  "profile": "report.json",
  "selected": { "codec": "libx265", "crf": 28 },
  "ffmpeg_argv": ["ffmpeg", "-y", "-hide_banner"],
  "output": "encoded.mkv",
  "exit_status": 0,
  "encode_size_bytes": 148213,
  "encode_time_ms": 1843.2,
  "encoder_version": "libx265-3.5",
  "ffmpeg_version": "n8.1",
  "stderr_tail": "..."
}

Exit status. On a real run the process exit status is FFmpeg's own, so a failed encode surfaces the encoder's code (for example 254) rather than a generic 1. The JSON payload is still written first, so a wrapper script can read exit_status and stderr_tail regardless.

Exit codes

benchmark, encode-profile and the sidecar group use the same exit-status convention as the Python CLI they replace:

Status Meaning
0 Success.
2 A usage or validation failure — a missing/unknown flag, an unparseable flag value, a missing input file, a filter that matches no recommendation, a baseline encoder absent from the corpus, an unknown --codec or unreadable feature / capture file on sidecar.
1 sidecar only: the cache directory, host UUID or state.json could not be written (an uncaught OSError in Python). See Exit codes and diagnostics.
other encode-profile only: FFmpeg's own exit status from a failed encode.

Note that the earlier ports (compare, ladder, report) still report every failure as 1; that pre-existing inconsistency is tracked separately.

Hardware-encoder caveats.

  • The emitted argv contains no -init_hw_device chain. FFmpeg's QSV bridge on Linux needs -init_hw_device vaapi=va:<node> -init_hw_device qsv=qsv_dev@va -filter_hw_device va before the first -i, plus a format=nv12,hwupload filter, or the encode fails with -22 (see ADR-0601). This matches the Python implementation exactly: vmaf-tune injects that chain in its compare sweep, never in encode-profile. Supply the flags yourself with repeated --extra-ffmpeg-arg, or drive QSV through compare.
  • Hardware encoders (NVENC / QSV / AMF) reject sources below roughly 320x240. A profile built from a smaller clip will fail at the encoder.
  • av1_videotoolbox is a placeholder: upstream FFmpeg ships no such encoder, so the adapter refuses to emit an argv shape it cannot verify (ADR-0339).

--model is a Go-side addition: the Python auto driver always constructs its predictor without a model path, which is the analytical fallback this flag defaults to. Supplying a model routes inference through the ONNX bridge in pkg/ai, which degrades back to the analytical curve when the ORT runner is not on PATH.

Example — plan only:

vmafx-tune-go auto \
  --src src.mp4 \
  --target-vmaf 93 \
  --allow-codecs libx264,libx265

Example — plan and realise the winner:

vmafx-tune-go auto \
  --src src.mp4 \
  --target-vmaf 95 \
  --max-budget-bitrate 6000 \
  --execute --runs-dir runs/

How the plan is built

  1. Probe the source: geometry and duration via ffprobe, HDR signalling via the colour-metadata classifier. Every probe failure degrades to conservative defaults (1920x1080, duration 0, SDR) rather than aborting the run.
  2. Apply the content recipe. The content class selects a small override set — a narrower or wider conformal-confidence gate, a forced single-rung ladder, a saliency intensity, and a target-VMAF offset. An HDR source carrying only a generic live_action label is promoted to the HDR recipe. The overrides load from ai/data/phase_f_recipes_calibrated.json when that file is reachable, and fall back to the documented placeholders with a one-line warning otherwise.
  3. Walk the ten short-circuits, recording each one that fires under metadata.short_circuits.
  4. Estimate each (rung, codec) cell: invert the predictor for a CRF, then estimate the VMAF and bitrate that CRF would produce.
  5. Pick a winner against the VMAF target and the bitrate budget.

Short-circuits

Each predicate names a stage the tree can skip. They are evaluated in this order, and the order is part of the output contract.

Name Fires when
ladder-single-rung Source height is below 2160, or the recipe forces a single rung.
codec-pinned --codec is set, or --allow-codecs resolves to one entry.
predictor-gospel The predictor verdict is GOSPEL; trust its CRF and skip the coarse-to-fine fallback.
skip-saliency Content class is neither animation nor screen_content.
sdr-skip The source carries no HDR signalling.
sample-clip-propagate --sample-clip-seconds is positive; propagate it verbatim to internal sweeps.
skip-per-shot The source is both shorter than 5 minutes and below 0.15 shot variance.
low-complexity The probe-encode bitrate is under 200 kbps. Dormant when no probe has run.
baseline-meets-target A default-CRF encode already meets the target. Dormant when no baseline was scored.
no-two-pass The resolved codec adapter does not support two-pass encoding.

Confidence-aware escalation

Each cell carries a conformal interval width, and that width decides whether the predictor's own verdict is overridden:

Interval width Decision
<= tight (default 2.0) skip-escalation — trust the point estimate even on a FALL_BACK verdict.
>= wide (default 5.0) force-escalation — escalate even on a GOSPEL verdict.
between Defer to the native verdict.
NaN (uncalibrated) Defer to the native verdict.

Without a calibration sidecar the interval is uncalibrated, so cells carry NaN and no override happens. The 2.0 / 5.0 defaults are an emergency floor, not a corpus fit.

Plan JSON

The plan is a {"cells": [...], "metadata": {...}} object with sorted keys, byte-compatible with the Python vmaf-tune auto output.

Each cell carries: rung, codec, verdict, crf, estimated_vmaf, estimated_bitrate_kbps, hdr_args, sample_clip_seconds, confidence_decision, interval_width, effective_predictor_target_vmaf, prediction_source, saliency_intensity, and selected.

metadata.winner.status is one of:

Status Meaning
budget_and_quality_met A cell satisfies both the target and the budget.
quality_met_budget_exceeded Quality is reachable, but every such cell is over budget; the smallest overage wins.
target_unmet No cell reaches the target; the closest miss is returned so you get a concrete next encode.
no_eligible_cells No cell carried finite estimates.

The plan JSON is not strict RFC 8259. An uncalibrated interval_width is emitted as the bare token NaN, exactly as CPython's json.dumps does with its default allow_nan=True. This is deliberate byte-compatibility with the Python emitter. Parse it with Python's json module or another permissive parser; jq --strict and Go's encoding/json will reject it. The --execute results log (tune_results.jsonl) is strict — non-finite values there are rendered as null.

Execute mode

With --execute the selected cell is encoded and scored, and one row per executed cell is appended to <runs-dir>/tune_results.jsonl. Encoded files land beside it as encode_<index>_<codec>_<preset>_crf<n>.mkv.

A failed encode is recorded in its row with a non-zero encode_exit_status, and scoring is skipped for that cell. The command exits non-zero only when cells were executed and none scored successfully.

sidecar — Local on-host predictor sidecar

Trains and inspects a bias-correction term on top of the shipped predictor:

sidecar_vmaf = predictor_vmaf + sidecar_correction(features)

The shipped predictor is never mutated, so model upgrades stay deterministic and reproducible across hosts.

vmafx-tune-go sidecar <status|predict|record|batch-record> [flags]

Flags shared by every nested subcommand:

Flag Default Description
--codec libx264 Codec bucket for the sidecar state. Must be one of the 19 registered codec names; an unknown name is a usage error whose message lists them.
--cache-dir ${XDG_CACHE_HOME:-~/.cache}/vmaf-tune/sidecar Sidecar cache root. A relative path stays relative in state_path.
--predictor-version predictor_v1 Predictor-version namespace.
--model (unset) Optional ONNX predictor — see the ONNX note. Unset uses the analytical fallback.
--json false Emit machine-readable JSON instead of the one-line text form.

Nested subcommands:

Subcommand Extra flags Purpose
status — Print state metadata: codec, host UUID, state path, predictor version, update count, residual RMS.
predict --features-json, --crf Predict VMAF with the correction folded in. Reports the base score, the correction, and the sum.
record --features-json, --crf, --observed-vmaf, --no-persist Fold one observed encode result into the fit.
batch-record --captures-jsonl Fold a JSONL capture file, one observation per row, persisting once at the end.

Example — inspect, train from a capture log, then predict:

vmafx-tune-go sidecar status --json
vmafx-tune-go sidecar batch-record --captures-jsonl captures.jsonl --json
vmafx-tune-go sidecar predict --features-json shot.json --crf 26 --json

sidecar status

Prints the state metadata for one codec bucket: the anonymous host UUID, the on-disk state path, the predictor-version namespace, the number of folded-in captures, and the RMS of the buffered residuals (the drift signal).

vmafx-tune-go sidecar status --codec libx264 --json
{
  "codec": "libx264",
  "host_uuid": "0123456789abcdef0123456789abcdef",
  "n_updates": 0,
  "predictor_version": "predictor_v1",
  "recent_residual_rms": 0.0,
  "schema": "vmaf-tune-sidecar-status/v1",
  "schema_version": 1,
  "state_path": "/home/u/.cache/vmaf-tune/sidecar/predictor_v1/libx264/state.json"
}

Without --json the same fields print as one line:

codec=libx264 predictor_version=predictor_v1 updates=0 residual_rms=0.000000 state=/home/u/.cache/vmaf-tune/sidecar/predictor_v1/libx264/state.json

sidecar predict

Predicts VMAF for one shot at one CRF with the correction applied.

Flag Description
--features-json (required) Path to a JSON object carrying the shot's feature values.
--crf (required) CRF to predict at.
vmafx-tune-go sidecar predict --features-json shot.json --crf 26 --json

The payload (schema vmaf-tune-sidecar-predict/v1) carries base_vmaf (the bare predictor), correction, and sidecar_vmaf (the sum, clamped to [0, 100]), plus codec, crf and n_updates. The text form is base=<b> correction=<c> sidecar=<s> updates=<n> with six decimals.

sidecar record

Folds one observed VMAF measurement into the ridge fit.

Flag Description
--features-json (required) Path to the shot's feature JSON.
--crf (required) CRF the observation was measured at.
--observed-vmaf (required) Observed libvmaf score for the encode.
--no-persist Update in memory only; mainly useful for tests.
vmafx-tune-go sidecar record \
  --features-json shot.json \
  --crf 26 \
  --observed-vmaf 91.75 \
  --json

The residual is computed against the bare predictor, never against the sidecar-corrected value, so repeated captures converge rather than compounding. The payload (schema vmaf-tune-sidecar-record/v1) is the status payload plus crf, observed_vmaf, base_vmaf and residual (observed_vmaf - base_vmaf); the text form is recorded updates=<n> residual=<r> state=<path>.

sidecar batch-record

Folds a whole JSONL capture file into the fit, one observation per line.

Flag Description
--captures-jsonl (required) Path to the JSONL capture file.
vmafx-tune-go sidecar batch-record --captures-jsonl captures.jsonl --json

Each line is a JSON object carrying the feature fields (either at the top level or nested under a features key) plus crf and observed_vmaf. Lines are split on \n, \r\n or a lone \r (CPython's universal-newline rule) with no length limit; blank lines are ignored but still numbered. A malformed row — not a JSON object, a missing or non-numeric required field, a crf string that is not an integer literal — is reported as vmafx-tune sidecar batch-record: skip line <n>: <reason> on stderr, counted in rows_skipped, and the run still succeeds. That is deliberate: a capture log is often partially corrupt after an interrupted run, and losing the good rows to one bad line would be worse. State is written once at the end, and only when at least one row was recorded. The payload (schema vmaf-tune-sidecar-batch-record/v1) is the status payload plus rows_recorded and rows_skipped; the text form is recorded=<rows> skipped=<rows> updates=<n> state=<path>.

Feature JSON

--features-json takes an object of shot features, or a {"features": {...}} wrapper so a capture row can carry crf and observed_vmaf alongside. The four probe_* fields are required — a zero probe bitrate would train the fit on a fabricated complexity barometer — and everything else defaults to 0. Values may be JSON numbers or numeric strings ("2400"), as CPython's float() accepts.

Key Required Meaning
probe_bitrate_kbps yes Average bitrate over the probe encode.
probe_i_frame_avg_bytes yes Mean I-frame size.
probe_p_frame_avg_bytes yes Mean P-frame size.
probe_b_frame_avg_bytes yes Mean B-frame size (0 for codecs without B-frames).
saliency_mean, saliency_var no Saliency signals; 0 when unavailable.
frame_diff_mean, y_avg, y_var no FFmpeg signalstats aggregates.
shot_length_frames, fps, width, height no Structural metadata.
{
  "probe_bitrate_kbps": 4200.5,
  "probe_i_frame_avg_bytes": 51234.0,
  "probe_p_frame_avg_bytes": 8123.25,
  "probe_b_frame_avg_bytes": 2011.75,
  "saliency_mean": 0.42,
  "saliency_var": 0.031,
  "frame_diff_mean": 7.5,
  "y_avg": 112.25,
  "y_var": 1830.5,
  "shot_length_frames": 240,
  "fps": 24.0,
  "width": 1920,
  "height": 1080
}

State and privacy

State lives at:

${XDG_CACHE_HOME:-~/.cache}/vmaf-tune/sidecar/
  host-uuid                                    # random 128-bit token
  <predictor-version>/<codec>/state.json       # ridge weights + inverse Gram

The host UUID is drawn from a CSPRNG on first use. It is never derived from a MAC address, hostname, /etc/machine-id, CPUID, or any other machine-identifying signal.

A predictor-version or schema mismatch on load discards the fit and resets to cold start, keeping only the host UUID. That is what makes a shipped-model upgrade safe: a stale correction can never be replayed against a refreshed predictor. At cold start the weights are zero, so the correction is exactly 0.0 and the sidecar returns the bare predictor's value untouched.

A corrupt state.json also cold-starts — including one whose weights or a_inv carry null, the residue of a NaN capture — and the corrupt file is left in place so you can inspect it.

Exit codes and diagnostics

The four subcommands follow the Python vmaf-tune sidecar exit contract:

Status Meaning
0 Success — including a batch-record run in which every row was skipped.
1 An I/O failure the Python CLI does not catch either: the cache directory cannot be created, the host UUID or state.json cannot be written, or stdout cannot be written.
2 A usage or validation failure: an unknown or unparseable flag, a missing required flag, an unknown --codec, an unresolvable --model, a --features-json that cannot be read / is not a JSON object / lacks a required key, or a --captures-jsonl that cannot be read.

Diagnostics go to stderr and are not byte-identical to Python's: cobra prefixes them with Error: where Python prints vmaf-tune sidecar <cmd>:, the batch-record skip lines carry the vmafx-tune prefix, and the reason text is Go's rather than CPython's exception message. Nothing is written to stdout on failure.

Byte compatibility with vmaf-tune sidecar

For the same inputs, every stdout payload (JSON and text form) and every byte of state.json written by the Go binary is identical to the Python CLI's: the analytical predictor, the Sherman–Morrison update, and the json.dumps(..., indent=2, sort_keys=True) rendering are reproduced to the last bit, and the atomic write leaves the same single state.json behind (no .tmp residue). cmd/vmafx-tune/cmd/testdata/sidecar/ holds fixtures dumped from the Python CLI by regen.sh (pinned host UUID, relative --cache-dir), and TestSidecarPythonParity replays the same 23-step operator sequence — cold status, three records, two batch-record loads, warm status and predict, a --no-persist record, a second codec bucket, and nine error paths — requiring identical stdout, identical state.json snapshots and identical exit statuses. Known, deliberate differences:

  • A capture row containing the non-standard JSON tokens NaN / Infinity is skipped by the Go binary and counted in rows_skipped. CPython's json.loads accepts the tokens: a NaN feature then silently skips the update while still counting the row as recorded, and a NaN observed_vmaf poisons the ridge weights (they persist as null and the next load cold-starts).
  • Hand-edited state files outside the shape either writer produces are not guaranteed to load identically (for example a numeric string where CPython's float() would coerce and Go's decoder will not).
  • --model is resolved differently — see the ONNX note.

ONNX predictor models

The two binaries interpret --model differently, and the Python behaviour itself depends on the host:

  • Python takes a filesystem path to predictor_<codec>.onnx. With onnxruntime importable, a missing file raises and the CLI exits 2; without it, the flag is silently ignored and the analytical curve is used (exit 0).
  • Go resolves the value through the model registry (pkg/ai): an absolute path that exists, else <model-dir>/<name>.onnx, else <model-dir>/<name>, where the model dir is $VMAFX_MODEL_DIR or /usr/local/share/vmafx/model. An unresolvable name exits 2. Inference then goes through the vmafx-ort-runner subprocess, which this repository does not build; when the runner is absent from PATH the predictor falls back to the analytical curve silently — a warning is logged only when the runner is present and inference fails.

Omit --model to use the analytical fallback — the default, and the only path the parity fixtures exercise — and the two binaries agree byte for byte.

Ported subcommands (per-shot tuning)

tune-per-shot — Per-shot CRF tuning

Cuts the source into shots, bisects a CRF against the VMAF target inside each shot, and emits an FFmpeg encoding plan: one command per segment plus the concat-demuxer command that stitches them into the final file.

The plan is emitted, not executed. Segment encodes are independent, so you can run them sequentially or in parallel, then run the concat command.

vmafx-tune-go tune-per-shot --src <video> [flags]

Pipeline:

  1. Shot detection — shells out to the fork's vmaf-perShot binary (vmaf-perShot.md), which wraps TransNet V2. A missing or failing binary degrades to one shot spanning the clip.
  2. Uniform-window splitter — any shot longer than --max-shot-duration is sliced into equal sub-shots, so an under-cutting detector (fades, short clips) still yields a usable timeline.
  3. Per-shot bisect — each shot is extracted to raw YUV and run through the CRF bisect against --target-vmaf.
  4. Plan emission — the recommendations become segment + concat commands.

Required flags:

Flag Description
--src Reference video: raw YUV, or any FFmpeg-readable container

Source geometry:

Flag Default Description
--width auto-probed Source width. Required for raw YUV (.yuv / .raw); auto-probed via ffprobe for containers.
--height auto-probed Source height. Same rule as --width.
--pix-fmt yuv420p Source pixel format.
--framerate auto-probed Source framerate. Falls back to 24.0 when the probe yields nothing.
--bitdepth 8 Source YUV bit depth: 8, 10 or 12.
--total-frames 0 Frame count for the single-shot fallback when vmaf-perShot is unavailable.

Shot detection:

Flag Default Description
--per-shot-bin vmaf-perShot Path to the shot-detector binary.
--scene-threshold detector default (12.0) Override the detector's mean-absolute-luma-delta cut threshold. Lower yields more shots.
--max-shot-duration 2.0 Uniform-window splitter, in seconds. 0 disables it.

Tuning:

Flag Default Description
--target-vmaf 92 Target pooled-mean VMAF per shot.
--encoder libx264 Codec adapter (see the codec table below).
--preset codec default (medium) Preset for the bisect encodes.
--crf-min / --crf-max codec absolute window Inclusive bisect search bounds. Pass both or neither.
--max-iterations 8 Maximum encode+score rounds per shot.
--vmaf-model vmaf_v1.0.16_3d0h Model passed to the vmaf binary.
--neg off Route the model to its NEG variant. There is no NEG counterpart to any vmaf_v1.0.16_* model, so --neg also selects the v0.6.1 generation (vmaf_v0.6.1neg). See vmaf-neg.md.
--score-backend auto libvmaf backend: auto, cpu, cuda, sycl, hip. An explicit backend that the host cannot provide fails fast rather than silently downgrading.
--vmaf-bin / --ffmpeg-bin vmaf / ffmpeg Binary paths.
--workdir $VMAFTUNE_WORKDIR or OS temp Scratch space for encode / decode artefacts. Raw YUV decodes are large — point this at a volume with room.
--max-concurrent-decodes 1 Concurrent reference-YUV decodes. 1 is safest on space-constrained volumes.

Output:

Flag Default Description
--plan-out stdout Destination for the JSON plan.
--output per_shot_encode.mp4 Final concatenated encode path named inside the plan.
--segment-dir <output dir>/segments Directory the segment commands write into.
--script-out — Also write the plan as a copy-paste shell script.

Example:

vmafx-tune-go tune-per-shot \
  --src src.mp4 \
  --target-vmaf 92 \
  --encoder libx264 \
  --plan-out plan.json \
  --segment-dir ./segments \
  --output final.mp4

Plan JSON schema

Byte-compatible with vmaf-tune tune-per-shot: keys are sorted, floats keep Python's repr() form, and a shot whose predicate never measured a bitrate carries null rather than NaN.

{
  "concat_command": ["ffmpeg", "-y", "-hide_banner", "-f", "concat", "..."],
  "encoder": "libx264",
  "framerate": 24.0,
  "predicate": "bisect",
  "segment_commands": [["ffmpeg", "-y", "-hide_banner", "-ss", "0.000000", "..."]],
  "shots": [
    {
      "bitrate_kbps": 1234.57,
      "crf": 22,
      "end_frame": 48,
      "predicted_vmaf": 92.5,
      "start_frame": 0
    }
  ],
  "target_vmaf": 92.0
}

start_frame is inclusive and end_frame exclusive (half-open). The vmaf-perShot sidecar uses an inclusive end frame; the tuner normalises it.

The concat listing (concat.txt) is written next to --plan-out when --segment-dir is not given, matching the Python. Pass --segment-dir explicitly to pin the listing and the segment commands to the same directory; the command logs a WARN when the two diverge.

Supported codecs

tune-per-shot accepts the ten codecs the Go encoder registry can construct. Each emits its own quality knob in the plan:

Codec Plan argv shape
libx264, libx265 -c:v NAME -preset medium -crf N
libsvtav1 -c:v libsvtav1 -preset 7 -crf N (integer preset)
libaom-av1 -c:v libaom-av1 -cpu-used 4 -crf N
h264_nvenc, hevc_nvenc -c:v NAME -preset p4 -cq N
h264_qsv, hevc_qsv -c:v NAME -preset medium -global_quality N
h264_amf, hevc_amf -c:v NAME -quality balanced -rc cqp -qp_i N -qp_p N

The Python registry carries seven more adapters — av1_nvenc, av1_qsv, av1_amf, the four VideoToolbox encoders, libvvenc and libvpx-vp9. They have no Go encoder implementation yet, so --encoder rejects them by name and points at the Python binary.

Flags with no Go implementation

Both fail fast with an actionable message rather than being accepted and silently ignored:

Flag Why Use instead
--predicate-module MODULE:CALLABLE Loads a Python callable at runtime; Go has no runtime import. The Go equivalent is the pershot.PredicateFn seam, available to library callers. vmaf-tune tune-per-shot --predicate-module ...
--fast-nr NR early-elimination runs the nr_metric_v1 ONNX model through onnxruntime; the Go binary has no ONNX runtime binding. vmaf-tune tune-per-shot --fast-nr

Python-only flags

Every vmaf-tune subcommand is ported. A few individual flags still need the Python implementation, because they depend on in-process ONNX inference or on importing a Python callable at runtime:

Flag Subcommand Why Use instead
--fast-nr tune-per-shot NR early-elimination needs an ONNX forward pass per bisect midpoint vmaf-tune tune-per-shot --fast-nr
--predicate-module tune-per-shot Imports an arbitrary Python MODULE:CALLABLE at runtime vmaf-tune tune-per-shot --predicate-module
--saliency-aware recommend-saliency Requires a saliency ONNX forward pass vmaf-tune recommend-saliency --saliency-aware

recommend-saliency --saliency-aware and predict --use-saliency are accepted by the Go binary. When the saliency session cannot be built, recommend-saliency proceeds without an ROI map (the report's saliency_aware field then reads false), and predict logs a warning and degrades saliency moments to 0.0, matching the Python behavior.

--model (on predict, sidecar, auto) routes inference through the vmafx-ort-runner subprocess (vmafx-ort-runner.md), which the dev container and the Go CI job build from cmd/vmafx-ort-runner (ADR-1134). When the runner is absent from PATH, or present but linked against a libvmaf built without ONNX Runtime (exit 3), the predictor logs a warning carrying the runner's stderr and falls back to the analytical curve — the same fallback the Python takes without onnxruntime, but reported rather than silent. On sidecar an unresolvable model name is a usage error (exit 2); see the ONNX note for how the two binaries resolve the flag differently.

recommend's encode-driven path writes the same schema-v3 corpus JSONL the corpus subcommand does, and every key is present. Five corpus features are not carried by this group's port and their fields hold the same zero / empty values the Python emits when the feature is unavailable, so a reader filters on them exactly as it already does: the content-addressed encode cache (ADR-0298), HDR detection (ADR-0295), TransNet-V2 shot metadata (shot_count stays 0), sample-clip windowing (clip_mode stays full), and the encoder-internal pass-1 stats (the ten enc_internal_* columns stay 0.0, which is what the Python aggregator returns for an empty frame list). Those belong to the corpus port.

Configuration and logging

vmafx-tune-go runs each subcommand inside the golusoris clikit (cobra + fx) framework. The framework injects a structured *slog.Logger and a config tree into every subcommand, so run diagnostics (sweep start/finish, ladder build summary, report rendering) are emitted as structured log lines on stderr, while subcommand output (JSON / Markdown / HTML) goes to stdout or the --output file. This keeps machine-readable output separate from logs when you pipe stdout.

Configuration is read from environment variables under the VMAFX_ prefix. golusoris maps each underscore in the variable name to a config-path delimiter (VMAFX_LOG_LEVEL → log.level):

Environment variable Config key Effect Default
VMAFX_LOG_LEVEL log.level Minimum log level: debug, info, warn, error info
VMAFX_LOG_FORMAT log.format Log handler: auto (tint on a TTY, JSON otherwise), tint, json auto

Examples:

# Quiet the per-run INFO diagnostics; keep warnings/errors.
VMAFX_LOG_LEVEL=warn vmafx-tune-go report results.json

# Force JSON logs for machine ingestion regardless of TTY.
VMAFX_LOG_FORMAT=json vmafx-tune-go compare --reference src.mp4 --targets 90

Migration roadmap

Stage Scope ADR Status
Stage 1 compare subcommand, libx264/libx265, single/multi-target bisect ADR-0705 Merged
Stage 2 ladder subcommand, hardware encoders (NVENC, QSV, AMF), convex hull + knee selection ADR-0730 Merged
Stage 3 Downscale plumbing — Merged
Stage 4 report subcommand, Markdown + HTML rendering ADR-0770 Merged
golusoris Migrate the CLI root + subcommands onto the golusoris clikit (cobra + fx) framework; VMAFX_-prefixed config + injected slog ADR-1119 Merged
ML-driven recommend, predict, recommend-saliency, prefilter; codec-adapter registry, encode/score drivers, predictor, saliency pipeline, native TPE — This PR

| Encoder introspection | benchmark + encode-profile subcommands; pkg/benchmark, pkg/codecadapter, pkg/encodeprofile, pkg/pyjson | ADR-0770 | This PR | | Stage 5 | tune-per-shot subcommand, conformal CLI wiring | Planned | — |

| golusoris | Migrate the CLI root + subcommands onto the golusoris clikit (cobra + fx) framework; VMAFX_-prefixed config + injected slog | ADR-1119 | This PR | | Stage 5 (corpus/sidecar) | corpus + sidecar subcommands; pkg/codecadapter, pkg/corpus, pkg/pyjson (corpus); pkg/tune/predictor, pkg/tune/sidecar, pkg/tune/pyjson (sidecar) | ADR-1125 | Merged (#1153) | | Stage 5 (per-shot) | tune-per-shot subcommand, conformal CLI wiring | Planned | — | | Stage 6 | fast subcommand (requires ONNX Go binding) | Planned | — |

| Stage 5 | tune-per-shot subcommand: pkg/pershot, pkg/scorebackend, codec-adapter table, raw-YUV scorer | ADR-0705 | This PR | | Stage 5b | conformal CLI wiring | Planned | — | | Stage 6 | fast subcommand + tune-per-shot --fast-nr (both require an ONNX Go binding) | Planned | — |

| Stage 6 | fast subcommand + pkg/conformal + pkg/scorebackend; smoke path complete, production path blocked on ONNX named inputs | ADR-0276 / ADR-0304 | This PR | | Stage 5 | tune-per-shot subcommand, conformal CLI wiring | Planned | — | | Stage 6b | fast production mode (needs a named-input ONNX seam — see Production-mode blocker) | Planned | — |

| golusoris | Migrate the CLI root + subcommands onto the golusoris clikit (cobra + fx) framework; VMAFX_-prefixed config + injected slog | ADR-1119 | This PR | | Stage 5 | auto (Phase F planner + execute mode) and sidecar subcommands | ADR-1125 | Merged (#1153) | | Stage 6 | tune-per-shot subcommand, conformal CLI wiring | Planned | — | | Stage 7 | fast subcommand (requires ONNX Go binding) | Planned | — | | Stage N | Feature parity; rename binary to vmafx-tune | Planned | — |

Correction. An earlier revision of this table listed pkg/conformal as merged under Stage 3. It was not: no such package existed on master. The package landed with Stage 6 above, and the roadmap row has been corrected.

Architecture

The CLI root and every subcommand are built with the golusoris clikit (cobra + fx) framework (ADR-1119): clikit.New builds the root, clikit.Command builds each subcommand, and a thin withGolusoris adapter boots a one-shot fx graph per invocation so the command receives an injected *slog.Logger and config, runs to completion, and propagates its error as the process exit code.

The Go binary uses an adapter pattern with these core packages:

  • pkg/encoder/ — Encoder interface + software and hardware encoder implementations. Each encoder shells out to ffmpeg; no libavcodec CGo dependency. Stage 5 adds the codec-adapter policy table (Adapter: preset vocabulary, quality windows, per-codec argv shape), AdapterEncoder, and ProbeSource.

Each encoder shells out to ffmpeg; no libavcodec CGo dependency. EncodeParams.InputArgs carries ffmpeg input-side options so raw-YUV sources (-f rawvideo -pix_fmt -s -r) and sample clips (input-side -ss / -t, for fast-seek) work; EncodeParams.OutputPath pins a deterministic destination and EncodeResult.OutputSizeBytes reports the encode size for size-over-duration bitrate maths. - pkg/bisect/ — Stateless Run(src, enc, scoreFunc, params) function. The score function is injectable, enabling unit tests without a live vmaf binary. Stage 5 adds YUVScoreFunc, which decodes a containerised distorted file to raw YUV and invokes vmaf with full geometry / model / backend flags. - pkg/ladder/ — Build(src, encoder, Params) function. Convex hull (upperConvexHull), knee selection (selectRenditions), min-bitrate-gap filter. - pkg/pershot/ — Stage-5 shot detection, uniform-window splitter, per-shot tuning, and encoding-plan construction + JSON emission. - pkg/scorebackend/ — Stage-5 libvmaf backend resolution: parses the vmaf --help backend line, probes each vendor independently, and honours an explicit --score-backend strictly. - pkg/report/ — Stage-1: EmitJSON / EmitMarkdown renderers (single-run emit). Stage-4: RenderMarkdownMulti / RenderHTMLMulti (multi-file report rendering). - pkg/fast/ — the fast path: Recommend (the flow), RunTPE (the goptuna-backed search), NewSamplePredictor / NewVerifier (the probe and verify pipelines), and ORTProxy (the fr_regressor_v2 seam). - pkg/scorebackend/ — libvmaf backend detection and strict selection, ported from the selection half of vmaftune/score_backend.py. Detect intersects what the local vmaf --help advertises with what nvidia-smi / sycl-ls / rocminfo report; Select honours auto via a fallback chain and never silently downgrades an explicit request. - pkg/conformal/ — distribution-free prediction intervals for the VMAF predictor (split conformal and CV+ / jackknife+), ported from vmaftune/conformal.py. The JSON sidecar is byte-compatible with the Python writer, so a calibration produced by either implementation loads in the other. Not yet wired into a CLI flag — that is Stage 5.

The ML-driven group adds:

  • pkg/codecadapter/ — the nineteen-codec registry (quality windows, preset mapping onto each encoder's native axis, probe argv, two-pass argv).
  • pkg/ffencode/ — the encode driver the tuning subcommands share: raw-YUV geometry, named presets, injected extra params, sample-clip windowing. Distinct from pkg/encoder, which models the narrower Stage-1 bisect abstraction.
  • pkg/scorecli/ — the libvmaf CLI driver with explicit geometry, the backend selector and the canonical-6 pooled aggregates.
  • pkg/predictor/ — ShotFeatures, the per-codec analytical curve, the binary-search CRF inversion, feature extraction and the validation harness.
  • pkg/pershot/ — shot detection via vmaf-perShot with the single-shot fallback and the uniform long-shot splitter.
  • pkg/saliency/ — the full saliency ROI pipeline: YUV to ImageNet tensor, four temporal aggregators, QP mapping, per-block reduce, five sidecar formats.
  • pkg/prefilter/ — the frozen Pelorus knob contract plus a native TPE sampler.
  • pkg/recommend/, pkg/conformal/, pkg/uncertainty/, pkg/corpusrow/ — the predicate pickers, split-conformal intervals, confidence bands and the schema-v3 corpus row.
  • pkg/pyjson/ — renders payloads the way Python's json.dumps does, so the emitted JSON is byte-identical to the Python binary's. It is the one CPython-JSON encoder in the tree (ADR-1137); the former internal/pyjson, internal/pyjsonstrict and pkg/tune/pyjson copies were folded into it.

The corpus and sidecar subcommands add five more:

  • pkg/codecadapter/ — the ADR-0237 adapter registry: every codec's quality knob, preset vocabulary, validation rule, and ffmpeg argv slice. The encode driver never branches on codec identity.
  • pkg/corpus/ — the Phase A orchestrator: encode + score drivers, the pass-1 encoder-stats parser, HDR detection and codec-arg dispatch, shot metadata, scoring-backend selection, the JSONL reader / writer, and the coarse-to-fine search.
  • pkg/predictor/ — ShotFeatures plus the per-codec analytical VMAF curve and its CRF inversion.
  • pkg/tune/sidecar/ — the online-ridge bias-correction model (Sherman-Morrison rank-1 updates) and its cache-dir persistence. The port produced two implementations of this; the other (pkg/sidecar/) was removed because nothing imported it. The sidecar subcommand pairs it with pkg/tune/predictor/ (the analytical curve, with log10 routed through pkg/tune/pymath for libm parity) and pkg/tune/pyjson/ — not the pkg/predictor copy that predict uses or the pkg/pyjson copy that corpus uses; ADR-1125 records which consumer owns which copy.
  • pkg/pyjson/ — a CPython-compatible JSON encoder. The corpus JSONL and the sidecar --json payloads are cross-implementation artefacts, so the writer reproduces json.dumps byte-for-byte: bare NaN / Infinity tokens, repr()-style float rendering, and ensure_ascii escaping. pkg/corpus also ports CPython's Neumaier-compensated sum() and statistics.pstdev() so the aggregate columns match to the last bit.

The encoder-introspection subcommands add four more:

  • pkg/benchmark/ — corpus loading (LoadCorpusJSONL), per-encoder summarisation (Summarize) and the three renderers. No subprocess at all.
  • pkg/codecadapter/ — the argv-shaping half of vmaftune.codec_adapters: 19 codecs, each answering "what is the -c:v ... slice for this (preset, quality)?" so nothing else branches on codec identity.
  • pkg/encodeprofile/ — profile loading from JSON / HTML / Markdown, recommendation selection, EncodeRequest construction, FFmpeg argv composition and the encode driver (with an injectable Runner seam so tests never spawn ffmpeg).
  • pkg/pyjson/ — the same encoder, here rendering Go value trees byte-identically to CPython's json.dumps(..., indent=2, sort_keys=True) and jsonio.dumps_strict. Go's encoding/json differs on key ordering, HTML escaping, non-ASCII escaping and float formatting (float64(92) renders 92 in Go and 92.0 in CPython), so a shared encoder keeps the ported payloads diff-clean against the Python originals.

Stage 5 adds the auto / sidecar stack under pkg/tune/:

  • pkg/tune/auto/ — the Phase F decision tree: source probing, the ten short-circuit predicates, the recipe table, the confidence policy, winner selection, and the plan emitter.
  • pkg/tune/sidecar/ — the online-ridge bias-correction model, its Sherman-Morrison rank-1 update, and the cache-dir persistence layout.
  • pkg/tune/executor/ — --execute mode: the libvmaf CLI driver and the JSONL results log; its ffmpeg argv is pkg/ffencode's under the executor's name.
  • pkg/predictor/, pkg/codecadapter/, pkg/hdr/, pkg/pyjson/, pkg/pymath/ — the shared layers auto and sidecar consume: ShotFeatures and the per-codec analytical curve with the optional ONNX session and the PickCRF binary-search inversion; the codec-adapter registry (quality windows, probe knobs, preset vocabularies, per-encoder ffmpeg argv); HDR detection from ffprobe colour metadata plus the per-codec HDR flag dispatch; the CPython-compatible JSON emitter; and the correctly-rounded Exp2 / Log10 kernels that keep estimated_bitrate_kbps and estimated_vmaf on the platform-libm value CPython emits (Go's math.Pow / math.Log10 land a ULP away; the package docs record the measured residual). Each has exactly one implementation (ADR-1137); the pkg/tune/{predictor,codec,pyjson} paths survive only as thin aliases until the in-flight sidecar parity fix (#1187) lands, after which the sidecar imports move and the aliases are deleted.

See ADR-0705 for the migration rationale, ADR-0730 for Stage-2, ADR-0770 for Stage-4, and ADR-0702 for the Phase 4 umbrella.