Skip to content

Exploratory testing pass — mutago CLI, 2026-09-26

Scope: three CLI journeys on the current remote default branch: coverage-aware reports and score gates, per-test filtering, and PR line filtering with dry-run and execution.

Build and setup

  • Built mutago dev from 4df2956, the origin/main head used for this pass.
  • Go 1.26.6 on macOS arm64.
  • Used two disposable Go modules under /tmp; the coverage fixture has three tested functions and one intentionally uncovered function. The Git fixture changes one production line and its matching assertion.
  • Mutation invocations used GOMAXPROCS=1, --workers=1, --exec-timeout=10, and a disposable GOCACHE.
  • Help, setup, fixture sources, command output, and generated reports are in evidence/.

Results

No confirmed bugs were found, so no issues were filed.

Coverage gate and reports

Goal: run a small package as a CI user, inspect covered and overall MSI, and keep machine-readable reports. The expected result was a successful run with internally consistent totals, uncovered code excluded from covered MSI, and all requested reports written.

Ran --coverage --quiet --no-diffs --workers=1 --exec-timeout=10 --logger-summary-json --logger-agentic-json --html-output ./calc against the isolated example.com/xtest/calc fixture. It returned 0 and reported 14 mutants: 10 killed, 4 not covered, no escapes or errors, 71.43% overall MSI, and 100% covered-code MSI. The summary JSON totals reconcile; the agentic JSON and HTML report were non-empty. Evidence: run output, summary JSON, agentic JSON, HTML report (gzip-compressed to preserve the generated bytes), and fixture.

Then exercised both gate outcomes. Floors of 71% overall and 100% covered passed with exit 0; raising the overall floor to 72% returned exit 4 and stated that 71.43% was below the threshold. Evidence: passing gate, failing gate.

Per-test filtering

Goal: reduce mutant test work while retaining the ordinary coverage result. The expected result was a per-test coverage map and the same mutant classifications as the non-per-test run.

Ran the same fixture with --coverage --per-test. Mutago reported a map for all three tests and returned the same totals and MSI as the ordinary run: 14 total, 10 killed, 4 not covered, 71.43% overall, and 100% covered. Evidence: per-test output. Fixture source hashes remained unchanged after the runs: hashes.

Changed-line filtering

Goal: check a PR-sized change before running the full suite. The expected result was that dry-run candidates and executed mutants would be limited to the changed production line.

In the isolated example.com/diff Git fixture, changed math.go:4 from n + 1 to n + 2 and updated the test expectation. With --git-diff-lines --git-diff-base=HEAD, dry-run listed four possible mutations, all at math.go:4; the real run killed four mutants, also all at line 4, and returned 0. Evidence: fixture diff, dry-run, real run, and fixture sources.

A first setup attempt used HEAD~1, which did not exist in the one-commit fixture. Replaying with set -o pipefail showed the documented tool-error exit 3 and a clear Cannot load git diff message. This was a fixture-reference error, not a product bug. Evidence: invalid-base replay.

Usability observations

  • The score summary clearly distinguishes overall MSI from covered-code MSI, and the compact JSON exposes flat totals that are easy to reconcile.
  • --per-test prints how many test entrypoints it mapped before the mutation summary.
  • Dry-run and real execution agreed on the changed source line in this fixture.

Unexplored areas and limits

Live TTY progress, custom --exec, baseline updates, annotation filters, and signal interruption were outside these three journeys. The fixtures are deliberately small and do not represent large or multi-package repositories. No product source was changed during the pass.

Cleanup

The disposable modules and cache were removed after evidence was copied. The report and evidence remain here.