Exploratory testing pass — mutago CLI, 2026-09-26¶
Scope: three CLI journeys on the current remote default branch: coverage-aware reports and score gates, per-test filtering, and PR line filtering with dry-run and execution.
Build and setup¶
- Built
mutago devfrom4df2956, theorigin/mainhead used for this pass. - Go 1.26.6 on macOS arm64.
- Used two disposable Go modules under
/tmp; the coverage fixture has three tested functions and one intentionally uncovered function. The Git fixture changes one production line and its matching assertion. - Mutation invocations used
GOMAXPROCS=1,--workers=1,--exec-timeout=10, and a disposableGOCACHE. - Help, setup, fixture sources, command output, and generated reports are in
evidence/.
Results¶
No confirmed bugs were found, so no issues were filed.
Coverage gate and reports¶
Goal: run a small package as a CI user, inspect covered and overall MSI, and keep machine-readable reports. The expected result was a successful run with internally consistent totals, uncovered code excluded from covered MSI, and all requested reports written.
Ran --coverage --quiet --no-diffs --workers=1 --exec-timeout=10 --logger-summary-json --logger-agentic-json --html-output ./calc against the isolated example.com/xtest/calc fixture. It returned 0 and reported 14 mutants: 10 killed, 4 not covered, no escapes or errors, 71.43% overall MSI, and 100% covered-code MSI. The summary JSON totals reconcile; the agentic JSON and HTML report were non-empty. Evidence: run output, summary JSON, agentic JSON, HTML report (gzip-compressed to preserve the generated bytes), and fixture.
Then exercised both gate outcomes. Floors of 71% overall and 100% covered passed with exit 0; raising the overall floor to 72% returned exit 4 and stated that 71.43% was below the threshold. Evidence: passing gate, failing gate.
Per-test filtering¶
Goal: reduce mutant test work while retaining the ordinary coverage result. The expected result was a per-test coverage map and the same mutant classifications as the non-per-test run.
Ran the same fixture with --coverage --per-test. Mutago reported a map for all three tests and returned the same totals and MSI as the ordinary run: 14 total, 10 killed, 4 not covered, 71.43% overall, and 100% covered. Evidence: per-test output. Fixture source hashes remained unchanged after the runs: hashes.
Changed-line filtering¶
Goal: check a PR-sized change before running the full suite. The expected result was that dry-run candidates and executed mutants would be limited to the changed production line.
In the isolated example.com/diff Git fixture, changed math.go:4 from n + 1 to n + 2 and updated the test expectation. With --git-diff-lines --git-diff-base=HEAD, dry-run listed four possible mutations, all at math.go:4; the real run killed four mutants, also all at line 4, and returned 0. Evidence: fixture diff, dry-run, real run, and fixture sources.
A first setup attempt used HEAD~1, which did not exist in the one-commit fixture. Replaying with set -o pipefail showed the documented tool-error exit 3 and a clear Cannot load git diff message. This was a fixture-reference error, not a product bug. Evidence: invalid-base replay.
Usability observations¶
- The score summary clearly distinguishes overall MSI from covered-code MSI, and the compact JSON exposes flat totals that are easy to reconcile.
--per-testprints how many test entrypoints it mapped before the mutation summary.- Dry-run and real execution agreed on the changed source line in this fixture.
Unexplored areas and limits¶
Live TTY progress, custom --exec, baseline updates, annotation filters, and signal interruption were outside these three journeys. The fixtures are deliberately small and do not represent large or multi-package repositories. No product source was changed during the pass.
Cleanup¶
The disposable modules and cache were removed after evidence was copied. The report and evidence remain here.