Expanded mutation campaign

benchmarks/expanded_mutations.py is a second, bounded output-mutation campaign for the alpha release. It is deliberately separate from the original synthetic benchmark and from benchmarks/mutate_guards.py, which mutates source copies. The campaign uses the public MutationCase/audit() API and writes deterministic JSON plus a short Markdown summary to benchmarks/expanded-results/.

Run it from the repository root with the development interpreter:

./.venv/bin/python benchmarks/expanded_mutations.py

To compare two runs without overwriting the checked-in artifact directory:

one=$(mktemp -d)
two=$(mktemp -d)
./.venv/bin/python benchmarks/expanded_mutations.py --output "$one"
./.venv/bin/python benchmarks/expanded_mutations.py --output "$two"
cmp "$one/campaign.json" "$two/campaign.json"
cmp "$one/corpus.json" "$two/corpus.json"
cmp "$one/summary.json" "$two/summary.json"
cmp "$one/summary.md" "$two/summary.md"

The corpus is authored with seed 20260929 and contains 21 paired cases. State invariants and text heuristics are reported independently. The current checked-in run has 5 valid invariant faults (4 detected) and 8 valid heuristic faults (4 detected). It also has 5 valid preservation controls (4 preserved); the negated formula control is retained as a heuristic false positive. Exact branch-claim mismatches and typed fact changes are attributed to invariant rules, while long lexical copies and literal forbidden phrases exercise the heuristics.

Several cases are intentional guarantee-boundary challenges. Undeclared prose contradictions, paraphrased forbidden phrases or premises, semantic restatements, and paraphrased history are labelled from transparent operator intent and remain visible as survivors when the configured rule cannot establish the claim. A missing state snapshot is explicitly excluded because the rule reports undetermined, not a violated check. Unicode normalization is an explicit equivalent exclusion, and arbitrary prose meaning is an explicit not-applicable exclusion. audit() never infers these labels from its own output.

These numbers are construction checks for the configured synthetic cases. They are not natural-output semantic-accuracy estimates, independent annotation results, or evidence of generalization to model output.