Turn a survivor into a regression check

This guide shows how to preserve a real audit failure while fixing the validator that caused it. It uses the audit_validator API from mateprobe==0.1.0a4; it does not require replacing the validator or adding a new policy language. The example is self-contained and uses no private application data.

1. Keep the survivor and a valid control

Suppose refunds are allowed only in USD and only up to the supplied limit. The existing validator checks the amount but forgets the currency rule:

from mateprobe.mutations import Relation, Validity
from mateprobe.validator_audit import (
    AuditCase,
    Obligation,
    Verdict,
    audit_validator,
)


def validate_refund_before(data):
    violations = ()
    if data["amount_cents"] > data["limit_cents"]:
        violations = ("amount_exceeded",)
    return Verdict(accepted=not violations, violations=violations)


baseline = {"amount_cents": 5000, "limit_cents": 6000, "currency": "USD"}
cases = (
    AuditCase(
        id="over-limit",
        obligation="refund-policy",
        baseline=baseline,
        variant={"amount_cents": 6001, "limit_cents": 6000, "currency": "USD"},
        relation=Relation.VIOLATION,
        expected=("amount_exceeded",),
        validity=Validity.VALID,
        provenance="Reviewed policy case: 6001 exceeds the limit of 6000.",
    ),
    AuditCase(
        id="unsupported-currency",
        obligation="refund-policy",
        baseline=baseline,
        variant={"amount_cents": 5000, "limit_cents": 6000, "currency": "EUR"},
        relation=Relation.VIOLATION,
        expected=("currency_not_supported",),
        validity=Validity.VALID,
        provenance="Reviewed policy case: refunds are limited to USD.",
    ),
    AuditCase(
        id="at-limit-usd",
        obligation="refund-policy",
        baseline=baseline,
        variant={"amount_cents": 6000, "limit_cents": 6000, "currency": "USD"},
        relation=Relation.PRESERVE,
        validity=Validity.VALID,
        provenance="Reviewed control: the inclusive USD limit permits 6000.",
    ),
)
obligations = (Obligation("refund-policy", "Refunds must be USD and no greater than the limit."),)

before = audit_validator(
    validate_refund_before,
    cases,
    obligations=obligations,
    validator_id="refund-policy/1",
)
assert {case.id: case.outcome for case in before.cases} == {
    "over-limit": "detected",
    "unsupported-currency": "survived",
    "at-limit-usd": "preserved",
}

unsupported-currency is the survivor: the authored invalid variant was accepted. over-limit is a real detected fault, and at-limit-usd is the preservation control. Keep both when fixing the validator; a reject-everything change must not make the audit pass.

2. Fix the validator, then run the same corpus

Add the missing blocking finding to the validator. The adapter still receives only the sample; it must not receive case.expected or copy labels into the result.

def validate_refund_after(data):
    violations = []
    if data["amount_cents"] > data["limit_cents"]:
        violations.append("amount_exceeded")
    if data["currency"] != "USD":
        violations.append("currency_not_supported")
    return Verdict(accepted=not violations, violations=tuple(violations))


after = audit_validator(
    validate_refund_after,
    cases,
    obligations=obligations,
    validator_id="refund-policy/2",
)
assert {case.id: case.outcome for case in after.cases} == {
    "over-limit": "detected",
    "unsupported-currency": "detected",
    "at-limit-usd": "preserved",
}
after.assert_thresholds()
print(after.to_markdown())

The AuditCase labels and rationales are the user's responsibility: review the baseline, decide whether each variant is a violation or preservation, name the obligation, and provide the expected stable finding IDs. The user also owns the adapter's mapping from real validator output to Verdict, the scope convention for IDs, and any domain authorization or backend checks.

The library's responsibility is narrower: isolate and run paired inputs, verify the baseline, compare the returned Verdict with the authored expectation, classify outcomes such as survived, detected, preserved, regressed, unattributed_rejection, undetermined, and error, and retain those results in the report. It does not infer truth labels, inspect arbitrary prose, prove business correctness, or turn an exception into a detection. A rejection with the wrong or missing finding ID is unattributed_rejection.

3. Keep it in pytest and CI

Save the Python blocks above and the following test in tests/test_refund_policy.py. In your application, import the actual fixed validator instead of copying this illustrative implementation. Keep the cases unchanged when revising the validator.

def test_refund_policy(mateprobe):
    mateprobe.audit_validator(
        validate_refund_after,
        cases,
        obligations=obligations,
        validator_id="refund-policy/2",
    )
python -m pytest tests/test_refund_policy.py --mateprobe-report=refund-audit.json

The default gate requires targeted detection of every eligible fault and preservation of every eligible control, and fails on broken baselines. Removing the currency check makes this test fail again; replacing the validator with reject-all also fails. The plugin retains its JSON report on threshold failure. These thresholds cover this corpus, not every possible policy input.

4. Share sanitized feedback

Use this template when reporting a useful survivor outside the private project. Replace placeholders with public, synthetic, or redacted values. Do not paste raw samples, customer identifiers, tokens, receipt contents, stack traces, or evidence strings that contain application data. The report stores digests rather than full inputs, but evidence and exception messages still require review.

Title: survivor in <public obligation or rule name>

Library checkout/version: <commit or local version>
Validator adapter version: <public identifier>
Audit validator ID: <public validator_id>

Expected policy (sanitized): <one sentence>
Case ID: <stable public case id>
Relation: VIOLATION or PRESERVE
Validity: VALID
Expected finding IDs: <public IDs only>
Case rationale: <why the baseline and variant are justified>

Observed outcome: survived | detected | preserved | regressed | ...
Observed public finding IDs: <IDs, or none>
Reproduction: <minimal synthetic input or pseudocode with private values removed>
Fix status: <unfixed / fixed locally / fixed in public change>
Regression result after fix: <outcome and control outcome>

Include a sanitized to_markdown() excerpt only after checking evidence and exception text. A survivor report is a focused reproducible case, not a claim of coverage or a claim that the library labels cases automatically.