How to test AI output validators in Python¶
You already validate structured responses from an LLM or agent. How do you know the validator catches the policy violations you care about, while accepting valid changes? MateProbe by Mate4B runs your validator against paired samples and reports what it detects, misses, or rejects for an unrelated reason.
The independent audit_validator API is available in published 0.1.0a4 for
Python 3.11+. It works with an adapter around your existing function; adopting the
library's built-in state contracts or Document/Claim representation is optional.
Start with a specific failure¶
Suppose an agent returns refund_completed=True. Your schema accepts the boolean,
but application policy also requires a completed backend receipt for the same
request. Test those obligations separately:
| Sample | Expected behavior |
|---|---|
| Completion claim with a matching completed receipt | Accept the baseline |
| Completion claim without supporting receipt evidence | Reject under the declared completion policy |
| Completion claim with another request's receipt | Reject for the receipt mismatch |
| Honest pending declaration when completion is unconfirmed | Preserve the valid control |
Your application supplies the authoritative receipts and justifies the labels. A missing receipt does not prove an operation never happened. The audit checks the specified evidence policy on these samples.
Run the one-file refund audit from PyPI without cloning this repo. It includes a prose-only contradiction that still passes, so the known limit remains visible. No API key or model is needed for that example.
Why isn't any rejection enough?¶
If a case expects customer_mismatch but gets malformed_input, a generic
assert not validate(sample) passes. The audit records an unattributed
rejection because the intended finding was not observed. A crash is an execution
error; missing evidence can be reported as an incomplete verdict. Neither is
credited as a targeted detection.
A valid control checks the other direction: a wording change or another permitted value must stay accepted. Rejecting every input cannot establish useful coverage.
Input audits, source mutation testing, and pytest¶
| Approach | What changes | Question it answers |
|---|---|---|
| This validator audit | Caller-supplied input samples, with justified fault/control labels | Does the validator detect these faults and preserve these valid cases? |
| Source mutation testing | Program source, with tests run against altered implementations | Do the tests detect these changes to the implementation? |
| Ordinary pytest assertions | Whatever fixtures and assertions you write | Does the code satisfy these explicit test expectations? |
These approaches can be used together. MateProbe by Mate4B adds paired execution, finding attribution, obligation inventories, and JSON/Markdown reports to a supplied corpus. It does not edit validator source or automatically invent and label mutations. Every individual expectation can also be expressed directly in pytest.
Connect your validator and keep a regression¶
- Wrap its actual result in
Verdict; preserve its finding IDs and completeness. A boolean-only adapter is supported, but a rejection alone cannot identify the cause. - Define an
ObligationandAuditCasepairs: an accepted baseline, justified invalid variants, and valid controls. Expected findings belong to cases, never inside the adapter's output. - Call
audit_validator, inspect individual outcomes, and retain survivors and errors. - Use the optional
mateprobe.audit_validatorpytest fixture to set thresholds and write--mateprobe-report=contract-results.json. Assert important case outcomes so a new miss cannot silently replace an acknowledged gap.
See the adapter API, real Pydantic and JSON Schema adapters, and the regression walkthrough for executable integrations.
Can this replace an LLM judge?¶
It can audit checks whose behavior is explicitly defined by your policy and cases, without introducing a judge. The built-in checks run offline; an arbitrary wrapped validator retains its own dependencies and behavior. Reproducible results require a deterministic validator and fixed inputs/configuration.
It cannot certify unrestricted prose as true or helpful. A mutation score measures the supplied corpus, not production accuracy or total policy coverage. Use the evaluation selection guide to choose checks for the properties you actually need. The a3 integration observation shows one agent implementing a fresh inventory audit, with its limits documented.