Integrate MateProbe by Mate4B in an agent project¶
This guide targets alpha 0.1.0a4, Python 3.11+. Use it when an application has authoritative state and needs to test explicit declarations in generated outputs against that state. The check itself makes no model or network calls.
Already have a validator? Start with testing AI output validators
and the one-file audit. audit_validator accepts a sample-only
adapter and caller-authored baseline/variant pairs; you do not need to migrate to
the Document/Claim representation used in the state-check example below.
Install the alpha API¶
python -m venv .venv
. .venv/bin/activate
python -m pip install mateprobe==0.1.0a4 pytest-mateprobe==0.1.0a4
The second package supplies pytest and the automatically discovered mateprobe
fixture. Do not also load it with -p. Pin the alpha versions rather than assuming
an unqualified pip install selects a prerelease. See API reference.
Choose the source of authority¶
- Obtain state from trusted application code, a database projection, or a receipt.
- Validate the output schema and required fields. The Pydantic recipe illustrates this step; Pydantic is optional and is not a core dependency.
- Map an explicit allowlist of output fields to claims. Preserve JSON scalar types.
- Choose
state_refin trusted code. Never let the generated output pick whichever snapshot makes its own claims pass. Do not populate expected state from the output. - Evaluate the configured rules, inspect
acceptedandcomplete, and handle rejection in the application. Evaluation does not execute actions or enforce backend permissions.
RequiredFact only checks state. DeclaredClaimsConsistent checks supplied claims;
it does not ensure that every required output field was declared or that prose agrees.
A complete pytest example¶
Save as test_reply.py. This includes both a detected state mismatch and the
prose-only failure that remains outside the guarantee.
import pytest
from mateprobe import (
Claim,
Context,
DeclaredClaimsConsistent,
Document,
Surface,
evaluate,
)
def test_reply_contract(mateprobe):
# This snapshot is provided by the application, independently of the output.
context = Context({"current": {"ticket.status": "resolved"}})
rules = (DeclaredClaimsConsistent("ticket-claims", ("reply",)),)
def reply(status, text):
return Document(
(
Surface(
"reply",
text,
state_ref="current",
claims=(Claim("ticket.status", status),),
),
)
)
accepted = mateprobe.check(reply("resolved", "Your ticket is resolved."), context, rules)
assert accepted.accepted and accepted.complete
wrong = reply("open", "Your ticket remains open.")
rejected = evaluate(wrong, context, rules)
assert rejected.complete and not rejected.accepted
assert any(
check.rule_id == "ticket-claims"
and check.code == "CLAIM_STATE_MISMATCH"
and check.status.value == "violated"
for check in rejected.checks
)
# The fixture records the report and raises on rejection.
with pytest.raises(AssertionError):
mateprobe.check(wrong, context, rules)
# Correct declaration, contradictory prose: deliberately still accepted.
prose_gap = mateprobe.check(
reply("resolved", "Your ticket remains open."),
context,
rules,
)
assert prose_gap.accepted and prose_gap.complete
python -m pytest -q test_reply.py --mateprobe-report=contract-results.json
Expected: one passing test with three recorded evaluations. The prose challenge passes its test because the test explicitly asserts this documented limitation. Use a serial pytest run; distributed report merging is not supported.
Diagnose a result¶
| Result | Meaning | Application response |
|---|---|---|
satisfied |
The configured predicate holds | Continue only if the full report meets policy |
violated |
The configured predicate fails | Inspect rule, code, scope, and evidence |
undetermined |
Evidence or support is missing | Supply the missing information; do not call it success |
error |
Rule execution failed | Fix the rule/integration; do not count it as fault detection |
Default policy blocks violations, unknowns and errors. Heuristics can be advisory
with Policy(block_heuristics=False). complete only means no unknown/error result;
it does not mean every sentence or required business obligation was checked.
If schema validation fails before evaluation, report that stage separately
(for example, schema_valid=False, evaluation_ran=False, accepted=False).
Do not set an engine-style complete=True when no contract report was produced.
Audit the checks¶
Use the runnable mutation example and mutation semantics. Include accepted baselines, attributable faulty variants, valid controls, and label provenance. Report detection and preservation separately, with exclusions and missed faults visible.
API version¶
check_fields, CompareFields, AllowedTransition, and the independent
validator-audit API are included in 0.1.0a4. Older a2 wheels do not provide them.
See the current a4 API index and adoption levels. Pin the
package version rather than silently switching a consumer to main.
For free-form factual correctness or writing quality, this library alone is insufficient. See choosing an evaluator.