Choosing an evaluation method

Use MateProbe by Mate4B when the application can express a checkable predicate over trusted state, explicit declarations, or a deliberately narrow text pattern. It evaluates those predicates without inference calls. This is a scoped guarantee, not general factuality or writing-quality evaluation.

Which check answers your question?

Question Appropriate mechanism Remaining boundary
Are the output fields present and correctly typed? A schema validator, such as strict Pydantic A schema-valid value may still contradict application state
Does the declared refund status match the trusted result? A state/declaration contract State provenance and field completeness belong to the application
Does the text contain a configured forbidden phrase? A lexical rule It does not interpret quotation, negation or semantic paraphrases
Does unrestricted prose contradict the database? A separately validated extraction/entailment method, or constrained rendering Supplied claims alone do not establish what the prose says
Is the response helpful or stylistically appropriate? A task-specific evaluation protocol, potentially including model or human judgments The protocol needs its own reliability evidence
Does my Python validator catch invalid AI outputs and preserve valid ones? A validator audit with paired input samples, fault targets and controls Results depend on the validity and coverage of the authored cases

These mechanisms can coexist. For example, validate the JSON schema, evaluate exact state predicates, then route selected prose properties to another evaluator. Keep the resulting evidence separate rather than treating every pass as proof of truth.

In this audit, mutations are supplied input variants. Source mutation testing instead changes program source to test whether the test suite catches altered behavior. See how to test AI output validators for the distinction and a workflow for an existing validator.

Deterministic contracts and LLM judges

For this library's configured offline rules, identical inputs and implementations produce the same report. Evaluation uses local compute but no inference tokens or provider round trip. This does not mean zero latency or zero operating cost.

An LLM judge can be asked about properties that are difficult to encode as explicit predicates. Its usefulness must be established for the task, model and evaluation protocol. This project has not run a controlled head-to-head study and claims no blanket superiority in accuracy, coverage, speed or cost.

Guardrail frameworks can contain several kinds of checks; compare actual configured validators and versions, rather than assuming every framework uses a model judge. This page is a selection guide, not a product ranking.

What our evidence establishes

The captured-output mutation audit includes 32 model responses and 384 authored variants. Exact declarations, lexical faults, schema faults and preservation controls are reported separately. All 64 inserted prose-only contradictions pass when the supplied declarations stay correct. Related variants are not independent natural-output samples.

This supports reproducible testing of the configured checks on the published corpus. It does not establish natural semantic accuracy, independent validation, production adoption, or better performance than a model judge.

Start with the agent guide or Pydantic recipe if the state/declaration boundary matches your application.