Agent integration trial

Historical evidence/reference from Narrative Contracts. The current project is MateProbe by Mate4B; see the migration guide. Original package names, versions and recorded results below identify that earlier work.

This page preserves the historical a2 trial. See the separate discovery and a3 integration observations for the newer tasks.

We tested whether an agent could build a small integration using the public documentation and published packages, without inheriting the implementation conversation. The task and criteria were fixed first.

One gpt-5.6-luna agent started at llms.txt and implemented a shipping-assistant validator using narrative-contracts==0.1.0a2, its pytest plugin, and strict Pydantic validation. The application selected the trusted branch; generated fields could not choose another snapshot. Core library and docs source access were prohibited by instruction, not by an OS sandbox.

Observed results

Check Observation
Installable API Published core/plugin a2; no unreleased imports
Consumer's tests 8 passed
Additional maintainer checks 16 passed against unchanged consumer code
Wrong declaration Attributable CLAIM_STATE_MISMATCH
Missing state Rejected, incomplete/undetermined report
Invalid types or branch injection Schema rejection
Correct declarations with contradictory prose Accepted, with the limitation explained

The 24 tests exercise related cases in one prompted task. They are not 24 independent integrations or an agent success-rate estimate. The trial does not test whether a model recommends the library without being told about it.

Problems retained

The agent initially invoked pytest before creating its test file; that exit 4 is retained in the evidence. It also saved only llms.txt from the requested documentation snapshots, so the maintainer recovered the reported public pages and actual site build metadata afterward. We do not count artifact capture as fully completed by the agent.

The consumer wrapper labels schema failures complete=true although the core engine was not evaluated. Rejection remains accepted=false, but the field is ambiguous. The integration guide now explains how to distinguish schema results from engine completeness. The original consumer code remains unchanged in the evidence.

Reproduce

The frozen evidence and replay instructions include the source, reports, docs revision, hashes and maintainer checks. CI runs the replay against installed PyPI packages. Replaying code tests its behavior; it does not rerun the agent or demonstrate general discoverability.

This is maintainer-run usability evidence, separate from the real-output mutation audit and from independent human evaluation. Search visibility, third-party adoption and unprompted recommendations remain unmeasured.