Controlled validation · Synthetic DEV record
Evidence to one approved update.
Documentary context and a governed action, each checked in its own session.
Knowledge retrieval and Gateway execution were validated separately. This case joins the observations into a reviewable method; it is not a transcript of one continuous autonomous agent run.
Observed outcome
One narrow change, then verification.
In a private, human-approved validation, Gateway read one authorised synthetic DEV record and found one non-sensitive demonstration field empty. It prepared a temporary, non-mutating plan for one fictional change. A human reviewed the exact proposal and approved that same pending plan in a later message. Gateway applied it once; a bounded read confirmed the proposed state.
The human later cleared the demonstration field through the authorised DEV interface, and a separate bounded read confirmed the initial empty state. The MCP validation made no QA or PROD write. This account omits private mapping, values, identifiers and raw results.
Capability boundaries
Different evidence for different claims.
Documentary context · Knowledge
An isolated local session retrieved version-tagged official CMDB evidence and a separately selected authorised project procedure. The chosen matches were lexical and their sections were expanded. The project procedure supplied local context; neither documentary source proved the live record's state or authorised a write.
Live observation · Gateway
In a separate controlled DEV validation, Gateway selected the environment explicitly, checked readiness and read exactly one authorised synthetic record. The bounded result established its then-current state.
Plan and human review
Gateway policy admitted only the narrow demonstration update. Planning stored one proposal with a digest, expiry and current-state precondition without changing Helix. The exact proposal was reviewed by a human; later approval applied to the same unchanged pending plan.
Apply, verify and clean up
One apply call returned a known applied result and an immediate bounded read confirmed the state. Sanitised audit metadata recorded operation outcomes, while the reviewed tool result and read established value equality. Human cleanup was separate from the MCP write.
What to repeat
Review the exact plan before acting.
The method requires an installation's own authorised synthetic DEV record, private project procedure, matching product/version context, narrowly scoped Gateway policy and a human reviewer. Confirm the live target's version before using version-specific evidence. Stop if the read is ambiguous or unexpected; stop if the plan changes or expires. Investigate an uncertain write outcome instead of retrying automatically.
The check on 15 September 2026 used Knowledge 1.31.5 and Gateway 0.10.3. Separate hermetic tests rejected expired plans and mismatched digests before a fake client write; those paths were not attempted as live DEV writes. This validation does not benchmark continuous agent autonomy, demonstrate a concurrent-change conflict or establish production suitability.
For a fully public, fictional Knowledge-only first answer, read the source-grounded answer case.
For the later uninterrupted two-server read path that deliberately stopped before planning, read the integrated read-only preflight case.
