Freezing the tool catalogue

N17Q pinned server identity, tool schemas, descriptions, mappings, and the exact offered subset so replay and approval referred to the capabilities a run actually saw.

I replayed an N17Q run against an MCP server with one familiar tool name and a different schema.

The historical trace said the agent had called search. The live server still exposed search. One required field had changed, the description now encouraged broader use, and result identity behaved differently.

A name was not enough to reconstruct what the run had seen.

MCP discovery needed a durable registry snapshot between server catalogue and agent context.

Discovery observed one moment

When N17Q connected to an MCP server, it recorded local connection identity, observed server metadata, protocol version, tools and resources, schemas, descriptions, and retrieval time.

The snapshot represented what one configured connection reported then. It did not claim the server would remain stable or truthful.

Later discovery produced another snapshot and a diff. Old runs retained their original reference.

The catalogue became evidence instead of ambient current state.

A server-reported name could collide or change. N17Q assigned a registry identity to the configured transport, installation provenance where known, endpoint or executable evidence, credential scope, owner, and review state.

Remote metadata remained inside that record. High-consequence mappings could pin stronger implementation evidence.

No identity proved a remote server safe. The local anchor prevented a self-described replacement from inheriting an old mapping automatically.

Human names stayed readable while authority used exact references.

Schema and description had separate digests

A type change could break parsing. A prose change could alter model selection without changing execution shape.

N17Q stored canonical schema and description digests separately, with raw reviewed material under retention. Registry diffs classified added, removed, required, optional, enum, default, and description changes.

Meaning-bearing differences moved a mapping to review required. Cosmetic changes could follow an explicit lighter path.

Tool language entered version control because it shaped behavior.

An MCP server could advertise many capabilities. N17Q did not pass them directly to the model.

Candidates appeared in an integration review surface. A reviewed adapter mapped one candidate into a narrow product capability with effect class, normalization, authority, data policy, limits, idempotency, query, compensation, and replay strategy.

Unknown candidates stayed disabled. Connecting a server did not grant every tool to every run.

Interoperability ended where product activation began.

The mapping had its own revision

The same remote schema could map differently as N17Q learned. A broad path argument might become an opaque workspace handle. A remote success might require local receipt verification.

Every product mapping received revision and digest. Requests, policy, approval, execution, fixtures, and traces referenced it.

A server snapshot and mapping snapshot together explained both what was offered and what N17Q believed it meant.

Neither remote schema nor local adapter alone told the complete story.

Registry snapshot could contain twenty tools while one run received two. N17Q recorded the compiled catalogue for each model attempt.

The manifest named product capability, local description, normalized schema, mapping revision, policy context, and relevant remote candidate. Disabled and irrelevant tools remained absent.

This let evaluation answer whether the model selected poorly among offered choices or whether orchestration exposed too much.

Discovery and disclosure were distinct events.

Tool order and descriptions entered context provenance

Selection could change when tools were reordered or descriptions shortened to fit context. N17Q stored the actual compiled representation and its digest.

Provider adapters could transform names or schemas to meet platform constraints. Those transformations appeared in the manifest.

A replay did not assume the registry's ideal representation was what the model had read.

Context provenance extended down to capability vocabulary.

An MCP resource URI could return changed content. N17Q source observations added connection, server, URI, content digest, retrieval time, media type, and provider revision metadata where available.

The registry snapshot established discovery identity. Each read established content identity. Existing claims remained attached to the earlier revision.

The protocol could locate context; the evidence graph made it durable.

Discovery alone never certified source truth or freshness.

Prompts and metadata remained untrusted

Server-supplied text could contain instructions directed at the model. N17Q treated names, descriptions, resource contents, and errors as untrusted metadata or data.

Local capability descriptions were reviewed. Tool availability came from policy. Source content entered quoted observations. Unknown fields and hostile defaults could not broaden arguments.

The snapshot preserved what the server said without giving the statement authority.

Provenance made untrusted text inspectable, not safe by declaration.

If discovery found a changed schema or server identity, affected mappings entered review required. New calls stopped.

In-flight attempts retained their starting snapshot. Unknown effects followed the historical contract for reconciliation and could require manual inspection if the current server no longer supported query.

Pending approvals did not flow through the new mapping. Compatible changes required an explicit rule.

Connection refresh became a migration boundary.

Replay never rediscovered live state

A historical run used its recorded registry and fixture manifest. Replay mode had no live MCP connection or credentials.

The model received the original offered subset or a declared counterfactual variant. Tool calls resolved into deterministic fixture adapters. A missing schema or result stopped.

This prevented a server update from changing yesterday's benchmark and prevented replay from causing today's effects.

Registry snapshots made isolation possible but did not replace fixture behavior.

To test a clearer description or smaller tool set, N17Q branched from a checkpoint with a new compiled-catalogue digest.

The branch manifest named the change. The original trace remained. If the new tool schema required different requests, fixture coverage had to exist.

Results could show selection and budget differences without claiming the model had originally seen the improved catalogue.

Tool-interface experiments became explicit rather than retroactive edits.

Approval referred to product and remote identity

The reviewer saw the product consequence: exact artifact, destination, effect, and recovery. The receipt also pinned mapping and registry snapshot.

If the remote schema or local mapping changed meaning, approval invalidated. The UI did not ask the person to interpret protocol JSON, but technical evidence remained available.

Human consent stayed stable at the product layer and exact at the integration layer.

The snapshot prevented a familiar tool name from carrying old approval into new behavior.

A server could change error codes or result shape without changing the request schema. N17Q contract fixtures tied mappings to the registry and adapter revision.

Unknown current responses remained unknown. They did not inherit retry semantics from a superficially similar historical error.

Replay returned recorded adapter events and product outcomes. A counterfactual mapping change created a new contract.

Recovery depended on observed behavior, not protocol familiarity.

Credentials stayed outside the snapshot

The registry recorded safe credential identity, scope, and configuration revision. It never stored secret values in context or export.

Replay used synthetic or no credentials. A current run resolved credentials after policy. Rotating one did not rewrite historical registry evidence or broaden old approval.

A catalogue could show a capability the current credential could not exercise; eligibility still required both sides.

Discovery and authentication remained separate inputs.

Two MCP servers exposed search. Their schemas overlapped and their source semantics did not.

N17Q kept server registry identity on every candidate and mapping. The model saw distinct local product capabilities only when both were relevant. A fallback re-evaluated data policy and behavior rather than swapping names silently.

The trace could follow one result back to one server snapshot and resource revision.

Composition became possible without alias confusion.

Registry exports supported audit

A run package included safe registry snapshot, offered catalogue manifest, mapping contracts, schema and description digests, and fixture references. It excluded credentials and unapproved raw content.

A reviewer could inspect which capabilities existed, what the model saw, and what product policy allowed without reconnecting to the server.

Historical schema artifacts followed retention and license policy. Missing optional material appeared as tombstones.

The integration boundary remained explainable after removal.

A tool schema change affected invocation and effect semantics. A resource-list change affected what context could be selected. N17Q tracked both in one snapshot and reviewed them through different policies.

New resources did not enter every search scope automatically. Removed resources made current retrieval unavailable while historical source observations remained. Changed resource metadata could alter classification or selection and required review where relevant.

The registry preserved a common server boundary without pretending every discovered object carried the same consequence.

Connection health had a snapshot too

A catalogue could be valid while the server was unreachable, partially degraded, or failing one capability. N17Q recorded health observations separately from discovery identity.

Eligibility checked the current route needed by the run. A source read could remain available while a tool mapping paused. Rate limits scheduled later consideration rather than mutating the registry.

Historical traces retained what was known at execution. A green discovery event was not permanent proof of service health.

Identity, compatibility, and availability stayed distinct.

Request schema alone did not describe output size, media types, resource identity, error mapping, idempotency, or query behavior. The local adapter contract recorded those assumptions and fixtures under its revision.

A server returning a newly shaped result caused a typed incompatibility. Unknown content did not flow directly into model context or evidence graphs.

The registry view linked request schema, result observations, and behavioral contract.

A tool remained usable only while both sides of the mapping held.

Data policy shaped eligibility

One MCP server could be suitable for public documentation and ineligible for private drafts. Registry connection state named data classes, approved destinations, credential scope, retention assumptions, and local or remote execution characteristics.

The run compiler evaluated selected inputs before offering a capability. A fallback to another server repeated that decision.

The snapshot preserved which policy metadata was current without storing secrets.

Interoperable transport did not make data interchangeable.

Sending every discovered schema and description to a model consumed tokens and increased accidental selection. N17Q chose a small capability set after task planning and policy.

The compiler recorded why each tool was included, excluded, or deferred. Description truncation and provider-specific name transformation entered the offered manifest.

If required tool context could not fit, the request failed or the task was decomposed. Required effect warnings were never truncated silently.

Progressive disclosure applied to capabilities as strongly as sources.

Discovery order was not priority

Servers could return tools in implementation order, and that order could influence a model. N17Q canonicalized the registry for comparison and deliberately ordered the offered set by product relevance and stable rules.

The actual order entered context provenance. A counterfactual could change it and measure selection behavior.

No remote server gained prominence by listing a broad tool first.

Presentation was part of the adapter surface and remained local.

Opaque resource handles bound run, requester, server mapping, resource identity, action, and expiry. A handle from one workspace or registry snapshot failed in another.

The model did not receive arbitrary tenant, repository, or destination identifiers merely because the remote schema accepted strings. Resolution rechecked current policy immediately before execution.

The snapshot made the mapping exact; handles made each use narrow.

Catalogue visibility never became cross-workspace authority.

Registry updates had rollout states

A reviewed mapping could move through candidate, fixture-tested, limited, active, degraded, review required, retiring, and historical. New runs selected only eligible states.

Limited rollout used synthetic or non-sensitive scenarios first. Meaning-bearing failures paused expansion. In-flight intents followed pinned recovery contracts.

This avoided one binary enabled flag for capabilities whose schemas and behavior evolved independently.

Lifecycle became visible enough to operate rather than assumed from connection success.

An MCP server could conform to the protocol and still be unsuitable for N17Q's source identity, privacy, output bounds, or effect-recovery needs.

Protocol tests established message compatibility. Product fixtures established the behavior one mapping relied on. Policy established whether one run could use it.

The registry displayed all three evidence classes without compressing them into a Trusted badge.

Standards reduced integration variation; product fitness remained local judgment.

Registry corrections preserved what the model saw

I once misclassified a result as safely retryable. Correcting the contract created a new mapping revision and marked the earlier one erroneous.

Historical runs retained the old offered description and policy evidence so their behavior remained explainable. Re-evaluation could flag the risk under current understanding. Pending work stopped before another attempt.

The snapshot was not edited to make the past look better.

Provenance included mistakes in the integration layer.

The review interface connected four questions

For each mapping, N17Q showed: what did the server advertise, how did the product normalize it, what behavior did fixtures establish, and where could current policy use it?

Schema diffs, descriptions, effect class, credentials, data handling, limits, and replay coverage sat beneath those questions. Technical raw material remained available without being the only explanation.

This helped a reviewer see why a syntactically compatible update still required work.

The registry was a product boundary, not an inventory spreadsheet.

Removal closed every capability path

Deleting a connection removed it from discovery, eligibility, offered catalogues, scheduled retries, credential resolution, and new recovery actions. Existing unknown effects reached known or manual states first.

Historical snapshots and fixtures remained under retention for audit and replay. A stale checkpoint could not rediscover a same-named server and resume automatically.

Retirement was more than hiding a row in settings.

The registry tracked capability lifecycle from candidate to historical evidence.

Large catalogues and descriptions could bloat every run. N17Q content-addressed repeated schemas and stored references in the run manifest.

The compiled offered subset remained small and directly inspectable. Retention could deduplicate identical snapshots across runs while preserving each observation time.

Performance optimization never replaced identity with “latest.”

Historical precision did not require copying the entire server catalogue into every event.

Diffs focused review on consequence

The registry UI grouped schema shape, description, effect contract, scope, data handling, retry, and result changes. A one-word broader purpose could be prominent even when the JSON remained stable.

Reviewers could run contract fixtures before activation. High-consequence tools required stronger evidence. Unknown behavior kept the mapping disabled.

The system supported saying no to a valid server tool.

Interoperability increased choice; review preserved judgment.

N17Q could poll or refresh discovery under connection policy, but a long run did not adopt every new snapshot as soon as it appeared.

The registry compared current observation with the run's pinned catalogue and classified impact. Unrelated additions could wait for a later run. Removal or unsafe change affecting a pending capability paused it. Stable mappings could continue under explicit compatibility rules.

This prevented catalogue churn from perturbing an agent halfway through a controlled scenario. It also prevented pinning from concealing an urgent retirement.

Fresh observation informed a deliberate migration decision.

Snapshot access followed least privilege

Raw schemas and descriptions could reveal internal resource names or capabilities unavailable to ordinary run owners. N17Q exported only the subset and metadata the requester was authorized to inspect.

Protected reviewers could access full integration evidence. Models received the compiled catalogue. Historical traces used safe digests and local names where raw detail had expired or was restricted.

Provenance did not require exposing the entire server inventory to every audience.

With the historical snapshot pinned, the run loaded the old search contract and matching fixtures. With the live snapshot selected as a counterfactual, the adapter reported schema incompatibility until I reviewed and versioned a new mapping.

Neither path pretended the two same-named tools were identical.

The failure was inconvenient and exactly the evidence the registry had been missing.

Discovery is a fact about a boundary

MCP gives clients and servers a shared way to expose capabilities and context. In a long-running evaluated system, the moment of discovery becomes part of provenance.

N17Q preserved server identity, schemas, descriptions, mappings, resources, and the offered subset under stable digests. Policy and effect contracts still decided authority. Fixtures still decided replay.

A live catalogue answers what the server says now.

A registry snapshot answers what this run actually encountered—and lets every later approval, trace, and experiment refer to the same thing.