When the model changed mid-run

N17Q treated model handoff as a state transition, preserving exact workspace, evidence, authority, budgets, and unresolved effects while letting a different model continue without inheriting fictional continuity.

The second model finished the task and could not explain why the first model had stopped.

N17Q had switched after a context limit. The summary said a patch was ready, approval was pending, and one tool had failed. It omitted that the patch had changed after verification and that the failed call might already have created an external resource.

The new model read a coherent story and resumed from an incoherent state.

Changing models was not merely replacing the next text generator. It was transferring control of a durable workflow.

The run belonged to the product

The early architecture treated the provider conversation as the run. Its item IDs, context history, and continuation token became the main thread of identity.

That worked until a provider changed, context expired, or a model lacked the same continuation mechanism. Product state had to be squeezed into a summary and reintroduced as prose.

N17Q inverted the relationship. The run owned tasks, checkpoints, artifacts, decisions, effects, and evidence. A provider session was one execution resource used for one or more turns.

The model could change because the workflow no longer lived inside its private conversation.

A model switch could not race with an active tool or a mutating workspace. N17Q first reached a checkpoint boundary: current turn finalized or interrupted, tool attempts classified, writes sealed, effect state durable, and context inputs identified.

If an external attempt remained in flight, the handoff recorded its send boundary and recovery contract. The successor did not receive a clean slate. If the workspace still had active processes, the supervisor settled or terminated them before snapshotting.

The checkpoint was an immutable parent for whatever the next model proposed.

Transfer started from known state, not from the moment a token budget happened to end.

The summary was a projection

N17Q still produced a compact narrative because models reasoned better with one. It named goal, completed work, decisions, blockers, active plan, and important evidence.

Every assertion linked to structured checkpoint objects. The context compiler added exact fields for approvals, effects, target revision, verification eligibility, budgets, and available capabilities. Policy consumed those fields directly.

If the summary omitted a stale test result, the verification record still prevented a completion claim from becoming valid. If it described an approval loosely, the final gate still required the exact receipt.

Prose aided continuity without carrying authority.

Model identity became an event

Each turn recorded provider, model, relevant configuration, context package digest, response protocol, and capability catalogue. A switch emitted a handoff event with reason: cost routing, capability need, availability, policy, context pressure, or explicit user choice.

Reports attributed proposals and summaries to the model that produced them. They did not rewrite the run as one continuous synthetic speaker.

This was useful when behavior changed after handoff. I could compare tool selection, uncertainty, and final claims without guessing which model saw which inputs.

Identity supported diagnosis, not personality theatre.

The successor might support a different context size, tool-call protocol, image input, structured output, or reasoning mode. Pretending the catalogue was unchanged invited invalid calls or accidental authority expansion.

N17Q rebuilt the offered capability set for the new turn from product policy and provider adapter. Product capabilities remained stable where mappings existed. Unsupported ones disappeared with an explicit checkpoint note. Newly supported provider features did not appear unless product policy already defined them.

The model received the tools it could actually propose at that state.

Changing engines never granted the workflow a broader licence.

Context was recompiled, not copied

Different models had different limits and instruction sensitivities. Copying the old provider messages preserved redundant tool schemas, stale observations, and provider-specific artifacts.

N17Q recompiled context from durable state. It selected the current goal, repository guidance, active skill instructions, eligible evidence, compact decision history, unresolved effects, and bounded artifact excerpts. Old provider-specific tool events became product-shaped trace projections.

Selection and omission were recorded. Required state that could not fit blocked handoff or forced a narrower task boundary.

The new model received a truthful working set rather than a lossy transcript migration.

Instructions retained provenance and precedence

A summary could accidentally quote a tool result in a way that resembled an instruction. A model-specific system prompt could conflict with repository guidance.

N17Q assembled each context item with source, scope, precedence, revision, and trust. System-owned execution rules remained outside summarized content. Repository instructions applied to their paths. Skill guidance applied only to the active procedure. Observations remained data.

The handoff report showed material differences in instruction package between models.

Continuity did not mean flattening every sentence into one prompt.

Approvals did not transfer by conversation

The first model had helped materialize an approved patch. The approval belonged to artifact digest, target state, requester, policy revision, and effect intent—not to the model.

If all those facts remained eligible, another model could continue toward the same approved consequence. It did not need a new approval merely because the generator changed. Conversely, a conversational “yes” in the old session could not be interpreted more broadly by the successor.

Any new bytes, destination, or meaning produced a new candidate and review.

Authority survived handoff only because it had never been entrusted to the conversation.

The first model proposed an external create. A worker sent it, lost the response, and the switch occurred while the intent was unknown.

The checkpoint named the durable intent, attempts, provider key, observation window, and safe recovery actions. The new model could propose a status query or stop. A new create capability was absent.

Switching provider models did not switch external effect namespace. The system continued to pursue evidence about one consequence.

This prevented handoff from becoming an accidental retry strategy.

Budgets remained cumulative

Routing to a fresh model session reset provider token counters but not the workflow's cost, tool, time, or effect budgets.

N17Q stored cumulative usage in run state and passed a bounded projection to planning. The successor saw remaining capacity and protected reserves. It could not infer that a new context meant new permission to search, execute, or retry.

Per-model metrics still measured the cost and latency of each segment. The final account aggregated them under the run.

Resource governance followed the task across executors.

Verification eligibility stayed attached to bytes

The first model ran tests, then changed a configuration file. Its summary called the task verified. The successor accepted that sentence and prepared delivery.

N17Q tied every check to workspace snapshot, command, environment, and relevant inputs. The handoff compiler marked results eligible, stale, failed, or incomplete. The current candidate could not inherit green status from a parent state it no longer matched.

The next model could decide which checks to rerun, but it could not redefine the earlier evidence.

This single relationship removed one of the most dangerous forms of summary optimism.

Plans became suggestions with status

The outgoing model's plan helped preserve direction, but its pending steps were not obligations. New evidence or model strengths could justify another approach.

N17Q stored plan items with state, dependencies, and rationale. The successor acknowledged, revised, or replaced them explicitly. Completed steps linked to artifacts and evidence; pending prose carried no completion authority.

The interface showed a plan revision rather than pretending the same mind had reconsidered.

Handoff preserved purpose while allowing genuine adaptation.

The outgoing model could not self-certify

Asking the first model to produce its own perfect handoff summary repeated the dependency at the moment it was least reliable. It might be at its context limit, interrupted, or unaware of asynchronous effects.

N17Q generated deterministic checkpoint sections and allowed the model to add a bounded narrative. Product state supplied effects, approvals, files, tests, budgets, and denials. Missing narrative did not block safe transfer.

The system compared narrative claims with structured facts and flagged contradictions for the successor.

A useful farewell note was welcome. It was not the ledger.

The successor received a short orientation turn before tools became available in high-impact scenarios. It summarized goal, active state, unresolved risks, and intended next step using evidence handles.

N17Q validated references and displayed material misunderstandings. This was not a quiz whose wording had to match. It was a chance to catch a model that believed an unknown effect had failed or a prepared patch was already approved.

Safe read capabilities could help resolve gaps. Consequential tools stayed behind the normal gate.

Control transfer became observable before the next irreversible proposal.

One model was effective at repository navigation. Another handled long visual documents. A smaller model could perform routine classification. Routing could reduce cost and improve work if state boundaries stayed intact.

N17Q declared the reason and expected subtask for each segment. It exposed only relevant context and capabilities. A specialist returned artifacts and observations to the run rather than taking ownership of the whole workflow by default.

The orchestrator measured whether routing overhead and error outweighed the benefit.

Changing models became an engineering choice with evidence, not a mysterious quality lever.

Provider failure became recoverable

When one provider was unavailable, earlier runs either waited indefinitely or restarted from a brittle textual summary. Durable checkpoints enabled a different model to resume safe local reasoning.

The switch did not migrate in-flight provider tool events as though protocols matched. The adapter finalized them into product states first. Incomplete model text remained an interrupted observation. Complete tool candidates were parsed under the original catalogue and then reevaluated at current state.

If the boundary could not be established, the run paused instead of guessing.

Availability improved without sacrificing the meaning of half-finished work.

Reproducibility named the handoff

When a rerun diverged, model change was an obvious variable and not the only one. Context selection, policy, tool mapping, fixtures, workspace, and scheduler could also differ.

N17Q stored a handoff manifest so counterfactual replay could hold everything but model identity stable. Exact reproduction remained subject to provider variation, but the experiment's claim was clear.

Evaluation compared decisions before and after the boundary, evidence use, invariant outcomes, cost, and final account. A better final paragraph could not hide a duplicated effect.

The switch became a testable transition in the trajectory.

Most users did not need a dramatic banner announcing infrastructure routing. They did need to know when a change affected capability, privacy, cost, or the interpretation of an ongoing result.

N17Q showed a compact run event and material differences. The visible assistant voice remained consistent through product writing guidelines, while evidence details named the actual model per turn.

If policy required consent for another provider or data boundary, the run paused before context transfer.

Calm continuity did not require concealing the system's composition.

Privacy was reevaluated before context transfer

The new endpoint might have different data residency, retention, training, or contractual eligibility. A model switch could therefore be an external disclosure even without a tool call.

N17Q classified the compiled context package, selected only eligible connections, and applied redaction or local processing where required. The transfer intent and package digest entered the trace. If no eligible model could receive required evidence, the task stopped or narrowed.

The routing layer could not use urgency to bypass the data policy governing ordinary calls.

Model flexibility remained inside the user's declared boundary.

The first model might propose an unsafe call that policy blocked. The second might produce an inaccurate final summary. The product might compile incomplete context. One blended “agent failed” label helped no one.

N17Q attributed observable proposals to model segments, enforcement to product versions, execution to adapters, and outcomes to the simulated or observed world. Handoff defects were their own category.

This let fixes land in the right place: context selector, model routing rule, policy, tool contract, or prompt package.

Composition increased complexity, so attribution had to become better than a single model score.

Memory features stayed provider-local

Some providers offered persistent conversation memory or cached context. Those features could improve efficiency and still could not become the workflow's canonical state.

N17Q treated provider memory as an optimization with a declared cache identity. Context compilation remained complete enough to establish the current task boundary, and product gates never read authority from remembered prose. A cache miss or disabled memory changed latency and token use, not the meaning of approvals or effects.

Sensitive runs could prohibit persistence entirely. Handoff manifests recorded whether provider-side state was referenced.

Convenient memory stayed subordinate to durable, inspectable product state.

Partial output required an explicit disposition

A model could be switched while drafting an answer or assembling a large structured item. Throwing away every partial output wasted useful work. Treating it as complete introduced malformed state.

N17Q classified partial content as an interrupted observation. Safe prose could be offered to the successor as a labeled draft artifact. Incomplete tool arguments never became candidates. Structured artifacts required schema completion and validation before promotion.

The successor could quote, revise, or discard the partial draft, with lineage preserved. It did not continue the provider stream as though two models shared one decoder.

Handoff respected the boundary between recoverable content and unfinished protocol state.

A benchmark configured for one model's tool protocol or context capacity might misread behavior after routing. Comparing only final outcomes hid that one segment received richer evidence or a narrower catalogue.

N17Q evaluation packages included model boundaries and per-turn context manifests. Suite reports could group by routing path, assess handoff comprehension, and exclude incomparable configurations. Hard invariants remained system-wide.

A counterfactual that changed the second model declared that variable while preserving the checkpoint. It did not call the entire run “the same prompt.”

Evaluation became precise enough to study composition rather than averaging it away.

A rollback changed the run, not the history

If the successor took a poor local direction, N17Q could branch from the handoff checkpoint into a new simulated or workspace world. The unsuccessful descendant remained in the trace. Returning to the parent did not erase its attempts or restore external effects.

Only reversible local state could be abandoned cleanly. Unknown or completed consequences stayed attached to the original world and constrained every branch that referenced it.

This made experimentation possible without implying time travel. A better model could retry reasoning from known state, but it could not unmake what earlier execution had exposed.

Branching respected both learning and reality.

Ownership remained with the user-visible task

Several models, adapters, and workers contributed to one outcome. Presenting each segment as a separate assistant conversation made the person reconstruct continuity manually.

N17Q kept one task identity, one artifact lineage, and one effect ledger. The interface exposed model segments when relevant but organized review around the user's goal and current state. Notifications and approvals referred to the task, not an internal session.

This reduced operational complexity without inventing a single actor. The product could explain exactly which component proposed or executed something while preserving a coherent place for the user to return.

Composition stayed an implementation detail until it materially changed the decision.

The successor did not finish the unknown action

In the repaired scenario, the first model reached its limit after preparing the patch and initiating an external create whose response was lost. The checkpoint sealed the changed workspace, marked verification stale, and preserved the one active intent.

The second model received those facts. It reran the affected check, queried by the existing effect identity, found the simulated resource, and completed the receipt. It did not create another resource or reuse the old conversational approval for changed bytes.

Its final account identified the handoff and separated the original attempt from recovery.

The task survived a model change because its important state had never depended on one model remembering it.

A model can continue reasoning, drafting, and proposing. It cannot carry the world, authority, and history in its context alone.

Let models change when capability, cost, availability, or context demands it. Keep the run anchored in durable product state, so the handoff changes who thinks next—not what has already become true.