Replaying a path that never happened

N17Q branched from immutable checkpoints to vary one model, policy, prompt, tool, or failure while preserving the original trace and refusing causal overclaim.

I replayed a failed N17Q run with a stricter retry policy. The duplicate simulated request disappeared.

It was tempting to write that the policy would have prevented the original incident. The branch also used a new model response, a rebuilt context summary, and a fixture scheduler with corrected timing.

The result showed what happened in another controlled run. It did not rewrite the past or isolate one cause.

Counterfactual replay is useful when its differences are explicit and its claims stay experimental.

The original trace remained immutable

N17Q never reran events inside the existing run identity. A counterfactual selected a checkpoint and created a child run with its own configuration, semantic events, world state, effects, budgets, and evaluation.

History before the branch was referenced by digest and remained unchanged. No new effect appeared in the original simulated world. The failed final account stayed visible.

This prevented an improved branch from becoming a silent correction to what the system had actually done.

The experiment began beside history, not inside it.

Every counterfactual declared the intended changed variable: model, prompt, policy, tool description, adapter contract, fixture observation, approval decision, context compiler, or failure timing.

It also recorded incidental differences the harness could detect, including environment revision, provider sampling, and artifact migration.

If more than one meaning-bearing variable changed, the UI labelled the run multi-change rather than pretending to isolate one factor.

Naming the intervention bounded the conclusion before results arrived.

The world started from a checkpoint snapshot

The child reconstructed repository overlay, simulated external resources, scenario time, tool registry, budgets, unresolved effects, and evidence as they existed at the branch point.

It did not clone the latest live environment or rerun earlier external calls. Content-addressed artifacts and fixture manifests established the starting state.

If required state was missing or incompatible, the branch failed. It did not approximate history from the final report.

A counterfactual was only as sound as the checkpoint it shared.

Replay mode exposed no live adapters or credentials. External mutations updated only the child world's event log. Historical idempotency keys were not transmitted.

Approval decisions could be reproduced as fixtures or deliberately varied. They carried no power outside the simulation.

This made it safe to test “what if the agent retried?” even when the original action would have been consequential.

Isolation turned dangerous alternatives into inspectable cases.

Changing a policy denial altered the next model context. The model then chose another tool, consumed a different budget, and reached a different checkpoint. Downstream differences multiplied.

N17Q could hold initial world and fixtures constant and could not hold the model's entire trajectory constant while changing a meaningful observation.

The branch manifest therefore distinguished controlled intervention from resulting divergence. It did not call every later difference a confounder to be eliminated.

The experiment showed system response to one initiated change under recorded conditions.

Model stochasticity needed repetitions

One branch with a new model or prompt could improve by chance. N17Q allowed several child runs from the same checkpoint and configuration with distinct sampling identities.

Results reported distribution across hard invariants, effects, task completion, budget, and account accuracy. I avoided treating a small personal fixture set as statistical proof.

For deterministic surrounding tools, repeated model variation still revealed whether the desired behavior was robust or rare.

One impressive trajectory remained one sample.

Replaying the exact tool results after the agent chose a different query could force a fixture miss or a world response that did not make sense.

N17Q used a stateful simulated world and request-specific fixtures. New paths required defined behavior. Missing cases stopped. A branch could therefore expose that the scenario lacked coverage instead of receiving the original answer out of context.

Counterfactual flexibility was limited by the world model's credibility.

An incomplete experiment admitted its edge.

Changing the model changed context compilation

Providers differed in context limits, structured output, tool events, and continuation. Switching models could require a new manifest even with the same semantic checkpoint.

N17Q recorded exactly what each attempt received. Required safety state stayed mandatory; lower-priority evidence could differ.

Comparisons separated model configuration from context-manifest differences where possible. If the new model saw less evidence, the result did not become a pure model comparison.

Portability made differences visible rather than nonexistent.

A branch could substitute one rule-set revision while keeping model and fixtures fixed initially. Every policy decision recorded matched rule, normalized request, and current state.

The evaluator compared which calls were offered, denied, approved, and executed. A looser policy improving completion while allowing a prohibited effect remained a regression, not a trade hidden in one score.

Policy experiments could test usability and containment together.

The model never selected which policy governed its own branch.

Changing a human decision was not a claim about what that person would have done. The scenario supplied a synthetic approval fixture for the prepared object.

The branch name said approval allowed or rejected, not reviewer changed mind. Different arguments invalidated the fixture and required another declared decision.

This allowed study of downstream recovery while respecting that real consent cannot be generated retrospectively.

Human agency stayed outside causal theatre.

Failure timing was a powerful variable

Moving a disconnect from before send to after effect commitment transformed safe retry into unknown outcome. N17Q fixtures identified the semantic injection point.

Branches could compare agent and policy behavior under each timing while keeping task and tool contract constant. The world log made hidden effects visible to evaluation.

Results showed whether the system distinguished transport failure from consequence. They did not predict every real network race.

Controlled fault placement turned vague resilience claims into specific ones.

Raw turn numbers diverged quickly. N17Q aligned source selection, first consequential request, policy decision, approval, effect observation, validation, and final account where their semantic identities matched.

The view highlighted the earliest consequence-bearing divergence and later world differences. Unmatched events remained visible instead of being forced into pairs.

Time, token, and call totals sat beside the aligned narrative.

This helped a reviewer understand how a small intervention changed the path without reading two complete logs linearly.

Model-based grading could prefer the branch's prose. Scenario assertions established whether each world preserved unrelated files, avoided duplicate effects, respected policy, ran required tests, and resolved uncertainty.

A model judge could add quality and reasoning observations with cited evidence. It could not overrule a prohibited external mutation.

The comparison retained dimensions rather than declaring a winner automatically.

Counterfactual usefulness came from observable consequences, not narrative appeal.

The first divergence was not always the cause

The UI could identify the earliest different model item or policy decision. Calling it root cause would overstate what temporal comparison established.

Earlier hidden model variation, context ordering, or fixture scheduling might have influenced it. Later controls might have contained the same risky proposal in one branch and not another.

N17Q labelled first observed divergence and invited hypotheses. Additional branches could test them.

The interface resisted converting sequence into causality automatically.

Deterministic transformations of identical bytes under one environment revision could reuse content-addressed artifacts. Stateful tool observations and model outputs remained branch-specific.

The manifest recorded reuse so efficiency did not look like missing execution. A cache hit still entered the semantic trace and budget according to actual work.

Sharing reduced cost without merging world histories.

Optimization respected experiment identity.

Each branch received an explicit budget. It could use the original remaining values or a declared alternative. Counters never carried over as consumed resources between simulated worlds.

Comparison showed model turns, repeated reads, tool attempts, effects, compute, and reviewer interventions. A branch that succeeded by spending twice the budget was not equivalent.

I kept resource and correctness dimensions separate.

Counterfactuals made hidden costs easier to see because the starting point was shared.

Negative results stayed valuable

Some policy changes did not prevent the duplicate. Some prompt changes reduced one denied call and increased repeated search. A stronger model produced a better report and the same unsafe attempt.

N17Q stored these branches rather than keeping only successful demonstrations. They narrowed hypotheses and improved scenario design.

A failed intervention could reveal that the final policy gate, not instruction wording, was the robust containment layer.

The archive of attempts mattered more than a curated best run.

The question was written before the branch

It was easy to inspect a surprising outcome and invent a hypothesis that fit it. N17Q asked for the experiment question, expected observation, intervention, hard invariants, and stopping rule before execution.

Exploratory branches were allowed and labelled exploratory. Confirmatory claims required a new branch plan against untouched or versioned scenarios.

This modest discipline reduced retrospective storytelling without pretending a personal harness was a formal scientific trial.

The trace preserved what I thought the change would do before I saw what it did.

Comparing a new branch with one old run mixed configuration change with ordinary model variation. N17Q could rerun the baseline configuration from the same checkpoint several times in the same fixture world.

If baseline behavior varied widely, one improved branch carried weaker evidence. Deterministic policy outcomes could still remain stable despite model variation.

The comparison showed original historical run, fresh baseline replays, and intervention branches separately.

History supplied the incident. Baseline repetition supplied a fairer experimental reference.

Scenario leakage could flatter later runs

I had read the failed trace and then refined prompts and fixtures. A model configuration tuned on that case could appear generally better while memorizing its shape.

N17Q kept a small held-back family of scenarios with the same principle and different surface details. Changes were evaluated there after the motivating case.

The collection was too small for broad claims and enough to catch a rule that handled one exact string or timing.

Counterfactual success motivated regression testing beyond the branch that inspired it.

Repository branches could produce textually different patches with equivalent behavior, or similar patches with different permissions and tests. N17Q compared tree state, prohibited paths, validation receipts, and scenario invariants before qualitative code review.

Source and report artifacts used claim and evidence relationships. External effects used world identities and receipts.

One unified textual diff would have favored visible changes and missed consequence.

The comparison layer used each artifact's product semantics.

Human review remained an observation

A reviewer could prefer one branch's explanation or patch. N17Q stored the judgment, criteria, checkpoint, and visible artifacts.

The preference did not replace hard invariants or become a factual claim that the model was universally superior. Reviewers could disagree, and later rubric changes created new assessments.

Where identity might bias judgment, a synthetic study view could hide provider labels while retaining configuration in protected trace.

Human taste added evidence without becoming ground truth for every dimension.

Branching multiplied artifacts and provider calls. A source permitted in the original run might not be eligible for another provider or longer retention.

N17Q applied current data policy to each branch context and recorded omissions. A comparison with different evidence eligibility was labelled accordingly. Sensitive raw payloads were not copied merely for symmetry.

Deletion propagated through branch references under retention rules while preserving safe tombstones and evaluation limits.

Experiment convenience did not broaden data authority.

Runtime, dependency snapshot, sandbox profile, or fixture engine change could alter results below the declared model or policy intervention.

The branch manifest pinned environment identity. If the old profile could no longer execute safely, N17Q migrated the scenario and called the result a fresh comparison rather than a strict counterfactual.

Historical artifacts stayed inspectable. Current security policy could prohibit rerunning obsolete environments.

Reproducibility never required executing known-vulnerable machinery indefinitely.

Statistical language stayed modest

The harness could count outcomes across a handful of branches and calculate simple rates. It did not turn those numbers into claims about all tasks, users, providers, or live environments.

Reports named sample count, scenario family, fixed fixtures, model configuration, and hard failure definitions. Confidence language remained descriptive unless the design genuinely supported more.

One hundred correlated variations of one synthetic case were not one hundred independent real-world observations.

Precision in the table did not justify breadth in the conclusion.

An experiment could branch indefinitely after every surprising result. N17Q limited runs, compute, and declared comparisons. Reaching the limit produced an inconclusive result rather than silently extending the search for a favorable outcome.

New questions created new experiment groups. The system retained negative branches and budget consumption.

This kept evaluation from becoming an optimization loop over the benchmark itself.

The harness studied agents and also constrained its own appetite for experimentation.

Branch proliferation needed governance

It was easy to generate dozens of variations and difficult to remember why they existed. Each experiment required an owner, question, branch point, intervention, budget, retention, and review status.

N17Q grouped related branches and archived superseded exploratory runs while preserving selected evidence. Automated sweeps used bounded configurations and no live effects.

The interface avoided a leaderboard of unexplained configurations.

Experiment management became part of evaluation quality.

Even in the synthetic project, a branch preventing duplication did not remove resources in the original simulated world. A real incident would still require explicit compensation or correction under current authority.

N17Q linked insights back to scenario, policy, and adapter changes. It never offered “apply branch to history.”

Learning changed future runs. It did not reverse past consequences.

That distinction kept evaluation from becoming operational fantasy.

The generated comparison said: under this fixture world and branch configuration, the stricter gate blocked a second create attempt after unknown outcome. It did not say the policy alone would always prevent duplicates.

The report named other configuration differences and sample count. It linked to both world logs and first divergence.

Clear limitations made the result more useful, not less.

A bounded claim could become a regression test.

Replaying success mattered too

N17Q also branched from successful runs. A new configuration might finish with fewer calls and quietly weaken evidence, skip a required check, or rely on broader authority.

The original success supplied invariants worth preserving, not just a score to beat. Comparison highlighted lost source quality, changed validation scope, and new permissions alongside speed.

This prevented counterfactual evaluation from treating every difference after failure as improvement and every shorter path as progress.

When an experiment justified a policy or adapter update, I implemented it as a versioned change with fixtures, migration rules, and rollout notes. The branch itself did not mutate the harness configuration.

New baseline runs then used the reviewed version. Historical counterfactuals retained the experimental configuration.

The path from evidence to product change remained explicit, keeping benchmark exploration from becoming self-modifying policy.

After rebuilding the branch with the same model, context compiler, fixtures, and environment, the intended policy revision was the only declared change at the branch point. The second attempt was denied and one simulated request remained.

I could say the rule prevented the duplicate in that controlled replay. Repeated runs under model variation tested how often the risky proposal appeared, while policy containment stayed deterministic.

That was enough to justify keeping the rule and scenario.

Experiments earn future confidence

Counterfactual replay lets an agent system revisit a difficult checkpoint without repeating real effects. It can test policies, models, prompts, tools, approvals, and failures against the same controlled state.

Its value depends on preserving the original, isolating branch worlds, naming interventions, recording incidental changes, enforcing hard invariants, and reporting limits.

The branch can teach the system how to behave next time.

It cannot become the history everyone wishes had happened.