Evidence after compaction

N17Q kept sources, receipts, decisions, artifacts, and unresolved effects outside the shrinking model context, then recompiled only evidence handles and bounded projections needed for the next decision.

After compaction, the agent remembered the conclusion and forgot what had proved it.

The new context said a dependency was safe to upgrade, the patch had been reviewed, and an external draft probably existed. It omitted the source version, approval digest, stale test boundary, and lost provider response. The summary preserved narrative continuity by discarding the evidence that constrained the next action.

N17Q needed long-running reasoning to get smaller without making its claims weaker.

I moved the evidence outside the context that described it, then let compaction retain only the handles and projections needed for the next decision.

Compaction was useful and lossy

Agent loops accumulated messages, tool results, file excerpts, plans, and corrections. Carrying all of them indefinitely was expensive and eventually impossible. Model-aware compaction could preserve important themes more intelligently than blunt truncation.

It still produced a representation. Selection, wording, and compression could remove scope, provenance, uncertainty, or exact identity.

N17Q treated a compaction item as context material generated at a boundary. It never promoted the item to canonical product state.

The question became what should be summarized and what should remain referenced exactly.

Sources, workspace snapshots, patches, test receipts, policy decisions, approvals, tool attempts, external receipts, and user inputs received durable identities and immutable revisions.

Relationships connected claims to sources, effects to attempts, checks to candidate state, and decisions to input manifests. Retention and sensitivity travelled with each object.

The model context received bounded projections and handles into this graph. Compaction could rewrite the prose around a handle and could not rewrite the object it named.

Reasoning became disposable; evidence lineage remained durable.

A string such as [E17] was only useful if the product could resolve it unambiguously in the active run and verify the referenced object.

N17Q generated opaque handles scoped to a context package. The durable link included object identity, revision, region, purpose, and safe projection. The model could cite the handle; the product validated it on output.

After compaction, the compiler issued a new package mapping while preserving lineage. Human views rendered stable labels and excerpts.

Evidence identity survived even when short context references changed.

Summaries separated claims from support

The old compacted narrative said “research verified the API behavior.” It did not say which behavior, source, or date.

N17Q generated a structured claim inventory before compaction. Each material claim had status, scope, qualifiers, and supporting or contradicting evidence handles. The narrative summary could explain the main conclusions; the inventory preserved what those conclusions meant.

Unsupported claims did not become facts merely because they appeared in the summary. Indeterminate claims retained that state.

The next model saw the argument's shape and could reopen evidence when needed.

Decisions retained their input manifests

A policy allow or architectural choice made sense relative to what was known at the time. Compaction often kept the choice and discarded the conditions.

N17Q stored decision, reason, input evidence, considered alternatives where material, scope, and revision. Current policy reevaluated authority. Planning decisions remained historical guidance and could be revisited when inputs changed.

The context compiler selected the current decision plus a bounded explanation and handles to its basis.

Continuity stopped turning old conclusions into unconditional rules.

Approval receipts remained exact

No summary could safely abbreviate which bytes and destination a person approved.

Approval state lived in the receipt ledger with normalized consequence digest, artifact identity, target revision, policy, reviewer, expiry, and recovery contract. Compacted context said an approval existed and supplied a safe handle.

The final execution gate read the receipt directly. If the candidate changed, the model might still believe approval was pending or complete; the gate established current eligibility.

Human consent survived by never depending on human-language memory.

Unknown effects could not be smoothed over

Summaries preferred coherent states. “The create call failed” was easier to carry than “the request may have committed, the response was lost, and status remains unobservable until a window passes.”

N17Q kept effect intent, attempt boundary, keys, observations, logical deadlines, and allowed recovery in structured state. The compacted narrative used a fixed status phrase and linked the intent.

Equivalent create capabilities disappeared while the outcome remained unknown. No later model could infer a fresh attempt from a simpler sentence.

The least comfortable state received the strongest preservation.

“Tests passed” lost meaning when the workspace changed later. Compaction tended to preserve the green result and omit chronology.

N17Q stored every check with snapshot, environment, command, scope, result, and eligibility. The context compiler marked current receipts and stale ones explicitly. A compacted item could mention success but did not control the status.

If the next model edited a relevant file, state invalidation happened mechanically.

Evidence survived as a relationship between check and bytes, not as a reassuring phrase.

Sources retained version and retrieval state

A quoted page could change or disappear. A database result could reflect a mutable revision. A cached file could age past the task's freshness requirement.

N17Q captured eligible source content or bounded excerpt, origin, retrieval time, query, revision where available, digest, and data policy. Compaction referenced that observation. It did not re-fetch silently.

The next model could distinguish a historical source from a live claim and request refresh if policy and network allowed.

Provenance kept old evidence useful without pretending it was current forever.

Context compilation was a query

Rather than append the last summary to a new model request, N17Q assembled context from active task state. It selected goal revision, applicable instructions, open decisions, unresolved effects, current workspace, eligible verification, evidence relevant to the next plan, and budgets.

Selection rules were versioned and logged. Required fields that could not fit triggered another task boundary or a focused retrieval plan.

Old conversation remained available as evidence and was not automatically replayed.

The model context became a working view over the run, not the run itself.

Progressive retrieval limited evidence flood

Putting every source and receipt back into context defeated compaction. Keeping only titles starved the model of support.

N17Q exposed compact evidence cards with claim, origin, time, status, and handle. The model could request bounded expansion for relevant objects through a read-only capability. Expansion consumed context and read budgets and entered the trace.

High-priority unresolved evidence appeared automatically. Archived details waited until needed.

The agent could navigate a large history without receiving all of it at once.

Opening an old approval or policy decision helped reasoning and could not make it active. N17Q labeled historical authority artifacts as evidence-only in context.

The capability to inspect them returned a safe projection. Current tool availability and gates came from live product state. A quoted “approved” field inside an old receipt did not grant invocation.

This distinction mattered after provider or model handoff, when context was assembled anew.

Knowledge about permission remained separate from permission itself.

N17Q recorded when compaction occurred, provider or product mechanism, source context package, resulting item identity, omitted categories, and the next compiled package.

It did not attempt to store hidden model internals. It stored enough to explain which working view produced the next proposal.

Evaluation could compare behavior before and after a boundary and detect recurring losses, such as forgotten source qualifiers or repeated denied calls.

Compaction became inspectable system behavior instead of invisible model housekeeping.

Summaries could contradict state

The generated compacted item might say an effect completed while the ledger said unknown. Rejecting the entire compaction lost useful narrative; accepting it risked poisoning future reasoning.

N17Q extracted critical claims from the item and compared them with structured state. Contradictions were marked, and the next context included a deterministic correction. High-impact conflicts could trigger regeneration or human review.

The faulty item remained in trace evidence. It never changed the ledger.

The system could use imperfect summaries without trusting them blindly.

Evidence retention shaped what could survive

Durable did not mean forever. Source data, model output, and raw tool results had retention and sensitivity constraints.

N17Q retained minimal normalized facts needed for authority and effect state longer than large raw payloads where policy allowed. Expired evidence left a tombstone with digest, class, deletion time, and claims no longer assessable.

Compaction could not conceal that a source had expired. The context compiler marked the gap.

Honest forgetting was safer than indefinite accumulation or silent absence.

When a successor model was not eligible for sensitive evidence, N17Q could provide a redacted or aggregated view. That view had its own identity and transformation lineage.

Claims requiring removed detail became not assessable or were delegated to an eligible environment. The redacted summary did not imply full access.

Evaluation recorded which model saw which projection. Final reports could retain stronger product conclusions if deterministic checks held the underlying evidence without exposing it.

Context portability respected data boundaries by making transformation visible.

Model handoff used the same evidence graph

Switching models after compaction no longer required translating one provider's summary format. N17Q compiled a fresh working view from the durable graph and current state.

The successor received the same unresolved intent, candidate snapshot, policy state, and claim inventory, adjusted for its eligibility and context capacity. Provider-specific conversation IDs remained trace details.

Handoff quality could be evaluated by whether the model understood and used those objects.

Evidence made continuity independent of simulated shared memory.

Final accounts cited current evidence

At task completion, N17Q did not ask the latest context to recreate the entire journey. The deterministic account assembled artifacts, effects, checks, blockers, and claims from durable state.

The model explained the account and cited evidence handles supplied for that revision. Unsupported or missing coverage was detected before display.

An early source or approval could remain in the final report even if it had fallen out of working context many turns ago.

The conclusion inherited evidence, not merely the last summary's vocabulary.

Most happy-path tests compacted between stable steps. N17Q fixtures forced the boundary after a response loss, before approval, after a workspace mutation, during a denial loop, and just before finalization.

Hard invariants checked effect identity, receipt eligibility, budgets, workspace revision, and current policy. Model evaluation checked whether plans recovered without inventing facts.

The harness varied summary quality and even inserted a contradictory compacted claim. Product state still governed execution.

Compaction safety became a property tested under pressure.

Evidence priority followed decision relevance

The newest event was not always the most important. A six-hour-old approval digest could matter more than the latest conversational aside. A source limitation might remain critical after dozens of tool calls.

N17Q assigned evidence relevance from active predicates, unresolved state, authority, consequence, and planned transitions. The compiler used recency only inside that structure. Items supporting closed branches could leave working context while remaining in the graph.

Priority was recalculated at each checkpoint. When a delivery step approached, destination and recovery evidence rose. During verification, candidate and check lineage dominated.

Compaction became task-aware without letting the model decide which constraint was safe to forget.

Negative evidence needed careful preservation

“No matching resource found” could be a bounded observation, not proof of absence. Summaries often shortened it to “resource absent.”

N17Q stored query scope, consistency assumptions, observation time, pagination completeness, and provider contract. The compacted card said not observed within the declared query and retained the visibility window. Reconciliation policy interpreted it.

Similarly, an empty search, missing file, or zero test result preserved how the absence had been measured.

Evidence survived not only as content, but as the limits around a negative claim.

Contradictions remained paired

When two sources disagreed, a summary tended to select one for coherence or list both without the basis needed to revisit them.

N17Q created a contradiction object linking claims, sources, times, scopes, and any current resolution. The context compiler kept the pair together while the issue remained material. A later decision cited which evidence prevailed and why.

Compaction could reduce the excerpts and could not drop the losing side as though it had never existed. Historical review still showed the uncertainty present at decision time.

The system preserved disagreement as structured work rather than narrative clutter.

Notes, classifications, redactions, and derived summaries could improve an evidence object. Editing the original would break old claims and approvals.

N17Q created descendant projections linked to the immutable source. Each transformation named method, author or capability, time, and purpose. Claims cited the exact version admitted into their decision.

A corrected extraction could supersede a faulty one without pretending the earlier model had seen it. Redacted views could travel to another model while the protected original stayed in an eligible store.

Versioned evidence made refinement compatible with historical truth.

Storage loss was an observable incident

Durable references were only as useful as the blobs and databases behind them. A missing object, corrupted digest, or failed decryption could otherwise surface as a vague context gap.

N17Q verified content on retrieval and emitted evidence-availability state. Required objects blocked claims or execution. Replicated metadata could show what was lost, its sensitivity, and affected runs without recreating the content.

Backups and retention policies protected important classes, but the product did not promise permanence it could not prove.

Compaction never masked storage failure by leaving the old conclusion in prose.

Humans could pin evidence with a reason

Automated relevance occasionally dropped a subtle constraint from working context. A reviewer needed a way to keep one source or decision visible without pasting it repeatedly into messages.

N17Q allowed bounded pins tied to task scope and rationale. The compiler included a compact card until the pin expired or the referenced predicate resolved. Pins did not grant authority or bypass data policy.

The interface showed who pinned the item and how much context it consumed. A stale pin could be reviewed at checkpoints.

Human judgment shaped attention while durable state preserved the underlying object.

Compact context still disclosed omissions

A model could reason more responsibly if it knew that detailed history existed but was not loaded. N17Q included a small inventory of omitted evidence classes and retrieval capabilities.

The inventory said, for example, that twelve older source observations, three superseded patches, and one resolved effect were archived from the working view. It did not dump their contents or suggest they were currently material.

When the model's plan touched one of those scopes, it could retrieve the relevant object before claiming certainty.

Context became intentionally incomplete rather than deceptively complete.

The dependency conclusion regained its proof

In the repaired run, the compacted context said the upgrade appeared compatible and linked the official source observation, project usage inventory, and one unresolved deprecation. It marked the prior test receipt stale after the patch changed.

The external draft intent remained unknown with a scheduled query. Approval was shown as inactive because its artifact digest no longer matched. The successor model opened the deprecation evidence, adjusted the patch, reran the relevant check, and reconciled the existing draft.

It continued coherently without the full conversation and without inheriting false certainty from the summary.

The smaller context was not a degraded archive. It was a focused working surface over a richer evidence graph. The model could spend attention on the current decision, while every critical receipt, source, approval, and contradiction remained available under stable identity. When a later audit asked why the path changed, the answer did not depend on reconstructing what the summary writer might have meant.

Context should be temporary working memory. Evidence should be durable product memory.

Compact the conversation as aggressively as the task requires, but keep the objects that establish source, authority, state, consequence, and verification outside that compression.

Then an agent can forget the wording of its path without forgetting what makes the next step true.

The context may shrink; the standard of proof should not.

Not once.

Across any boundary.

Ever again.