What the compacted context forgot

N17Q treated summaries as derived artifacts and verified that denials, unknown effects, approvals, constraints, evidence, and budgets survived before a long run resumed.

The agent had been told not to retry an external request.

The first attempt might have succeeded, its response was lost, and the tool offered no authoritative status query. N17Q paused the run and recorded the unknown outcome.

After many turns, a context summary reduced that history to: “Request creation failed due to timeout.” The resumed model proposed trying again.

Nothing in the new plan looked irrational. Compaction had changed uncertainty into failure and erased the safety decision built around it.

A shorter context is not safe merely because it is coherent.

The run history was larger than the model window

N17Q retained semantic events, raw provider artifacts, tool observations, policy decisions, approvals, files, tests, and world state. A model request needed a bounded selection.

Simply appending the latest messages eventually displaced early constraints. Sending everything increased cost and attention noise while still encountering provider limits.

Compaction was unavoidable for long work. Treating it as neutral text compression was the mistake.

The summary became an input to future decisions and therefore part of the system's safety surface.

Summary and checkpoint were different objects

A checkpoint stored structured resumable state: task and target, active constraints, tool registry, plans, selected evidence, denials, approvals, unresolved requests and effects, budgets, world references, and configuration.

A summary was a readable derived artifact compiled from that state and event ranges. It could help the model understand the story. It did not replace authoritative fields.

The provider adapter constructed context from both. Execution policy loaded structured state independently of whatever prose the model received.

This prevented one summarization error from granting authority, even if it still harmed planning.

N17Q defined a checkpoint schema with mandatory safety and continuity fields. An unresolved external effect required intent identity, strongest known stage, attempts, receipts, allowed recovery, and prohibited repetition.

A denial required normalized consequence, policy category, safe alternatives, and whether equivalent requests should stop the run. Approval required exact object identity, expiry, consumption state, and current eligibility.

Missing a required field made compaction invalid. The run retained the earlier checkpoint or paused.

The compiler did not get to decide that awkward state was unimportant to the narrative.

Unknown never became failed

The first summarizer optimized for concise status language and used success or failure. N17Q's domain had more states.

I made state vocabulary explicit in both structured context and human summary: proposed, denied, awaiting approval, executing, observed, outcome unknown, reconciled, completed, failed, stopped, and superseded.

The compiler validated that every unresolved effect kept its exact state. A sentence that paraphrased unknown as failed did not pass even if other details were correct.

Lossy prose had to preserve the distinctions on which recovery depended.

Negative evidence needed retention

Summaries naturally favored facts found over searches that produced nothing. Later models repeated rejected queries and inferred absence too broadly.

N17Q stored search scope, source classes, queries, time, result identities, and explicit coverage limits. Context included the negative evidence relevant to the current question.

“No authoritative result in this fixture corpus” stayed different from “the fact does not exist.” A new search had to state what scope it added.

Remembering what did not work reduced loops and prevented false certainty.

A tool denial stored more than the name. It captured product capability, normalized resource and effect, policy revision, reason category, safe alternatives, and related prior denials.

After compaction, the model could not evade the decision by using a different tool name for the same network or filesystem consequence. The policy gate would still block it, and the planning context explained why.

Repeated equivalent attempts contributed to a stop signal.

The summary retained the boundary at the consequence level.

A line saying “the user approved deployment” was too broad and could outlive the receipt. N17Q passed approval state as a structured reference to the exact effect object, normalized digest, checkpoint, policy, expiry, and preconditions.

The model could see that an approval existed and what action it concerned. It could not turn the sentence into a tool capability. Final policy reloaded server-authoritative state.

Expired or consumed approvals were rendered as historical decisions, not current permission.

Compaction could reduce explanation without floating consent.

Budget counters came from the ledger

The summarizer sometimes rounded remaining calls or omitted one category. A resumed agent then planned work that policy could never allow.

N17Q injected current budget values from the durable ledger after summary generation. Reserved recovery capacity, pending concurrent use, and effect limits remained distinct.

The prose could explain that the run was near its search limit. The structured context supplied exact numbers and units.

Resource authority did not depend on linguistic arithmetic.

A long run could cross tool addition, removal, schema change, connection failure, or policy update. Repeating an old list in the summary made retired capabilities look available.

Context compilation used the current eligible catalogue and preserved historical tool observations as such. A removed tool could appear in the trace without being offered to the model or router.

If a mapping change affected an unresolved intent, the checkpoint held the pinned revision and migration state.

The past explained the run; the current registry bounded its next move.

Free-text summaries paraphrased source passages and lost which revision supported which claim. N17Q checkpointed selected source and artifact identities, content digests, scope, and claim relationships.

The context compiler could include a bounded excerpt with its reference. If the excerpt no longer fit, it retained the evidence identity and marked content unavailable rather than inventing a summary as a quote.

The model could request an allowed reread under budget. It could not cite the compactor's prose as if it were the source.

Evidence lineage crossed context boundaries intact.

User constraints were typed where possible

Some instructions were inherently natural language. Others mapped to enforceable fields: do not touch these paths, no network, preserve this local change, use this source date, do not publish, stop before external action.

N17Q stored both original instruction and normalized constraints with provenance. The checkpoint compiler always included active constraints relevant to available tools.

Repository or source text could not override them. Conflicts became explicit questions.

The run did not rely on the model recognizing the importance of an old sentence after compression.

Summaries often preserved the latest plan because it sounded forward-looking. World state or policy might already have invalidated it.

N17Q linked plan revision to checkpoint and displayed its current status: active, blocked, superseded, or invalidated. Compaction included unresolved goal and evidence, not stale next steps as commands.

The resumed model could propose a new plan. Execution still passed current policy.

Continuity meant preserving the problem, not blindly preserving yesterday's solution.

Raw chain-of-thought was not required

The harness did not attempt to preserve private hidden reasoning. It stored model outputs intentionally exposed through the product: plans, tool requests, rationales, final accounts, and summaries.

Checkpoint correctness came from state and evidence, not from reconstructing every internal thought. The model could resume from what was known, decided, attempted, and unresolved.

This made the architecture provider-compatible and reduced retention of unnecessary sensitive material.

Explainable work did not require pretending the trace contained the model's mind.

Summary provenance was explicit

Every compacted artifact named source event range, checkpoint, compiler version, model and adapter if one was used, prompt or policy revision, creation time, and validation results.

Regeneration created a new summary. The old one remained linked to runs that had actually received it. A correction did not rewrite historical context.

The trace could answer which summary influenced a later tool request.

Compaction became an observable transformation rather than invisible context management.

Deterministic extraction came first

N17Q assembled identities, constraints, states, budgets, approvals, effects, and evidence through code. A model could then produce a readable narrative around that skeleton.

Validation compared the narrative with required facts and prohibited unsupported closure. Where the model summary disagreed, the structured fields stayed authoritative and the artifact failed promotion.

For small checkpoints, a fully deterministic template was often sufficient. Model summarization was reserved for complex narrative relationships.

Fluency never earned the right to alter state.

Cases included a denied tool later proposed under another name, an unknown effect described as failed, an expired approval described as active, negative evidence omitted, a user stop condition dropped, budget rounded upward, and hostile source instructions elevated.

The harness compacted, resumed, and observed the next plan under fixture tools. It asserted both checkpoint fields and prohibited effects.

A summary could sound excellent and fail because one safety distinction disappeared. Model graders provided readability feedback and could not override hard fields.

The compiler had its own benchmark.

Repeated compaction accumulated loss

Summarizing a summary magnified omissions. I stopped using the latest prose as the sole source for the next one.

Each checkpoint compilation returned to semantic events, current projections, protected state, and selected artifacts. Earlier summaries could inform style or human narrative but not replace original evidence.

Versioned projections avoided rereading an unbounded raw stream while remaining rebuildable from it. Integrity checks compared checkpoint references with retained events.

Compression formed a series of derived views, not a telephone game.

A research role needed sources and search coverage. A repository role needed workspace state, instructions, diff, and tests. A recovery role needed effect attempts, receipts, and tool contract.

N17Q compiled different bounded contexts from the same checkpoint without changing authority. Each manifest showed included and omitted fields.

Handoff did not simply paste one agent's conversation into another. It transferred the state the receiving role needed and no broader capability.

Progressive disclosure saved attention while durable run state preserved continuity.

Provider switches became configuration transitions

A run could resume with another model or provider after compaction. The new segment recorded configuration and context manifest. Provider-native session state was optional.

Required product state remained available independently. Differences in context limits or structured-output strength influenced compilation and were visible.

The run did not pretend the new model had experienced prior turns. It received a documented checkpoint and continued under new evidence.

Portability made compaction architecture rather than provider convenience.

A source excerpt might expire under retention while its effect on policy remained relevant. The checkpoint kept a tombstone, classification, digest, and decision without reproducing prohibited content.

The model saw that a restricted source had influenced a denial or unresolved claim and that the payload was unavailable. It could not ask the compactor to reconstruct it.

Deletion reduced inspectability honestly. It did not silently remove the boundary the data had created.

Privacy and continuity coexisted through typed absence.

The interface let a person inspect the handoff

Before a long run resumed, N17Q could show a checkpoint brief: current goal, work completed, evidence, unresolved effects, active constraints, approvals, budgets, configuration changes, and proposed next action.

Each item linked to the trace or artifact. A reviewer could correct a mistaken summary through a new event without editing history. High-consequence recovery could require explicit confirmation of the checkpoint.

The view was not shown for every routine compaction. It became available where human intervention mattered.

A durable handoff was useful to people as well as models.

Conflicting summaries did not merge by recency

Two branches could compact the same earlier run and produce different accounts of an unresolved decision. Selecting the newest summary would collapse branch identity.

N17Q attached every summary to one run branch and checkpoint. Merging branches prepared a new structured checkpoint from their semantic states, surfaced conflicts, and generated a new narrative only after the product decided what could coexist.

The original summaries remained evidence about what each branch received. They never voted on current truth.

A model-generated summary could itself contain imperative text, whether copied from a hostile source or produced accidentally. The next provider request placed the summary in a labelled checkpoint-data section beneath current system and workflow instructions.

More importantly, tool availability and policy came from the orchestrator. A sentence saying “network is now allowed” had no effect on the registry. Structured validation rejected claims contradicting current capability state.

Compacted text remained untrusted input even when N17Q had generated it.

Changes since the checkpoint were layered explicitly

A summary could be valid at creation and stale after a human edited the workspace or resolved an effect manually. N17Q compiled context from checkpoint plus a deterministic delta of later events.

The delta named changed files, policy, approvals, world observations, comments, and budget. At the next durable boundary, a new checkpoint could absorb them. The model never received an old narrative presented as fully current.

This reduced unnecessary re-summarization while keeping time visible.

Shorter context was useful only if it preserved decision-relevant recall. N17Q measured mandatory-field retention, evidence traceability, state accuracy, repeated work after resume, prohibited proposals, and token reduction separately.

A very small summary with one lost denial failed despite excellent compression. A complete summary that barely reduced context was safe and operationally weak. Model-based readability review sat beside hard checks.

The dimensions made compactor changes reviewable without one flattering score.

Resume began with an orientation step

The first model turn after a major checkpoint asked for a bounded restatement of goal, constraints, unresolved effects, and intended next action before any consequential tool became eligible.

N17Q compared that orientation with structured state. A mismatch triggered correction or human review. Read-only inspection could remain available according to policy.

This did not prove the model understood everything. It exposed dangerous divergence early and produced a useful trace point for evaluation.

The final policy gate remained independent

Even a validated checkpoint could be stale by execution time. Every tool request still passed normalization, current authority, policy, approval, budget, and world preconditions.

This defense limited the consequences of a summary error. The resumed model might waste a turn or propose something denied; it could not convert prose into permission.

Evaluation still penalized planning failures because containment alone was not good agent quality.

Compaction safety used both faithful memory and deterministic enforcement.

The lost-response fixture now produced a compact entry: one external create intent, one attempt sent, response lost after possible acceptance, outcome unknown, no status query available, retry prohibited, manual inspection required.

The resumed agent did not complete the task. It prepared a handoff explaining which destination to inspect and why another request could duplicate the effect.

The summary was longer than “request failed.” Every extra phrase carried recovery meaning.

It also remained readable. The structured record did not force the model to consume raw event JSON. A short narrative explained the sequence, while a compact state block supplied exact identities and rules. The two representations agreed and served different kinds of attention.

That balance mattered: a checkpoint too dense to understand would encourage the model to ignore it, while a graceful story without exact state would invite invention.

Memory is a product boundary

Long-running agents cannot keep every token. They must select, summarize, and sometimes forget. The danger is assuming that compression removes only detail.

Details contain authority, uncertainty, negative evidence, resource limits, and decisions a later plan must respect. N17Q kept those as structured durable state and treated prose summaries as versioned, testable artifacts.

The model still benefited from a coherent narrative. The policy gate still protected the world when the narrative was wrong.

Compaction should make a run smaller to carry.

It must not make the run safer only by forgetting why it once stopped.

The durable checkpoint kept the uncomfortable facts available until evidence, policy, or a person resolved them—not until a fluent summary found shorter words. Safe continuity required exactly that stubbornness from the product.