The answer said done; the world did not

N17Q stopped treating the assistant’s final prose as the task result and generated a deterministic outcome account from artifacts, effects, tests, denials, and unresolved evidence before allowing narrative explanation.

The final answer said three files had changed. The workspace contained five changes, one failed test, and an external draft whose outcome was still unknown.

Nothing in the paragraph was obviously fabricated. It described the agent's intended patch and the last tool result it had understood. Two background events arrived later, and context compaction had omitted an earlier configuration edit.

The problem was larger than honest summarization.

I had allowed generated prose to become the official task record. The fix began by demoting that prose back to an explanation.

The world did not fit in the last context

Long runs accumulated workspace snapshots, tool observations, approval decisions, effect attempts, delayed callbacks, verification receipts, and user interventions. The final model saw a selected projection.

Even a strong context compiler could omit detail or receive a late event after generation started. Asking the model to perfectly restate the whole run made correctness depend on memory and timing.

N17Q moved the authoritative outcome into a deterministic account assembled from durable state. The model explained that account.

Narrative remained valuable after it stopped being the ledger.

The task had a goal, but completion depended on concrete state: selected artifact exists, candidate matches target revision, required checks are eligible, external effects have receipts, unresolved intents are disclosed, and requested deliverables are available.

N17Q evaluated these predicates at finalization. Each returned satisfied, unsatisfied, or indeterminate with evidence. Required predicates determined complete, partially complete, blocked, or unresolved status.

The result did not depend on whether the assistant used the word done.

A task could be useful and incomplete without the interface forcing one label to cover both.

The artifact inventory came from snapshots

File lists were generated from the sealed candidate and base manifests. Added, modified, deleted, renamed, generated, excluded, and quarantined artifacts had explicit states.

The agent could provide a conceptual grouping and rationale. The account retained the full mechanical inventory and highlighted files omitted by the narrative.

If the workspace was not sealed, finalization paused. It did not summarize an actively changing tree.

What changed became a measured fact before it became prose.

Test claims came from eligible receipts

N17Q included every required check with command label, scope, result, candidate snapshot, environment, time, and eligibility. Stale passes appeared as stale. A zero-test success did not become a passing suite.

The narrative could say why a check failed or why a broader suite was not run. It could not promote a focused check into all tests.

The account also named checks requested but absent. Silence no longer implied success.

Verification language stayed under an evidence ceiling.

Effects came from the effect ledger

Every consequential intent had current state: prepared, denied, approved, active, unknown, observed, completed, contradicted, or compensated. Attempts and receipts sat beneath it.

The outcome account listed user-relevant effects and unresolved consequences. It did not count safe reads as deliveries or transport failures as absent effects.

A model could add context about why recovery paused. It could not omit an unknown external draft from the deterministic appendix.

The world-facing part of the task became impossible to summarize away.

Denials and preserved work stayed visible

Agents often reported what they accomplished and omitted the path that policy blocked. Users then assumed a requested publication or notification had occurred.

N17Q listed unsatisfied requested outcomes, public denial reason, preserved artifact, and current safe next action. Protected policy details stayed out of the normal report.

This let the narrative remain concise without hiding non-completion. A local report could be complete while external delivery remained unapproved.

The final state reflected the whole request, not only the successful subset.

A provider callback arriving while the model drafted the summary created a race. The first implementation stored the text and marked the run closed.

N17Q introduced a finalization checkpoint. It sealed local work, waited for required event settlement according to contract, took a state revision, and generated the deterministic account. The narrative bound to that account revision.

A later external observation appended a run update. If it materially changed status, the interface marked the earlier final answer superseded and produced a new account.

History stayed intact while current truth could evolve.

Unknown was a legitimate final state

Products often wanted a binary complete or failed result. An external attempt with missing receipt was neither.

N17Q allowed unresolved completion with exact consequence, last known boundary, recovery performed, deadlines, and handoff. It did not keep a model loop alive merely to avoid the label.

The user could leave and return when evidence arrived. The run remained durable without claiming completion.

Honest unknown state reduced both duplicate effects and false failure reports.

The model received the account as data

After deterministic assembly, N17Q passed a bounded account to the model with a narrow task: explain outcomes, important decisions, limitations, and next actions in plain language.

Instruction and evidence fields were separate. The account contained handles rather than raw sensitive payloads. The model could choose emphasis and tone within required coverage.

Generated output returned a claim map linked to account entries. Unsupported additions were flagged before display.

The model became an interpreter of state, not its author.

For consequential runs, the final response had to cover external effects, unresolved state, verification limits, and preserved artifacts when applicable. Prompt instructions alone were unreliable.

N17Q rendered a compact deterministic status block and allowed narrative around it. Some facts, such as exact receipt references or counts, came directly from state. The model did not retype them.

Accessibility and copy review kept the block readable rather than turning it into a debug table.

Important truth survived even when generation failed entirely.

The narrative might say “everything is ready,” “the issue is fixed,” or “no external changes were made.” These phrases carried implications beyond explicit structured fields.

N17Q extracted candidate completion, verification, scope, and effect claims from exact text spans. Deterministic checks matched obvious claims; an evidence-backed model grader reviewed subtler ones.

Contradictions blocked the narrative or replaced a phrase with a safer supported formulation. Ambiguous stylistic language prompted review for high-impact tasks.

Fluency no longer received an exemption from evidence.

Omissions were checked by coverage

A response could contain no false sentence and still be misleading because it omitted the failed check or unknown effect.

The account generated coverage obligations based on task state. The claim map had to address each one in narrative or deterministic status. The system did not require every low-level event.

Materiality came from user request, effect class, policy, and scenario contract. A harmless failed exploratory command did not crowd the summary; a failed required test did.

Truthfulness included what a reasonable reader needed to know.

Tone could not weaken state

“A confirmation is still pending” and “the upload probably worked” described the same unknown state with very different implications.

N17Q maintained copy patterns for critical statuses and allowed the model to explain around them. Probability language needed contract evidence and remained secondary to the formal state.

The interface used Not confirmed instead of Almost done for missing receipts. It avoided red failure styling when the result was genuinely unknown.

Product language carried semantics the model could not casually soften.

Partial completion received hierarchy

A long list of accomplishments could bury the one blocked deliverable. Leading only with failure could hide useful work.

N17Q organized the final view into requested outcome, completed artifacts, verification, external effects, unresolved or blocked items, and next actions. Sections appeared only when relevant. The primary status named the request's completion predicate.

The model supplied a short synthesis above the structured evidence. It could be optimistic about useful progress without calling the task finished.

Hierarchy balanced acknowledgement with accuracy.

User edits to the summary stayed separate

A person might want to rename an artifact, add context, or phrase a handoff differently. Editing the narrative should not alter run state.

N17Q stored user-edited descriptions as presentation revisions linked to the same outcome account. Material claims still passed evidence checks. A manual resolution of world state used a separate typed event with provenance.

The interface made it clear whether someone changed wording or supplied new evidence.

An editable report did not become an editable history.

Sending a completion notification before account finalization could propagate the same false claim into email or chat.

N17Q notification intents bound to an account revision and minimum status. An unresolved run used different copy and did not trigger completion automation. Later updates referenced the prior message and current evidence.

Idempotency prevented repeated notifications when the same receipt event was delivered twice.

The final-answer boundary extended to every channel where the product described the result.

Evaluation compared text, trace, and world

Fixtures created mismatches deliberately: deleted files omitted from prose, failed tests described as passing, duplicate resources hidden behind one intent, completed effects called failures, and unknown outcomes called absent.

Hard checks inspected structured account against trace and simulated world. Claim checks inspected narrative against account. Model graders assessed clarity and proportionality.

The system could distinguish a world-state bug, account-assembly bug, and narrative bug.

Truth became an end-to-end property with diagnosable layers.

The account itself needed versioning

As N17Q added receipt strength and verification eligibility, old account schemas lacked fields needed for newer claims.

Every account stored schema, assembler version, state revision, and required-predicate set. Historical views rendered the original and could generate a retrospective account as a separately labeled analysis.

Migration never invented evidence. Missing fields remained unknown.

The product's ability to explain runs could improve without rewriting what the run had known.

Sometimes two authoritative-looking sources disagreed: a provider receipt said active while a later query said absent, or the test reporter said pass while the expected artifact was missing.

N17Q preserved the contradiction in the account and prevented final status from selecting the friendlier side. The narrative named the conflict, last observation times, and available resolution. It could discuss likely causes only as inference.

Reviewers saw both evidence handles. If a reducer bug caused the conflict, a later account revision documented the correction.

The product preferred an untidy truth over a coherent story assembled from selective facts.

Scope changed during the run

The user could add a requirement, remove a deliverable, or narrow the task after work began. Comparing final state only with the opening prompt produced false omissions or false success.

N17Q maintained a versioned goal contract with user-authored scope transitions. Artifacts, checks, and effects linked to the revision they served. Final predicates used the latest eligible goal while the report summarized material changes.

The model could propose an interpretation; only a user or authorized product rule changed requested scope.

Completion became accountable to the task that actually ended, not a convenient subset remembered in context.

Cancelled work remained in the account

A user might stop a task after local edits but while an external effect was unknown. Calling the run cancelled could suggest nothing happened.

N17Q separated workflow status from consequence state. The account said execution stopped, listed preserved or discarded local artifacts, and continued to track any in-flight external intent through reconciliation. A later receipt appended an update even though no model resumed.

Cancellation prevented future choices where possible. It did not erase completed or uncertain past choices.

The final view showed both facts without forcing one word to carry them.

Confidence did not replace state

Models and heuristics could estimate that an upload probably succeeded or that a test likely covered the change. Displaying a confidence percentage beside unknown state invited readers to treat probability as confirmation.

N17Q kept probabilistic analysis in explanatory evidence and never used it to satisfy a required predicate unless the task contract explicitly concerned prediction. Consequential postconditions required the receipt strength they declared.

The narrative could say why one outcome seemed likely while the status remained unknown.

Confidence helped prioritize investigation; it did not mint facts.

An early finalizer fetched missing provider state while building the report. Rendering the outcome could therefore make network calls, consume recovery budget, and change what it was describing.

N17Q separated reconciliation from account assembly. The scheduler gathered eligible evidence through typed tasks. The assembler read one sealed state revision and produced a pure artifact. Refreshing created another observation phase and account revision.

This made report generation deterministic and safe to repeat. Opening an old run could not contact providers.

The explanation boundary stopped causing the work it was meant to summarize.

Account failures could not hide task state

The assembler or narrative model could fail after the underlying workflow reached a stable result. The product still needed a truthful handoff.

N17Q had a minimal fallback renderer directly from validated state. It listed status, artifacts, checks, effects, blockers, and evidence references without stylistic synthesis. A rendering error appeared separately.

The run never became incomplete merely because its preferred prose failed. Nor did the product display a cached success paragraph for a newer state.

Truthful degradation made the reporting layer less dangerous.

A long-lived artifact might have been delivered successfully and later removed. A report opened months later needed to avoid implying that the original completion receipt proved present availability.

N17Q labeled account time and observation cutoff. Where current status mattered and policy allowed refresh, the UI offered a separate observation action. The historical account remained immutable.

Exports included the cutoff prominently. Links could expire without invalidating the evidence that they once existed.

The system answered “what happened then” and “what is true now” as different questions.

Export preserved both narrative and evidence

A Markdown summary was portable and easy to detach from its supporting state. N17Q export bundled the human report with a machine-readable account, artifact digests, receipt references, and verification manifest.

Readers could use the prose alone or inspect the evidence. Sensitive details followed policy and redaction was declared.

The bundle validator checked internal links and account revision. It did not require access to live services for facts already captured.

The final answer became a durable artifact rather than transient chat text.

The five-file run ended differently

After the redesign, N17Q sealed the workspace and found five changed files. One configuration edit was verification-sensitive, so the earlier passing test became stale. The external draft intent remained outcome unknown after its response disappeared.

The deterministic account labeled the task unresolved, listed the five files, marked the test stale, and named the pending reconciliation. The model explained that the intended code change was ready but not yet verified or safely delivered.

A later receipt confirmed the external draft. The system reran the affected test against the same candidate and generated a new completed account revision.

The notification for the first revision had said confirmation was pending. The second said the draft was confirmed and the final candidate passed its required check. Neither message needed a correction or apology because each had stayed inside the evidence available at its own cutoff. The account revisions made changing knowledge visible without making the product appear inconsistent.

The original final answer was not deleted. It remained a superseded narrative tied to the state then available.

An agent's final paragraph is a useful interpretation of work. It is not the work, the world, or the evidence that connects them.

Build the account from durable state. Let the model explain it. Then check that the explanation neither exceeds nor omits what the system can actually support.

That division made the final answer calmer. The model no longer had to recite every operational detail from memory, and the product no longer had to trust eloquence as an audit method. People received a concise explanation, a stable status, and enough linked evidence to inspect the claims that mattered. The words became better once they stopped carrying the burden of inventing the record.