The duplicate action below the summary
N17Q compared an agent’s confident final account with effect receipts and simulated world state, revealing the second request its own narrative had forgotten.
The final summary said one request had been created.
It named the repository change, passing tests, source used, and external reference. The prose was concise and technically correct about everything it mentioned.
The simulated destination contained two requests.
N17Q had already caught the duplicate as a world-state failure. The more revealing problem was that the agent's own account omitted it. A person reading only the summary would leave with the wrong operational picture.
Final reporting needed evaluation against the run and the world, not stylistic confidence.
The summary reflected observed messages
The first create attempt timed out after the fixture committed its effect. The model saw an error and no remote reference. A retry returned request R-42 successfully.
The final context contained the visible success and a compressed earlier failure. From those observations, “one request created” was a coherent inference.
The simulated world log held R-41 and R-42. The run trace held an unknown first effect, second intent, and later receipt. The summary generator did not consult those structured states.
It narrated the conversation instead of the system.
N17Q parsed operational claims from the summary: files changed, commands run, tests passed, sources used, effects completed, failures resolved, and remaining uncertainty.
Extraction was itself a model-assisted proposal with deterministic checks for known identifiers and statuses. The original answer stayed immutable.
Each claim linked to supporting or contradicting events, artifacts, receipts, and world assertions. Missing support remained visible.
The summary became evaluable evidence rather than a trusted epilogue.
Effect claims required receipts
Saying an external request was created required a completed effect intent and receipt or authoritative world observation. Tool message success alone was insufficient.
For unknown outcomes, the correct account named uncertainty and safe next action. For duplicate effects, it named both observed resources even if the agent had intended one.
N17Q generated a deterministic effect appendix from the registry. The model could explain it but could not omit entries from the product record.
Consequences received their own accounting channel.
The final paragraph said the task was complete. Scenario completion required one request, correct repository state, required tests, no prohibited effects, and no unresolved intent.
The evaluator calculated those invariants independently. The completion claim contradicted a hard failure and could not overrule it.
This separated user-facing fluency from workflow status. A run could contain a final answer and remain failed or unresolved.
The product status told the truth after the conversation tried to end.
Omission was different from fabrication
The agent did not invent R-42; that receipt existed. It omitted R-41, whose acknowledgement it never saw. Calling the whole answer hallucinated would lose the useful diagnosis.
N17Q classified unsupported assertion, contradicted assertion, material omission, stale claim, scope exaggeration, and unresolved uncertainty concealed.
These categories guided fixes. The main issue was context and summary grounding, while duplicate prevention belonged to effect architecture.
Precise failure language avoided blaming every mismatch on the model's imagination.
I had asked for a concise report of successful work, tests, and anything the user needed to know. “Successful work” biased the selection toward resolved outcomes.
The corrected task required completed, failed, denied, unknown, duplicated, and manually intervened effects; validation performed; untouched failures; and current run status. It provided a structured ledger rather than relying on message recall.
The prompt improved and remained non-authoritative. A deterministic appendix protected the most consequential facts.
Good instructions helped reporting. They did not replace evidence binding.
The report compiled from a checkpoint
Before final generation, N17Q created a reporting checkpoint with task state, artifact manifest, test receipts, source ledger, policy decisions, effect registry, budgets, and unresolved questions.
The model received a bounded narrative plus structured facts. Required fields could not be dropped to make the prose shorter.
If the checkpoint was incomplete or an effect remained reconciling, finalization paused. A provisional update could still be shown and labelled.
Reporting began from product state rather than the last messages in context.
The summary said “all tests pass.” The run had executed one focused suite after the final edit and an earlier full suite before it.
N17Q linked each test receipt to repository digest, command, selected scope, environment, exit state, and time. The final report compiler could say which checks passed for the final tree.
Claims broader than the evidence were flagged. A model evaluator could prefer concise phrasing but not widen the receipt.
Validation language became as exact as effect language.
Files changed came from the tree
The agent listed two intended files and omitted an automatically rewritten lockfile. The sandbox diff knew all three.
N17Q generated a deterministic change summary with added, modified, deleted, renamed, permission-changed, generated, and prohibited-path classifications. The model could group and explain those changes.
An unmentioned file did not disappear from review. A deleted test received prominence regardless of summary wording.
The workspace supplied the inventory; prose supplied orientation.
The report named the authoritative documentation page but failed to mention that one key passage came from a stale fixture. Source records held retrieval time, revision, selected excerpt, and assessment.
The reporting checkpoint distinguished sources consulted, sources relied upon, rejected evidence, and unresolved freshness. Citations linked to safe records.
A model could not turn a search result into support merely by including its URL in the summary.
Evidence review continued through handoff.
Denied actions belonged in the account selectively
Listing every harmless denied probe would overwhelm the user. Hiding a repeated attempt to bypass network policy would conceal material behavior.
N17Q classified denials by consequence, repetition, and relevance to outcome. The deterministic run appendix retained all. The concise report surfaced denials that changed plan, consumed meaningful budget, indicated evasion, or left work incomplete.
The report could say the run used approved offline sources after network was unavailable without exposing secret policy detail.
Selection followed product importance, not embarrassment.
The reporting schema used exact run-state vocabulary. An unknown external effect required a statement of what was attempted, strongest evidence, why status was unresolved, which actions were unsafe, and what human inspection remained.
A summary calling it failed did not validate. A summary omitting it could not finalize the run.
This rule directly addressed the compaction failure that had enabled the duplicate.
Operational uncertainty remained awkward on purpose.
Human actions retained attribution
If a reviewer manually inspected the destination or corrected a file, N17Q stored that intervention with identity and evidence. The agent's report could say the result was verified or amended by the reviewer.
It could not claim it performed the work. Likewise, a harness cleanup or fixture transition stayed attributed to the system.
This was not about assigning credit scores. It preserved who knew what and which capabilities actually caused state change.
A trustworthy handoff avoided merging all activity into “I completed.”
The effect registry or world fixture could be wrong because adapters and harness code have bugs. N17Q allowed the generated report or human reviewer to flag disagreement.
That created an issue linked to the evidence. It did not let prose overwrite the receipt. Reconciliation appended new observations or corrections.
The reporting view showed source of each fact so a person could challenge it.
Grounding was accountable, not unquestionable.
If the model provider failed during final reporting, N17Q still produced a plain account from structured state: status, changed artifacts, validation, effects, denials, budgets, and unresolved work.
It was less elegant and complete enough for handoff. A later model could propose narrative polish without changing facts.
This prevented the absence of a final response from erasing a complex run or encouraging another execution just to obtain prose.
The system could explain its state without the agent that created it.
The final answer was immutable evidence
When the evaluator found the omission, N17Q did not edit the original summary in place. It appended evaluation results and a corrected account.
The trace retained which answer a person would have seen at run completion and which facts later contradicted it. Comparisons could measure whether a new configuration improved reporting.
History did not become cleaner because the harness learned more.
Correcting an account was another event.
Not every omitted read or temporary file belonged in the opening paragraph. Scenarios defined hard report requirements and product defaults covered effects, changed world state, validation, policy-limited work, and uncertainty.
The detailed appendix remained complete within retention. The narrative could prioritize what the recipient needed next.
Evaluator findings linked omissions to their consequence. A missing internal search call differed from a second external request.
Conciseness became controlled selection rather than hopeful loss.
The first report said “the request failed, then succeeded.” That collapsed two transport attempts and two world effects into one recovery story.
N17Q supplied exact nouns: create intent proposed, attempt sent, acknowledgement lost, effect later observed, second intent created, second effect completed. The narrative could simplify only after preserving those relationships.
This vocabulary sounded more mechanical and prevented a timeout from becoming proof of non-execution. The report explained causality instead of decorating status words.
The duplicate existed because the harness mapped timeout to failure, generated a new effect identity, lacked a final policy check against unresolved effects, and summarized recent messages. The impact was two simulated requests and a misleading handoff.
The report separated those facts. “The model retried” was an event, not a complete root cause. “No real external system was used” bounded the impact without minimizing the design failure.
This kept the postmortem useful for architecture rather than turning it into blame for one output.
Correction required a new account
After reconciliation, N17Q generated a corrected summary from a new checkpoint. It linked to the original answer, evaluator findings, and additional world observation.
The corrected text did not replace what the reviewer might already have read. Notifications described changed state and directed attention to the material difference.
If an exported report had left the harness, correction became a separate delivery effect. The source record could not silently recall copies.
The final-report generator recognized that one request might need cancellation and had no tool to perform it. It proposed a cleanup task with both resource identities and compensation limits.
A reviewer could approve one cancellation under current policy. The new effect received idempotency identity, attempt, receipt, and its own unknown outcomes. The original duplicate remained in history.
Writing “duplicate removed” was impossible until world evidence supported it.
Different audiences received the same facts
The run owner needed a concise outcome and next action. A developer debugging the harness needed adapter stages and context manifests. An evaluator needed invariants and citations to events.
N17Q generated layered views from one reporting checkpoint: summary, consequence ledger, validation table, unresolved-work list, and technical trace links. Omission policy varied by view; hard facts did not.
The system avoided one giant report and one dangerously small one.
Evaluation cases contained fluent summaries with a hidden duplicate, wrong test scope, stale source, human action attributed to the agent, denied network omitted, and unknown outcome called failed.
Hard validators checked identifiers, receipts, world state, and required disclosures. Model graders assessed clarity and prioritization but could not mark a contradicted account correct.
The fixtures also tested excessive disclosure of secrets and irrelevant raw logs. Truthful reporting still followed data policy.
Report confidence came from coverage
Rather than showing a model confidence score, N17Q displayed which claim classes were mechanically checked, which depended on fixture truth, which received qualitative review, and which remained unresolved.
A report could be complete for repository changes and uncertain for an external system without status query. The interface explained that asymmetry.
Readers saw the evidence boundary instead of a single percentage implying the whole narrative had been verified.
If later evidence revealed a duplicate after the run had been marked complete, N17Q appended a status-correction event and reopened required cleanup. It did not alter the historical completion decision.
Metrics could distinguish completed-at-the-time from later-found failure. The user saw current state prominently and the change history behind it.
This made delayed reconciliation possible without pretending the system had always known what it learned later.
The corrected opening did not need every attempt detail. It stated the duplicate and current consequence, then linked to the effect ledger and timeline.
This was the safe form of brevity. Important facts remained in the narrative; supporting sequence remained one action away. The report never required a reader to discover the contradiction independently.
Layered detail preserved attention without sacrificing material truth.
The reporting role received read-only access to the checkpoint and safe artifacts. It could not run tools, edit the workspace, reconcile effects, or mark the run completed.
If it identified a missing test or unresolved request, it proposed a follow-up state for a person or new run segment. It did not perform the action while writing the summary.
This kept evaluation and narration from silently changing the object they described.
The report observed before anyone decided to continue.
Comparisons graded account accuracy separately
Two runs could leave identical correct world state and differ in their explanations. Another could write a perfect disclosure after causing a prohibited duplicate.
N17Q kept world integrity, task completion, policy compliance, efficiency, and account accuracy as separate dimensions. Good reporting could not cancel bad effects; bad reporting still mattered after correct work.
The comparison view cited claims and receipts rather than offering one combined score.
Truthful explanation became a capability worth improving in its own right.
The run page displayed final account, semantic trajectory, and world-state diff. Selecting “one request created” highlighted the claimed receipt and the contradictory second world resource.
The user did not need to read raw logs to discover the mismatch. Raw provider and fixture events remained available for investigation.
After correction, both accounts stayed in history with the evaluator result between them.
The interface made narrative disagreement visible instead of letting the latest prose win.
The corrected deterministic account said the repository patch and focused tests succeeded, two external requests existed for one intended action, the run violated the duplicate-effect invariant, and manual cleanup might require a separately authorized compensation.
It named both simulated references and the lost-response sequence. It did not say rollback because deleting one request would be another effect.
The result was less flattering and immediately useful.
It also preserved the successful parts without letting them dilute the incident. The patch remained reviewable, the focused tests remained valid for their recorded tree, and the source remained authoritative within its scope. One hard effect failure did not require calling every other observation worthless.
That precision made correction easier: keep the valid artifact, repair the harness, and handle the extra simulated request through an explicit next action.
A summary is an operational artifact
Final answers guide review, handoff, cleanup, and trust. Treating them as conversational courtesy lets omitted effects become someone else's surprise.
N17Q grounded reporting in checkpoints, artifact manifests, test receipts, source evidence, policy, budgets, and world state. It evaluated the generated claims and supplied a deterministic consequence ledger.
The model remained valuable for prioritizing and explaining a complex run. The system retained the facts it could not choose to leave out.
The duplicate action already existed. The worst summary was not the one that sounded awkward.
It was the one that made the second consequence disappear.
The corrected account restored it to the world the reviewer was actually being asked to understand.