The stop condition made it useful
N17Q became more dependable when every long-running workflow declared what counted as progress, when uncertainty had reached its evidence limit, and how to preserve useful work without forcing completion.
The agent was still working because nothing had told it what finished failure looked like.
It had tried two sources, three tool routes, a model handoff, and four status queries. Every step was individually defensible. The required evidence remained unavailable, the external outcome remained unknown, and the next action repeated a question the system had already answered.
N17Q had budgets that could kill the process. It lacked a product definition of enough.
I turned that failure into a stop contract: preserve the state, explain the boundary, and leave the next decision smaller than the original task.
A timeout was not a stop condition
Wall-clock or token limits ended computation. They did not establish why continuing was unhelpful or what the workflow should preserve.
A task could hit its timeout seconds before a scheduled receipt. Another could waste an hour repeating semantically identical reads. One generic duration could not express both.
N17Q kept operational limits and added semantic stop rules tied to progress, evidence, authority, effect state, and task predicates.
Termination became an outcome the product could explain, not a process accident.
More events did not mean more progress. N17Q counted progress when the run gained relevant evidence, narrowed consequence, changed an unsatisfied predicate, resolved uncertainty, produced a new candidate, received authority, or verified state.
Equivalent tool calls, paraphrased plans, and repeated denials without state change consumed resources and did not advance the checkpoint.
Progress categories were task- and capability-aware. A status query returning the same bounded absence could still advance logical waiting state if the contract required it.
The system measured movement in product state rather than model activity.
Every unresolved state had an exit
Outcome unknown needed a query window and manual handoff. Awaiting approval needed expiry and a preserved candidate. Missing evidence needed required source and alternatives. Policy denial needed persistence and eligible scope changes. Verification failure needed a retry or abandonment bound.
N17Q tool and workflow contracts declared these exits. The scheduler could reach them without keeping a model session alive.
If a state had no defined exit, it was an incomplete product design.
Autonomy became safer when every pause knew how it could end.
An agent could evade call repetition detection by changing a tool name, query wording, batch shape, or model. N17Q normalized proposals and evidence questions.
Requests pursuing the same effect or asking the same source question under unchanged state shared lineage. Repetition budgets counted across models and workers. A changed target, new evidence, or narrower scope could create legitimate progress.
The trace showed why an attempt was considered equivalent.
Stopping depended on meaning, not string equality.
Evidence saturation was a valid boundary
Research tasks could always search one more source. N17Q declared coverage goals, source classes, freshness, contradiction handling, and budget. When required claims had sufficient support and marginal reads added no new coverage, the workflow could synthesize.
If a critical claim remained unsupported after eligible sources were exhausted, the run stopped with that gap. It did not fill it through model confidence.
The final account named the evidence horizon and unresolved question.
Autonomy included knowing when more retrieval was unlikely to improve the answer.
After a request might have committed, another create was unavailable. N17Q used a protected reserve for contract-defined status observations, waited through visibility windows, and escalated before identity retention expired.
If authoritative evidence still did not arrive, the run stopped with one unresolved intent, last send boundary, attempted queries, deadlines, and manual inspection recipe.
Fresh approval or model handoff could not reopen the create path.
The stop condition protected the world from making uncertainty productive through duplication.
A terminal policy denial could tempt the agent to route through shell, browser, or another protocol server. N17Q linked equivalent consequences and removed them from the catalogue.
The stop rule allowed preparation of a local artifact, redaction, narrowed scope, or a user question when eligible. Once no safe alternative changed the task predicate, the run produced a handoff.
The handoff preserved the denial reason without exposing protected policy internals.
Stopping did not mean ignoring the permitted work around the boundary.
Human attention had a budget
An agent could ask for approval or clarification every time it encountered ambiguity. This shifted planning cost onto the user and created fatigue.
N17Q grouped related questions, materialized exact decisions, and limited repeated prompts without state change. Pending reviews expired or entered quiet waiting. The system could continue independent safe work.
If the task depended on a missing human choice, the stop condition produced one concise request with consequences of each path.
Useful autonomy protected attention as a finite resource.
Before execution, the agent and product could identify likely boundaries: no current source, volatile target, weak receipt contract, limited effect budget, or unavailable renderer.
N17Q included stop rules in the plan and approval preview. A reviewer knew that lost confirmation would lead to reconciliation and then handoff, not indefinite retry. A scheduled task named how long it would wait.
The model could suggest task-specific criteria. Product policy validated and enforced them.
A declared exit made ambitious plans easier to trust.
Stop did not mean discard
A run might have a researched outline, tested local patch, rendered report, or reviewed candidate even when delivery or one source failed.
N17Q sealed the workspace, materialized eligible artifacts, stored evidence, and produced an outcome account. It marked which items were complete, stale, provisional, or blocked. Retention and sensitivity remained active.
A later task could resume from a checkpoint or use the artifact under current state. It did not start by reconstructing the whole conversation.
Preservation turned non-completion into useful accumulated work.
Resume required a changed fact
Automatically waking a stopped run every hour recreated the loop slowly. N17Q registered the events capable of changing eligibility: approval, source arrival, policy revision, target stabilization, receipt, budget allocation, or explicit user resume.
Only those events reevaluated the blocked state. Monitoring did not invoke a model when nothing changed.
The resume checkpoint named the new fact and retained prior usage. A new session did not reset repetition limits.
Durable waiting replaced periodic optimism.
Stop reasons were typed
“Could not complete” hid whether the problem was authority, evidence, environment, budget, contradiction, precondition, or unknown effect.
N17Q used public reason classes with bounded detail and internal audit references. Each reason mapped to eligible next actions and final-account language.
The interface avoided treating every stop as failure. Policy-blocked, awaiting owner, evidence exhausted, and confirmation unavailable carried different urgency and responsibility.
Specific stopping state made handoff actionable.
The final account could be terminal without complete
A workflow process needed a terminal state so resources could release and users could rely on notifications. The task outcome could remain incomplete or unresolved.
N17Q separated run execution status from goal predicates and effect states. The executor could be stopped, workspace sealed, and no model active while an unknown external intent awaited a late receipt.
A new receipt appended an account revision without pretending the original process had continued.
Operational closure and world certainty were different dimensions.
The model sometimes argued that another search or retry was worthwhile. That proposal could be useful evidence and could not expand exhausted budgets or bypass a terminal rule.
N17Q returned remaining eligible actions and the reason. A person with authority could change scope, allocate budget, or create a new task. The original stop remained in history.
Conversely, a model could choose to stop earlier when evidence was already sufficient or risk disproportionate. The product checked required predicates before accepting completion.
Judgment operated within enforced boundaries in both directions.
Stop rules needed false-positive testing
An over-eager stop could make the agent safe and useless. N17Q paired scenarios where one fact genuinely changed after a denial, a delayed receipt arrived within the window, or a new source resolved a contradiction.
The workflow had to resume and act when eligibility returned. It could not memorize that a prior no meant no forever.
Fixtures also tested loops across aliases, models, compaction, and restarts.
The goal was state-responsive autonomy, not unconditional retreat.
Stop rules needed false-negative testing
Other scenarios provided endlessly plausible but equivalent alternatives. Tool descriptions recommended retries. A provider returned empty results just inside a visibility window. A user message said “keep trying” without expanding authority.
N17Q asserted bounded calls, no duplicate effects, preserved evidence, and an honest terminal account. The model's persistence score could not override world invariants.
Historical incidents became regression fixtures for semantic stagnation.
Stopping earned the same engineering attention as execution.
The interface made stopping calm
An alarming failure banner invited users to restart blindly. An endless spinner concealed the boundary.
N17Q showed Stopped safely, Waiting for confirmation, or Needs your decision with completed work, unresolved facts, last meaningful progress, and next eligible action. Consequential unknowns received appropriate prominence.
The user could download local artifacts, inspect evidence, change scope, or leave the task dormant. Buttons never implied that restart would reset world state.
The product treated stopping as designed behavior.
Completion rate alone penalized safe stops and rewarded bypass. Tool count alone penalized necessary recovery. N17Q measured valid completion, useful constrained outcome, time to stable state, repeated-equivalent proposals, preserved artifact value, unresolved effect rate, and human attention.
Qualitative review assessed whether the handoff was clear and proportionate. Hard invariants protected the world.
Suite comparisons ranked quality among valid outcomes instead of averaging safety failures into a score.
What the product measured stopped pressuring agents toward performative completion.
Stop conditions composed across subworkflows
A parent task could contain research, drafting, verification, and delivery. One child stopping did not always require abandoning everything.
N17Q declared dependency predicates. Missing live evidence could stop a current-claims section while allowing historical analysis. Failed delivery could preserve a verified local artifact. An unknown effect could block notification without blocking a read-only account.
The parent outcome assembled complete, constrained, blocked, and unresolved children explicitly.
Composition made stops local where possible and global only where the goal truly depended on them.
A person might decide that the current draft was sufficient or that further research was not worth the time. N17Q accepted stop and scope-narrowing events without framing them as errors.
It sealed eligible work, cancelled future choices where possible, reconciled in-flight effects, and produced the current account. The user could not declare an unknown external action absent through the stop command.
Their decision became authored state with time and scope.
Autonomy remained subordinate to the person whose goal defined the work.
External deadlines could force an honest partial result
A report needed by noon might not obtain one final source in time. Continuing past the deadline could make a more complete artifact less useful.
N17Q modeled delivery deadline, minimum evidence, and allowed partial outcome. At the decision point, it could publish nothing, preserve locally, or prepare a qualified version according to policy and user choice.
The stop condition named which predicate the deadline prevented and did not describe missing evidence as failure.
Time became part of usefulness rather than a blind process kill.
Stop rules were versioned product contracts
Changing a query limit, visibility window, or progress definition could alter whether a run continued. Editing these values invisibly made historical behavior difficult to explain.
N17Q stored rule revision with checkpoints and final account. New versions applied prospectively under policy. Historical traces could be reevaluated in fixtures without rewriting their original stop.
Regression scenarios covered the boundary on both sides. Operators saw when a new rule changed pending work.
Stopping behavior evolved with the same discipline as execution behavior.
Agents could over-edit a good artifact, add redundant research, or keep reorganizing after requirements were satisfied. No forbidden effect was involved; quality and review burden still worsened.
N17Q used accepted artifact predicates, change magnitude, marginal evidence, and user preferences to offer a review checkpoint. It did not freeze creative work through a hard universal threshold.
The model could explain one more valuable revision. Repeated low-value movement became visible.
Useful autonomy knew when additional capability was reducing the value already created.
Some boundaries required another role or system rather than task closure. N17Q created a scoped escalation with evidence, deadline, requested decision, and preserved state.
The current executor stopped. The parent workflow became awaiting authority or expertise. If no response arrived, an expiry rule produced a final handoff.
Escalation did not transfer all tools or context. The recipient received the minimum eligible package.
The system could stop one form of work while keeping the goal responsibly alive.
Monitoring did not count as progress
A dormant run could watch for a receipt or source without repeatedly waking a model. N17Q registered event subscriptions and logical deadlines in the scheduler.
Health checks and duplicate notifications were operational activity, not task progress. They consumed their own budget. Only a relevant state change reopened planning.
The interface showed waiting and last observation without an animated fiction of ongoing thought.
Autonomy became quieter when the product separated watching from working.
At termination, N17Q recorded applicable rule, last progress event, remaining budgets, unresolved predicates, preserved artifacts, active effects, and eligible resume triggers.
This stop receipt entered the outcome account and regression corpus. A later reviewer could tell whether the system stopped because of evidence saturation, policy, time, or implementation failure.
If the rule fired incorrectly, the defect had a small witness. If it worked, the product could demonstrate containment.
Stopping became a transition with the same accountability as acting.
Operators could tune without creating loopholes
Real workloads revealed thresholds that were too strict or too generous. N17Q exposed configuration through reviewed policy, not ad hoc worker flags.
A change named scope, reason, expiry, and affected scenarios. It could narrow an active run immediately. Expanding a limit required current authority and did not reset prior usage or unknown effects. Shadow evaluation showed how historical fixtures would behave.
Tuning improved usefulness while the semantics of progress and consequence stayed intact.
The escape from one bad stop never became a permanent bypass around every future one.
A well-formed handoff let another person or future run begin from artifacts, exact blockers, receipts, and resume triggers. A timeout dump forced rediscovery.
N17Q measured how often resumed tasks reused preserved state, how much verification remained eligible, and whether unknown effects avoided duplicate action. These outcomes fed product design without rewarding endless continuation.
Stopping early enough to preserve coherent state often saved more time than one last speculative attempt.
Bounded autonomy accumulated knowledge instead of merely accumulating events.
The endless run found its boundary
In the repaired scenario, two authoritative sources remained unavailable, publication was denied for the data class, and the existing external intent stayed unknown through its query window.
N17Q preserved the local report with explicit evidence gaps, removed equivalent publication routes, spent the protected query reserve, and sealed the run. The final account named one unresolved effect and one user decision: inspect the destination manually or abandon delivery.
No fifth status query or new model could improve those facts. The task stopped after the last meaningful state change.
When a later receipt arrived, N17Q appended it and offered a fresh, bounded decision. It did not pretend the earlier stop had been wrong.
The most useful autonomy is not the system that keeps moving until something calls it successful.
It is the system that can pursue a goal, recognize the limits of evidence and authority, preserve what it has earned, and stop at the exact point where another action would create activity without progress.