N17Q gave stronger reasoning a precise task-scoped capability set instead of broader default power, measuring whether it could adapt, ask, and stop within that boundary.
Blog, page 3
N17Q combined simulated-world assertions and exact receipts with evidence-citing rubric review, preventing a persuasive model judge from overruling a hard workflow failure.
N17Q changed a target after review and treated the once-valid decision as historical evidence, not current authority for a well-formed but stale action.
N17Q pinned server identity, tool schemas, descriptions, mappings, and the exact offered subset so replay and approval referred to the capabilities a run actually saw.
N17Q graded repository and simulated external invariants independently of the agent’s final narrative, catching damage a correct response could neither see nor repair.
N17Q showed why accurate answers, correct tool selection, a valid patch, and passing focused tests could still produce an untrustworthy agent run.
N17Q returned bounded, consequence-aware policy reasons that helped an agent choose a narrower safe path without exposing hidden capabilities or turning refusal into negotiation.
N17Q branched from immutable checkpoints to vary one model, policy, prompt, tool, or failure while preserving the original trace and refusing causal overclaim.
N17Q compared an agent’s confident final account with effect receipts and simulated world state, revealing the second request its own narrative had forgotten.
N17Q rebuilt resumable agent state from semantic events, artifacts, policy, effects, and budgets so continuation did not depend on one opaque conversation identifier.
N17Q treated summaries as derived artifacts and verified that denials, unknown effects, approvals, constraints, evidence, and budgets survived before a long run resumed.
N17Q stopped sequences of individually reasonable model turns and tool calls whose repetition stayed below every per-call limit while exhausting the run as a whole.
N17Q revalidated identity, authority, arguments, approvals, budgets, tool mappings, and world preconditions after planning because all could change while an action waited.
N17Q constrained filesystem, processes, network, credentials, time, and writable state while documenting what its local agent sandbox could not prove about hostile code.
The May 2025 coding-agent launch made background repository work tangible; N17Q treated files, local changes, commands, tests, instructions, network, and review as governed run state.
A missing N17Q replay fixture once reached a live read service; the repair removed fallback capability and made incomplete simulated worlds stop visibly.