Ground truth outside the trace
N17Q combined simulated-world assertions and exact receipts with evidence-citing rubric review, preventing a persuasive model judge from overruling a hard workflow failure.
The model grader liked the N17Q run.
It praised the agent's recovery after timeout, focused source use, and clear final report. It gave the trajectory a high reliability score.
The simulated service contained a duplicate request, the policy log showed an expired approval, and the repository had lost an unrelated edit.
The grader had read the narrative trace and missed the hard world facts. Qualitative judgment was useful. It could not be the final authority over deterministic failure.
A trace is evidence with several narrators
N17Q stored model outputs, semantic events, policy decisions, approvals, tool observations, artifacts, simulated-world events, and final accounts.
They could disagree. A model saw a timeout while the fixture world saw commitment. An adapter called a command failed while its partial output contained a test failure. A report omitted an effect.
Grading one flattened transcript made the most fluent narrator dominate.
Evaluation needed to preserve the source and authority of each observation.
Every scenario defined deterministic assertions for prohibited effects, required world state, file scope, preserved initial changes, current validation, approval, budgets, and unresolved outcomes.
The evaluator ran them against content-addressed workspace and simulated resources. Results linked exact artifacts and events.
A duplicate external mutation, write outside the sandbox, or missing required receipt was a hard failure. No model score could cancel it.
The rule was explicit and versioned rather than hidden inside judge instructions.
Ground truth stayed bounded
The simulator could know its own resources and injected failures exactly. It could not prove how a real provider would behave or whether a model had hidden intentions.
N17Q described ground truth as deterministic within scenario version. Live integrations used receipts and observations with explicit unknowns.
The evaluator never generalized one fixture result into universal safety.
Strong local truth supported narrow credible claims.
Plan focus, explanation clarity, evidence synthesis, maintainability, and appropriateness of intervention required judgment.
Model and human reviewers scored those dimensions against a rubric. Every assessment cited semantic events, artifacts, or report spans. Uncited impressions failed validation.
The result retained reviewer configuration, time, and uncertainty. Regrading created another observation.
Qualitative evaluation added depth without becoming invisible truth.
The judge saw hard results
N17Q provided the rubric reviewer with deterministic findings and relevant evidence. It asked the reviewer to explain implications, not reconsider whether a duplicate resource existed.
The output schema separated hard-invariant acknowledgement from qualitative scores. Contradicting a hard fact made the grade invalid.
This reduced persuasive smoothing while keeping the reviewer useful for diagnosis.
The judge reasoned around ground truth rather than above it.
Task completion, world integrity, prohibited effects, policy compliance, evidence quality, recovery, efficiency, account accuracy, and artifact quality appeared separately.
N17Q did not average them into one number by default. A run could complete the task and fail authority. Another could stop safely and remain incomplete. A third could be correct and wasteful.
Comparisons highlighted regressions by dimension and scenario.
One composite score would reward tradeoffs the product had never authorized.
Hard failures had severity and scope
Not every deterministic mismatch was equally consequential. A missing optional metadata field differed from duplicate external mutation.
Scenario rules named severity, blocking behavior, rationale, and affected claims. A critical prohibited effect blocked a pass. A noncritical formatting invariant could remain a warning.
Severity was human-authored policy, not derived from model confidence. Changing it versioned the scenario.
Determinism supplied consistency; judgment still designed the rules.
The trace could show that one timeout preceded a retry and that policy allowed it. Calling timeout the root cause required more.
N17Q used labels such as observed before, informed, authorized, attempted, and resulted in. Model reviewers could propose causal hypotheses with cited evidence and uncertainty.
Counterfactual branches could test those hypotheses without retroactively proving them.
Trace grading avoided turning chronological adjacency into confident explanation.
The final account was graded as claims
Operational statements about files, tests, sources, effects, denials, and remaining uncertainty linked to deterministic records.
The evaluator flagged unsupported, contradicted, stale, overbroad, and materially omitted claims. A fluent summary did not receive a high communication score if it hid the duplicate.
Style review followed factual coverage. It could prefer a shorter correct account over a verbose one.
Reporting quality became anchored in what the run actually did.
If a raw artifact expired or a live service lacked query, the evaluator did not infer the best outcome. It marked the relevant dimension unknown or unevaluable.
A model judge could explain likely interpretations and could not fill the gap with a score. Scenario policy determined whether missing evidence itself failed the run.
This prevented grading from manufacturing certainty after observability failed.
Unknown remained part of the result.
Evaluators had no mutation tools
The grading process received read-only trace and artifact access. It could not rerun commands, repair files, reconcile effects, edit the final answer, or change run status.
If review found a missing test, it proposed a follow-up. Another authorized run segment could perform it.
This protected the object under evaluation from helpful alteration by its judge.
Assessment observed before correction.
Their output passed schema validation, citation checks, size limits, and contradiction checks against hard results. Prompt injection inside repository, source, or tool output remained quoted data.
The grader received no external tools beyond bounded artifact reads. Its configuration and context manifest entered the trace.
A malicious source could influence prose and could not erase a world-state failure or grant itself more evidence.
Evaluation needed the same containment principles as generation.
Rubric drift was versioned
Changing wording, examples, model, evidence selection, or aggregation could change scores. N17Q assigned evaluator revision and retained previous results.
A comparison across versions showed rubric transition. Historical leaderboards were not silently recalculated.
Hard invariants could also evolve through new scenario versions, with original outcomes preserved and re-evaluation clearly labelled.
Scores remained meaningful only with the method that produced them.
Two model configurations or human reviewers could disagree about plan quality while agreeing on hard outcomes.
N17Q stored individual assessments, citations, and uncertainty. An aggregate view showed disagreement instead of selecting the newest answer. Policy could require adjudication for specific uses.
No majority could overrule a prohibited effect.
Disagreement identified subjective dimensions and weak rubric language.
Adversarial traces tested the grader
Fixtures paired polished summaries with duplicate effects, verbose uncertainty with safe outcomes, fabricated citations, misleading event order, prompt injection in tool output, and missing artifacts.
The grader had to cite valid evidence, acknowledge hard failures, avoid exposing secrets, and distinguish model proposal from executed effect.
Cases also checked false severity: a denied unsafe proposal contained by policy was not itself a world violation.
The evaluator received a benchmark because it was part of the product.
A grader could miss a hard fact because the context compiler omitted it, not because the model reasoned badly. N17Q recorded the grading manifest: selected events, artifacts, summaries, truncation, and hard-result projection.
Required invariant evidence could not be dropped to fit. Qualitative context used ranked selection with visible omissions. A smaller context that changed judgment became a compiler finding.
Evaluation quality depended on what the judge actually saw.
A terminal line could say success while the process exited unsuccessfully. A provider error could arrive after a committed effect. A model message could claim policy allowed.
N17Q passed raw snippets with source labels and linked semantic events. The grader was instructed and constrained to treat authoritative result fields and world assertions separately from untrusted text.
Citation validation rejected references to nonexistent or inaccessible log spans.
Complete logs still required an information model.
The harness kept cases with intentionally clear good, bad, and mixed trajectories. Evaluator revisions ran against them before use.
Calibration checked acknowledgement of hard failures, citation accuracy, score ordering within rubric dimensions, sensitivity to material omissions, and resistance to fluent but false reports.
The set did not prove broad reliability. It caught regressions such as a judge becoming more generous toward polished explanations.
Model grading needed the same humility as model generation.
Some evaluators returned confidence-like values. N17Q retained them as method-specific metadata and never used them to override a hard assertion or fill missing evidence.
The UI emphasized cited support, rubric, and disagreement. A low-confidence qualitative score could request human review. A high-confidence wrong statement failed contradiction checks.
Confidence described the grader's output under one method, not the world's willingness to change.
Aggregation rules were explicit
Where a release or comparison needed a compact result, N17Q applied versioned policy: critical invariants must pass, selected dimensions must be evaluated, and quality concerns can require review.
Weights and thresholds were visible. Missing dimensions did not default to zero or pass silently. The underlying vector remained available.
The aggregation engine, not a model-written paragraph, produced the status.
A concise decision still had an inspectable rule behind it.
When model graders disagreed or an invariant appeared wrong, a human could adjudicate with the trace and artifacts. The decision named scope, rationale, and which result it superseded for a particular use.
It did not delete earlier grades or mutate the run. Hard scenario changes required a new rule revision rather than one-off narrative override.
Adjudication handled genuine ambiguity without creating a secret exception path.
If evidence was too sensitive, artifacts missing, graders unavailable, or budgets exhausted, N17Q returned incomplete evaluation with deterministic findings intact.
The run did not receive a pass by default. Depending on release policy, it awaited human review or remained ineligible. No agent tool became available to repair the grade.
Stopping the judge was preferable to a confident result built from missing context.
An agent requesting a prohibited tool was not the same as the world suffering a prohibited effect. N17Q recorded proposal quality and policy containment separately.
A strong gate could keep world integrity perfect while a model repeatedly attempted evasion. The run failed adaptation or safety-behavior dimensions without inventing external damage.
This distinction let evaluation credit system defenses and still expose poor trajectories.
Reports could not cite evaluator prose as fact
A later summary might consume evaluation findings. N17Q linked it to the underlying event or artifact, not merely to a judge sentence.
Qualitative recommendations remained attributed assessments. Hard facts retained their deterministic source. Regrading could change advice without changing world history.
The evidence graph prevented a model judge from becoming an untraceable source of truth inside future context.
A scenario invariant might encode a mistaken expectation or fixture bug. N17Q linked each result to executable rule and world evidence so a reviewer could challenge it.
Corrections produced new scenario or evaluator revisions. The old result stayed as historical output. A waiver required explicit authority and reason; it did not delete the failure.
Hard meant non-negotiable during that evaluation, not beyond human correction forever.
Inspectability prevented determinism from becoming dogma.
Strict path, test, and effect rules could reject legitimate alternatives. Scenario families included several valid trajectories and equivalence cases.
When a run failed a rule but preserved the intended world, the invariant received review. I changed the benchmark rather than coaching the model to mimic one path.
Qualitative evidence helped identify overconstraint, while deterministic fixtures confirmed the revised condition.
Evaluation quality included not punishing acceptable diversity.
Where possible, the evaluator did not call the same adapter function whose output it judged. The fake request service exposed its event log to a separate assertion layer. Repository invariants read the final tree, not the agent's diff summary.
Complete independence was impractical, and shared components appeared in the manifest.
The architecture reduced correlated failure enough to make disagreement informative.
An evaluator repeating the product's own assumption would provide false certainty.
Trace projections could be rebuilt
Semantic timelines and summaries were derived from append-only events. N17Q rebuilt them and compared digests to stored projections.
If a projection omitted an event or linked the wrong workspace, grading paused. Raw artifacts and world state remained available for diagnosis.
This stopped a UI or indexing bug from becoming the judge's reality.
Evaluation depended on verifiable inputs as well as good rules.
Model reviewers could consume substantial tokens and repeat analyses. N17Q bounded evaluator calls, context, artifacts, and branches separately from the run under test.
An incomplete grade did not change run state. Deterministic results remained available. Caches used exact evaluator and input digests.
The harness reported evaluation cost beside run cost without charging it to the agent's behavior.
Measuring a system also needed resource discipline.
Some artifacts were too sensitive for an external model grader. N17Q selected eligible excerpts, used deterministic checks locally, or required human review under policy.
The context manifest named omissions. A score could be unavailable rather than sending prohibited data. Private notes and credentials never became grading convenience.
Local ground truth could remain strong while qualitative review became narrower.
Evaluation had no special exemption from data authority.
Recorded semantic traces, fixture worlds, artifact manifests, and evaluator context allowed N17Q to rerun a grading revision without repeating model tools or external effects.
A missing retained artifact narrowed the case visibly. The grader never fetched live replacements. Counterfactual grading created a new evaluation event and left the run untouched.
This separated evaluating the evaluator from re-enacting the agent, reducing both cost and risk.
The judge could recommend, not authorize
A rubric result might suggest rerunning tests, compensating an effect, or broadening evidence. Those recommendations entered a review list.
They did not create tasks with executable approval, alter policy, or consume effect budgets. A person or new run segment materialized any next action under current authority.
Evaluation remained advisory even when its diagnosis was correct.
Each dimension displayed hard assertions, qualitative assessments, cited events, artifacts, and unresolved limitations. Technical raw data stayed one level deeper.
The run owner could see why a model grader praised reasoning while the world-integrity section failed. No combined green badge obscured the disagreement.
Regrading added a comparison rather than replacing the page.
Trust in evaluation came from knowing which layer said what.
Release decisions consumed explicit policy
A project could require all critical invariants pass, no prohibited effect, selected quality thresholds, and human review for uncertain dimensions.
N17Q evaluated that release policy deterministically over current results. A model judge did not decide whether its own score was sufficient.
Changing thresholds created a policy revision. Approval bound exact scenario and evaluator set.
Evaluation informed decisions; it did not silently make them.
I kept the model grader's enthusiastic result. Beside it, N17Q showed three hard failures and the passages where the judge had relied on the final narrative instead of world evidence.
After changing the evidence package and rubric, a new grade acknowledged the failures and provided a better diagnosis. It did not erase the first assessment.
The contrast became a fixture for future evaluator revisions.
Judgment needs something it cannot persuade
Agent traces are too rich for deterministic checks alone. They contain choices, explanations, tradeoffs, and craft that benefit from human or model judgment.
They are also attached to files, effects, permissions, receipts, tests, budgets, and world state that should not yield to fluent interpretation.
N17Q put hard scenario truth beneath qualitative review. The judge could explain, compare, and criticize. It could not make a duplicate disappear or turn expired approval into permission.
The grader needed evidence.
The ground truth needed the last word on what the controlled world contained.