N17Q began evaluating what agents do after a capability is refused, revealing whether they narrow scope, seek evidence, stop well, or merely search for another route to the same effect.
#evaluation
N17Q stopped asking a second model for an unsupported verdict and built compact evidence packages, citation checks, calibrated uncertainty, and human-readable disagreement into qualitative evaluation.
N17Q separated qualitative grading from deterministic safety checks so a persuasive result could never average away an unauthorized effect, missing approval, or false completion claim.
N17Q combined simulated-world assertions and exact receipts with evidence-citing rubric review, preventing a persuasive model judge from overruling a hard workflow failure.
N17Q graded repository and simulated external invariants independently of the agent’s final narrative, catching damage a correct response could neither see nor repair.
N17Q showed why accurate answers, correct tool selection, a valid patch, and passing focused tests could still produce an untrustworthy agent run.
N17Q branched from immutable checkpoints to vary one model, policy, prompt, tool, or failure while preserving the original trace and refusing causal overclaim.
N17Q’s first benchmark rewarded correct tool choices and a polished report while missing the duplicate external consequence created between them.
V0M3 separated evidence assessment from prose mutation so a reviewer could identify unsupported, contradicted, or overbroad claims without silently becoming their author.
K81R split end-to-end answer quality into corpus, retrieval, evidence, generation, and verification stages so improvements could not conceal new failure modes.
K81R stopped rewarding answers merely for linking sources and began checking whether each cited passage actually entitled the adjacent claim.
K81R combined lexical and semantic ranks without pretending their raw relevance scores shared one calibrated meaning.
K81R's polished mistake turned retrieval, evidence lineage, answerability, and refusal into core product behavior.
The November 2022 research preview made fluent but incorrect output an interface and evaluation constraint, not a footnote to model quality.