Refusal needs a useful page

K81R turned insufficient evidence into an investigative interface showing available sources, missing relationships, contradictions, and narrower questions.

K81R's safest early answer was also its least useful:

I don't have enough information to answer that question.

It avoided inventing a retry policy. It offered no evidence, no explanation of what was missing, and no path forward. The search system knew it had found a current workflow definition, a historical adapter note, and no capability contract for the named adapter. The refusal compressed all of that into a dead end.

I designed refusal as a page with state, evidence, and next actions rather than a sentence apologizing for the model.

Refusal was an answerability result

The product reached refusal before generation when evidence failed one or more requirements:

  • Necessary source coverage was missing.
  • Authority could not be established.
  • Current sources contradicted one another.
  • The question exceeded corpus scope.
  • Required evidence was inaccessible.
  • The index generation was incomplete or stale.
  • The question was too broad for one supported synthesis.

The refusal record stored the question, task clauses, evidence-set revision, corpus scope, answerability checks, and suggested research actions.

No model personality was needed to decide that the adapter capability was absent.

The page began with what was known

The first region summarized supported facts conservatively:

  • Z29C treats some timeout outcomes as uncertain.
  • The current policy depends on adapter idempotency and status-query capability.
  • The selected evidence identifies the workflow version.

Each fact linked to its source and claim-support state. The page did not withhold useful evidence merely because the final joined answer was unavailable.

This made refusal feel like bounded progress. A person could verify what the archive established before investigating the gap.

Missing relationships were specific

“Not enough context” could mean anything. The coverage model mapped task clauses to evidence.

For the retry question, the page showed:

RequirementStatus
Z29C policyCurrent source found
Adapter identityNamed in question
Idempotency capabilityMissing
Status-query capabilityMissing
Timeout state meaningCurrent source found

The missing capability contract was the reason an answer would be unsafe. The page linked searches for the adapter ID, contract registry, and related release manifests.

The system explained which evidence would change the result.

Two active documents could disagree. A generic insufficiency message would hide an important finding.

The contradiction page placed claims and sources side by side, preserved dates and authority status, and highlighted the exact conflicting conditions. It distinguished:

  • Genuine active disagreement.
  • Historical source superseded by current authority.
  • Different scopes that only appeared contradictory.
  • Parsing or metadata uncertainty requiring source review.

If authority metadata resolved the conflict, the answer could proceed with historical context. If it did not, K81R refused to choose and offered an evidence set suitable for a decision review.

Disagreement became useful output.

Asking “What is the best retry strategy?” was broader than “What does Z29C's current adapter contract say?”

The archive could support a description of its own decisions. It could not establish a universal best practice from a few project notes.

The refusal page identified the boundary and suggested narrower alternatives:

  • Describe Z29C's current rule.
  • Compare Z29C's two historical approaches.
  • Show evidence for adapter-specific retry decisions.
  • Switch explicitly to brainstorming outside the archive.

The last option changed task mode and removed the implication that archive citations supported general advice.

A useful refusal often improved the question.

Permission refusal avoided disclosure

When required evidence existed but was inaccessible, the page could not reveal its title, topic, count, or facts.

It said:

This question cannot be answered from the sources available in your current scope.

The page showed accessible related evidence and the effective corpus scope. It did not distinguish “restricted source exists” from “no source exists” where that distinction itself was sensitive.

Internal diagnostics recorded the policy decision and hidden candidate identity under appropriate access. The ordinary requester received no suggestion derived from inaccessible text.

Useful did not mean revealing why authorization said no.

If the source corpus was at revision 44 and the active index covered revision 42, a current-state query might be answerable from direct source navigation and unsafe through indexed evidence alone.

The page showed the freshness boundary and any failed ingestions. It offered to search the indexed corpus as historical evidence, wait for a complete generation, or open a known source directly if authorized.

It did not say “No answer exists.” It said the current retrieval view could not establish one.

Temporal accuracy changed refusal language.

I initially asked the model to suggest follow-ups. It produced attractive questions unrelated to available evidence.

I built suggestions from task clauses, found entities, missing relationships, and indexed facets. For example:

  • “What policy applies to adapters without status lookup?”
  • “Show the current Z29C recovery states.”
  • “Find the capability contract for adapter_no_query_v2.”

A model could help phrase them, but the underlying query had to map to a known scope and evidence need. The interface previewed filters before running it.

Follow-up suggestions became retrieval actions rather than conversation filler.

Refusal preserved the evidence set

The first chat implementation discarded failed context and started the next question from transcript history. The refusal page kept the selected evidence set and its gaps.

A new search could add the missing contract, create evidence-set revision six, and rerun answerability. Existing sources and reviewer notes remained. If the gap could not be filled, the set could close as Insufficient and support a documented negative result.

The research workflow moved forward through state, not repeated prompting.

I avoided one Retry button. Available actions came from the refusal reason:

  • Missing source: search or broaden an allowed filter.
  • Superseded-only evidence: open current relationship or ask a historical question.
  • Active contradiction: compare and start a decision record.
  • Inaccessible evidence: request appropriate access outside K81R if such a process exists, without revealing hidden detail.
  • Incomplete index: wait, inspect ingestion, or use direct source.
  • Overbroad question: select a narrower task.
  • Unsupported claim: return to evidence selection.

Retry generation was absent because the evidence had not changed.

The page made safe progress the default action.

A dense warning panel would be difficult to scan. The page used a stable heading hierarchy:

  1. What can be established.
  2. Why the full answer is unavailable.
  3. Evidence already found.
  4. Missing or conflicting requirements.
  5. Available next steps.

Status was conveyed through text and icons, not color alone. Tables had clear headers. Source links named document and revision. Focus moved to the refusal heading after submission and returned predictably from source inspection.

The page was a research document, not an error toast.

Tone followed specificity

I removed phrases such as “As an AI” and generic apologies. The product could state its evidence condition directly.

Useful language was restrained:

  • “The selected sources do not identify this adapter's status-query capability.”
  • “Two current decisions disagree about the deadline.”
  • “The available index ends before the requested date.”
  • “This question requires sources outside the selected project.”

The page did not perform humility. It exposed a bounded reason that a person could verify.

Refusal could itself be wrong

An overly strict answerability rule refused questions the evidence actually supported. A missing parser relationship could create a false gap. Permission metadata could be stale.

I added Challenge refusal. A reviewer could point to a source, identify an incorrectly required clause, or reclassify the task. The challenge produced an annotation and potential evaluation case; it did not mutate the old refusal silently.

Evaluation measured false refusal as well as unsafe answer. Safety achieved by declining everything was not product quality.

The useful-page design made mistaken refusals easier to diagnose.

For each no-answer fixture, I recorded:

  • Correct refusal category.
  • Facts safe to show.
  • Information that must remain hidden.
  • Specific missing requirements.
  • Misleading evidence to avoid.
  • Valid narrower questions or actions.

Reviewers scored whether the page preserved progress, not only whether generation stopped. A refusal that named the wrong gap could send research in the wrong direction.

The benchmark included answerable cases to prevent a policy change from improving refusal accuracy by refusing more often.

Refusal records had a lifecycle

A refusal against corpus revision 42 could become answerable at revision 45. The record remained immutable and showed when a newer complete corpus or evidence set became available.

Rerunning produced a new answerability record linked to the old one. The interface could say, “This gap was resolved by the adapter contract added in revision 45.” It did not rewrite the earlier result as if the evidence had always existed.

Closed insufficient sets remained useful historical artifacts when they explained why a decision could not be made at the time.

“No current source in this corpus defines a retry count” was different from “A source may exist, but retrieval did not find it” and “Two sources disagree.”

K81R used distinct states:

  • No supporting source under a complete bounded corpus.
  • Retrieval gap or incomplete indexing.
  • Missing required relationship.
  • Conflicting evidence.
  • Inaccessible scope.
  • Answer outside the corpus's intended authority.

The wording never upgraded one into another. Absence of retrieved evidence was not proof of absence.

For complex refusal records, a model could draft a concise explanation from structured reasons. The application still chose the reason, allowed facts, hidden details, and actions.

The generated explanation passed the same claim and access checks as an answer. A deterministic template remained available. If generation failed, the page lost polish, not meaning.

This was the right division of labor: language help at the edge, policy and evidence state beneath it.

A refusal could conclude the task

Some questions should end with a documented gap. If no surviving source established why a 2014 prototype changed one setting, repeated generation would not recover the history.

The page could close the evidence set as Insufficient, record corpus and searches checked, preserve related sources, and state the narrow negative conclusion:

No source in the reviewed 2014 project archive establishes the reason for this setting.

That was more useful than invented recollection and more durable than an empty search screen.

The result was less conversational and more helpful

The chat response had treated refusal as a break in the assistant's performance. The page treated it as a state in an evidence workflow.

It showed what was known, what was missing, why the gap mattered, which evidence conflicted, and what safe actions remained. It preserved the work already done and could evolve when the corpus changed.

K81R became less eager to produce a sentence and better at helping a person finish an investigation.

Refusal needs a useful page because uncertainty is not the absence of a product state. It is a state with evidence, limits, and possible next steps.

The page made that state actionable.

The final scenario asked for an adapter rule that did not exist in the complete authorized corpus. The page showed the current general policy, documented every contract collection searched, named the missing adapter-specific classification, and offered to close the inquiry as insufficient. No Retry answer button appeared. Saving the result produced a durable gap record that later ingestion could resolve without erasing what had been knowable at the time.

“Broaden search” sounded harmless and could add another project, historical documents, or a more sensitive collection. The action previewed the current and proposed scope, authorization effect, expected index generation, and why the change might fill the gap.

A person confirmed the change, and the new query result retained both scopes in history. The page did not automatically widen from current decisions to every draft merely to avoid refusing.

Scope expansion was a research decision with provenance, not a retry parameter.

Partial answers needed explicit boundaries

Some questions contained independently answerable parts. “What does Z29C do after timeout, and why was that policy chosen?” might have current policy evidence but no surviving rationale.

K81R could answer the first clause, refuse the second, and present both under one coverage map. The prose did not imply that the known behavior established the unknown reason. Each partial claim retained its source and review state.

This was more useful than refusing the whole question and safer than inventing connective explanation. Partial answers required clause-level structure; a generic chat response could blur the boundary easily.

If verification infrastructure failed, the system initially returned “insufficient evidence.” That misreported an operational failure as an epistemic result.

I separated states:

  • Evidence is insufficient.
  • Answerability check could not complete.
  • One retrieval lane is unavailable.
  • Source authority could not be reached.

Operational failure offered retry or later resumption with a stable job identity. It never became evidence that the corpus lacked an answer.

The page needed to say what the system did not know and why it did not know it.

“How did my approach to reliability change?” crossed many years, projects, and concepts. Refusing as too broad was safe and wasted the question's value.

K81R proposed a plan grounded in archive structure: define periods, retrieve representative decisions and postmortems, select one contradiction or revision per phase, and synthesize only after coverage review. Each step created or extended an evidence set.

The plan did not answer the question prematurely. It converted an oversized synthesis into inspectable retrieval tasks with completion criteria.

Refusal could therefore be the beginning of a larger inquiry rather than a smaller prompt.

Repeated refusal revealed corpus problems

If many queries failed on the same missing adapter relationship, the solution might be better documentation or metadata, not more search suggestions.

I aggregated structured gap categories from synthetic evaluation and explicitly saved refusal records, not arbitrary private queries. The report showed missing authority links, unsupported document types, recurring parser failures, and stale relationship metadata.

Fixes entered the corpus through reviewed revisions. K81R did not generate source documents automatically from gaps.

The refusal surface became feedback about the knowledge system while respecting bounded data collection.

A refusal could be bookmarked and shared

The page had a stable URL tied to its immutable answerability record. An authorized reviewer could open the exact evidence, scope, and gap later. Sharing revalidated every source and could produce a disclosure-safe view if some evidence was unavailable.

A later resolved answer linked back to the refusal. This made “we could not establish this at revision 42” a durable part of the archive rather than a transient assistant message.

Negative results gained provenance and a lifecycle.

If a model failed to produce the required claim schema, that was generation failure, not a reason to refuse the question. The evidence set could remain answerable.

K81R retried parsing only under a bounded policy, offered a deterministic source summary where possible, and preserved diagnostics. The page said that drafting failed while keeping sources and answerability visible.

This distinction prevented reliability problems in one component from masquerading as principled uncertainty.

Refusal did not become a moral performance

I removed congratulatory language about being safe or responsible. The page did not ask users to admire the system for declining.

It stated the evidence boundary, showed useful material, and offered actions. When no action existed, it documented the negative result plainly. The restraint matched K81R's role as an evidence tool rather than a conversational character.

Good refusal was measured by whether a person could understand and continue, not by how cautious the prose sounded.