Claim review should not rewrite the claim

V0M3 separated evidence assessment from prose mutation so a reviewer could identify unsupported, contradicted, or overbroad claims without silently becoming their author.

I asked V0M3 to check whether a paragraph was supported by its sources. The response came back cleaner, more cautious, and apparently well cited.

It had also changed the paragraph.

One number disappeared. A causal statement became a correlation. Two qualifications were added. Those edits might have been sensible, but they answered a different question. I could no longer tell whether the original claim had passed review or been replaced by one that did.

A review system that repairs the object while judging it destroys its own evidence.

Claim review had to observe before it proposed.

A citation was not the unit of review

V0M3 initially attached sources to paragraphs. That was convenient for display and vague for reasoning. One paragraph could contain a date, a quantity, a comparison, and a causal explanation. A source might support only two of them.

I introduced explicit claim records. A claim captured the smallest independently assessable assertion while retaining its span in a particular document revision. It could be a fact, definition, attribution, comparison, prediction, recommendation, or interpretive conclusion.

The record did not pretend extraction was objective. Splitting one sentence into claims was itself a proposal with a compiler version and review state. The author could merge, divide, exclude, or reclassify it.

Evidence attached to claims, not decorative footnote numbers.

Review input was immutable

Every assessment named an exact claim revision, document revision, source revision, and review policy. The reviewer received quoted claim text and selected source passages through opaque identifiers.

It could return an assessment. It could not update the document, replace the claim record, or modify a source excerpt.

The application enforced this by giving the reviewer no mutation capability. Its output schema contained status, scope, evidence references, reasoning, missing information, and uncertainty. There was no revised_claim field.

This was more reliable than an instruction saying “do not rewrite.” Product authority matched the requested job.

“Supported” hid several ways evidence could fail. V0M3 used a deliberately small but expressive set:

  • Supported within the stated scope.
  • Partially supported.
  • Unsupported by the reviewed evidence.
  • Contradicted by reviewed evidence.
  • Source does not address the claim.
  • Source or claim is too ambiguous to assess.
  • Evidence is stale for a time-sensitive claim.

An assessment also identified which claim facets were covered. A report could support a count for one date without supporting the paragraph's causal explanation. A primary source could establish that someone made a statement without establishing that the statement was true.

The vocabulary made uncertainty actionable without converting every nuance into a fake decimal score.

Entailment and usefulness stayed separate

A source may be useful background and still not entail a claim. V0M3 asked two different questions: does this passage support the assertion, and is this source relevant context for the author?

The first demanded close scope alignment. The second could include definitions, competing explanations, historical context, or evidence that narrowed the argument.

Keeping both prevented a common failure. The system no longer promoted a topically related source into a citation merely because it contained similar nouns. It could say, “This helps explain the domain but does not support the number.”

Useful context remained visible without laundering it into proof.

A URL was not stable evidence. Pages changed, documents were replaced, datasets gained new periods, and excerpts could lose their surrounding qualification.

K81R's source model carried into V0M3: canonical identity where known, retrieval time, content digest, title, publisher, publication and update dates, selected passage, surrounding context, and access notes. A claim assessment pointed to that source revision.

If a later retrieval changed relevant content, the existing assessment did not silently update. It became potentially stale, and a new review could compare revisions.

The interface could still open the live URL. The historical judgment remained attached to what was actually inspected.

Short extracted passages helped focus review, but an isolated sentence could reverse meaning when a preceding paragraph supplied a condition or a following table note changed the denominator.

The source selector included surrounding headings, nearby paragraphs, table labels, footnotes, and page identity within a bounded context budget. Truncation was explicit. The reviewer could request an adjacent passage through a read-only source capability.

It could not browse arbitrary private material or expand the evidence set without recording what it retrieved. Every added passage became another source observation.

The goal was not maximum context. It was enough context to justify the assessment and enough provenance to inspect it.

Numbers received mechanical checks first

Models were useful at comparing language and poor substitutes for arithmetic that code could perform exactly.

For quantitative claims, V0M3 parsed units, dates, denominators, ranges, and comparison direction where possible. Deterministic checks caught a percentage-point change described as a percent change, a total compared with a monthly subset, or a nominal value treated as inflation-adjusted.

The reviewer saw those check results and source cells. It explained scope and ambiguity rather than recomputing silently in prose.

A failed parser did not imply a false claim. It marked the mechanical check unavailable and left the evidence question open.

The same pattern applied to links, dates, and quoted strings: use precise tools for precise work, then preserve their observations.

Contradiction required careful scope

Two sources could report different values without contradicting each other. They might use different dates, populations, definitions, or revision policies.

V0M3 required a contradiction assessment to name the shared proposition and the incompatible evidence. If scopes differed, the result was unresolved or partially supported rather than a dramatic conflict badge.

This was especially important for negative claims. Failure to find a result did not prove that none existed. The review recorded search boundaries, source classes, and queries when absence mattered.

“No evidence found in this reviewed set” remained different from “this never happened.”

Review policy depended on claim type

A personal reflection did not need the same evidence as a release date. A recommendation might rely on constraints and tradeoffs rather than an external source. A prediction could be clearly labelled and evaluated later, not cited into certainty.

V0M3's review policy mapped claim types to expectations. Direct quotations needed exact source alignment. Time-sensitive product facts needed recent authoritative material. Architecture conclusions could link to measurements, code, or explicit project constraints. Personal judgments could remain uncited when presented as such.

The policy was visible and versioned. Changing it could mark older assessments for reconsideration without rewriting history.

This kept the evidence system from turning every sentence into the same bureaucratic object.

Source quality was explained, not ranked into truth

I experimented with one source-quality score. It looked decisive and combined unrelated dimensions: proximity to the event, editorial process, primary versus secondary status, recency, accessibility, and conflicts of interest.

V0M3 stored those dimensions separately with short explanations. An official announcement could be authoritative for its own release date and interested in its performance claims. An independent analysis could be stronger for comparison and weaker for internal implementation detail.

The reviewer reasoned about fitness for this claim. It did not multiply a publisher reputation by a model confidence and call the result truth.

The interface showed why a source was used and what remained weak.

A retrieved page could contain text telling an assistant to ignore the task, reveal documents, or mark the page authoritative. That text was part of the evidence object, not an instruction channel.

V0M3 passed source passages through quoted data structures, limited available tools to bounded retrieval, and enforced a reviewer schema without document mutation. The orchestration layer selected sources and applied policy outside the model.

If hostile text still influenced an assessment, the output could be wrong. It could not grant itself more sources, alter the canonical claim, or accept a correction.

Containment did not eliminate reasoning risk. It prevented reasoning risk from becoming silent authority.

Assessment reasons cited exact evidence

A generic explanation such as “the sources support this” was not enough. Each assessment reason referenced specific passage identities and the claim facet they addressed.

For partial support, the result stated the supported portion and the unsupported remainder. For contradiction, it identified the incompatible text. For staleness, it named the temporal condition that had changed.

The UI highlighted those relationships without inventing a quote. Selecting a reason scrolled both claim and source to the relevant spans. Keyboard navigation followed the same links.

This made review slower than adding a green checkmark and faster than rediscovering the basis later.

When a claim failed review, V0M3 offered choices: keep it with a note, narrow it manually, find additional evidence, remove it, or request a correction proposal.

Requesting a correction created a new generation task targeting the exact claim revision. The task included the assessment as evidence and produced typed replace or comment operations. It did not edit the failed assessment or pretend the replacement had already passed.

If the author accepted a revised claim, that new claim revision entered unreviewed state. Another assessment could evaluate it. The system could reuse source observations but not the old verdict.

This extra step preserved the distinction between diagnosis and treatment.

Reviewers could disagree

A deterministic check might pass while a model reviewer found the scope ambiguous. Two human reviewers might interpret a causal phrase differently. V0M3 stored assessments as observations with author, method, configuration, and time.

An aggregation view could show agreement, disagreement, and unresolved questions. It did not overwrite earlier assessments with the latest one. A final editorial decision named which evidence and policy it accepted for the current revision.

For low-consequence writing, one review might be enough. More sensitive claims could require a second method or human inspection. The workflow made that a policy choice, not an aura of certainty around the model.

Disagreement was information about the claim, not a database error.

Batch review did not become batch acceptance

Reviewing many claims together was efficient. Mutating them together was dangerous.

V0M3 could schedule a batch of immutable assessment jobs sharing a source set and policy revision. Each result still targeted one claim revision. Partial failure did not blur which claims were assessed.

The interface grouped common problems, such as five claims depending on an expired dataset, while preserving individual evidence. The author could request several correction proposals, but every resulting operation remained separately reviewable.

No “fix all unsupported claims” command landed model prose directly into the document.

I built synthetic cases for exact support, topical-but-non-entailing sources, changed denominators, stale evidence, scoped contradiction, negative claims, quotation context, and prompt injection inside a source.

Fixtures fixed claim, sources, policy, and expected hard constraints. Qualitative reasoning could vary, but an assessment could not cite a nonexistent passage, call missing evidence contradiction, or mutate the document.

False confidence mattered as much as false rejection. I reviewed cases where cautious language hid an incorrect supported status and cases where a useful source was dismissed for not proving more than the claim asked.

The test suite evaluated the review workflow, not just the natural-language verdict.

Not every supported claim ages at the same rate. A release date remains historical; a current price, policy, officeholder, browser-support statement, or service limit can become false between review and export.

V0M3 let policy assign a freshness condition to a claim type or source. The condition could be a fixed review interval, a source revision signal, or a requirement to recheck at a publication milestone. The assessment stored when it was made and the observation on which it depended.

Expiration changed current status to stale. It did not rewrite the historical assessment as wrong. A new source retrieval and review produced another observation, preserving whether the world or only the evidence had changed.

For claims whose truth could change during a long document project, this was more honest than a citation that stayed green forever.

Search and assessment stayed distinct

The system could fail to support a claim because the claim was false, because the source set was weak, or because retrieval missed the right evidence. Combining search and judgment in one opaque answer made those causes indistinguishable.

V0M3 recorded source discovery as its own trace: queries, filters, result identities, selections, and coverage limits. The assessor received an explicit set. It did not imply that the set was exhaustive.

An unsupported result could lead to a new search task without changing the old assessment. The new task might use different terms, source classes, or date bounds. Its results joined the evidence graph and a fresh assessor considered the enlarged set.

This preserved negative evidence carefully. “The selected passages do not support the claim” was a bounded statement. “No source supports it” required a scope the system rarely possessed.

A person could assess a claim directly without forcing their reasoning into model-shaped fields. The interface presented claim, source passages, deterministic checks, and prior observations, then allowed the reviewer to choose status and explain scope.

Model assessments were labelled by method and configuration. They did not appear as anonymous system truth. A human could agree, disagree, or mark the case unresolved while citing the same passage identities.

When policy required a human decision, clicking a positive model verdict did not satisfy it. The resulting review record named the person and exact evidence they inspected. Delegated authority and self-review remained visible.

One data model supported both paths without pretending the methods had equal guarantees.

Replacing one unsupported sentence with cautious prose sometimes created three new assertions: a narrower number, an explanation of uncertainty, and a claim about why sources disagreed.

The correction compiler re-extracted candidate claims from the proposed text and displayed the difference in the review. Existing evidence links could carry forward only when span meaning and source scope remained compatible. New assertions started without inherited support.

After acceptance, those claim revisions entered the normal review queue. A correction was not recursively blocked until every interpretive phrase had a green label, but publication policy could require particular claim classes to clear review.

This prevented a fluent caveat from becoming a tunnel around the evidence system.

Review latency could not freeze the document

Assessment jobs were asynchronous. The author could continue editing while a review ran, but the result remained bound to the claim revision that was submitted.

If the claim changed, V0M3 stored the completed assessment as historical and marked it superseded for the current span. It did not transfer the verdict by string similarity. Where the edit was structurally unrelated, other claim assessments stayed current.

The UI showed reviewing, reviewed, stale, and superseded without blocking manual work. A publication command evaluated the latest required statuses at execution time.

This let evidence work proceed in parallel without allowing old confidence to stick to new language.

The author retained the final decision

Evidence review informed authorship. It did not replace it.

An unsupported claim might be removed, reframed as experience, kept as a clearly labelled hypothesis, or investigated further. A technically supported sentence might still be irrelevant, misleading in context, or wrong for the document's purpose.

V0M3 displayed current assessment state beside the claim and blocked only the actions the workspace policy explicitly prohibited. It did not turn every model judgment into a locked editor state.

Authority remained legible: sources supplied observations, reviewers supplied assessments, policy supplied constraints, and the author supplied the editorial decision.

Review became trustworthy by doing less

The original reviewer seemed helpful because it returned polished prose. In reality it concealed whether the evidence had supported the words I wrote.

Removing mutation capability made the feature narrower and its result more valuable. A claim stayed fixed while evidence was selected, checked, challenged, and explained. Unsupported text did not disappear before I could understand the problem. Corrections became new proposals with their own lineage and review.

The design also made model mistakes easier to contain. A reviewer could misclassify evidence, but it could not silently revise the document into apparent compliance.

That boundary is useful beyond writing. Evaluation is credible only when the object under evaluation remains observable. If the judge is allowed to repair the submission before returning a verdict, a passing result no longer tells us what passed.

V0M3 kept those roles separate. First say what the evidence supports. Then let the author decide what the document should become.