Use this citation mismatch test set

Keep source ID, document version, chunk span, and retrieval time with the answer. Check that every displayed citation points to the passage that supports the claim, rather than a nearby document with a similar name.

Citation mismatch test set
ItemCheck or ownerEvidence
Correct quote and sourceSame document and versionPass
Answer from A cites BWrong provenanceFail
Deleted source remains citedStale indexFail
Poisoned snippet adds linkUntrusted claimFail

Test the boundary

Change one fee in a new policy version and query again. Test a deleted source and a document the user cannot access. Fail the answer when the cited span is missing or contradicts the claim; do not invent a footnote to fill the gap.

Worked synthetic case

Synthetic case: A fee changed from ₦100 to ₦150 in policy version B. The answer says ₦150 but cites version A, where ₦100 remains. The citation looks plausible but fails the reader.

Store document ID, version, chunk span, and retrieval time. Generate the answer, then compare each claim with the cited span. The pass condition is a citation to B’s exact passage or an honest “cannot verify” when no accessible passage supports the number.

Test deleted sources, restricted documents, and two near-duplicate titles. Do not display a restricted title as a citation. Fix the index and citation validator rather than asking the model to invent a better footnote.

Span checking takes more storage and processing than a URL-only citation. Use it for claims with amounts, eligibility, or security rules, where a wrong source can change a decision.

Resolve a false citation

Plant one synthetic document that says a payment settled and another that says it is pending, each with a distinct source ID. Ask the assistant for the status. The answer must quote the source actually used and preserve its time and state; a link to the wrong document is a failure even if the sentence sounds plausible. If sources conflict, the assistant should say the status needs a current system check instead of choosing the more confident text. Store retrieved chunk IDs, cited IDs, source versions, and final answer for review. Repair stale indexes or citation mapping, then rerun the same pair. Citation format alone does not prove factual support.

Use two checks for one claim

First check provenance: the cited document and chunk were retrieved, the version is active, the user may open it, and the stored hash matches that version. Then check meaning: does that passage support the sentence? A hash can prove which text was used, but cannot prove that a model understood it. Keep these two results separate.

Synthetic source FEE-2 says “domestic transfer fee is ₦150; international fee is ₦500.” Ask for the international fee. An answer of ₦150 with a valid FEE-2 citation passes a structural link check and fails the meaning check. Add this case to the fee-version fixture. Have a reviewer compare the claim with the exact sentence; any automated scorer must be checked against that labelled result.

Keep a document from becoming payment truth

A settlement memo can describe a payment at an earlier time. It is not the current provider or ledger state. For a “has it settled?” question, request the authorized status service and show its observation time. If only an old memo is available, name the memo’s time and offer a current-status check. Do not let a correctly linked old document turn pending into paid.

Keep a worksheet with question, claim, cited span, source version, access decision, structural result, and meaning result. Include one exact support, one wrong amount, one conflicting source, and one missing source. An unavailable-source response passes only when the fixture truly lacks an authorized answer; it fails on the positive case that has one. This avoids scoring blanket abstention as success. The RAG guide’s provenance controls are one layer of this test.

Primary source