Peichun Hua

arXiv preprint · 2026

Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

Suting Chen*, Peichun Hua*, Yunming Xiao

* Equal contribution.

Overview

RINSE asks whether retrieved passages contain enough evidence to answer a question before a generator runs. It targets cases where passages are relevant to the topic but omit the facts needed for a supported answer.

Method

The paper builds paired sufficient and insufficient evidence cases through passage substitution, deletion, and question swaps. RINSE combines question-part coverage, answer-span evidence, and a small language model's cross-passage sufficiency signal. It corrects scores for evidence-set size and sums three standardized signals with equal weights.

RINSE scores evidence with content coverage, answer-span evidence, and a cross-passage language model, then normalizes and combines the three signals.
RINSE combines three complementary evidence-sufficiency signals before answer generation. Figure 2 in the paper

Evaluation

Across six datasets, RINSE reaches 0.837 average pairwise accuracy, compared with 0.746 for the strongest prior method and 0.784 for a frontier model queried through an API. On 3,149 pairs formed from real retriever outputs, it averages 0.813 versus 0.674 for the strongest prior method. Median latency is 36.5 ms per record on one NVIDIA A100 GPU.

Evidence sufficiency on six paired datasets
MethodAverage pairwise accuracyLowest dataset score
RINSE0.8370.684
Self-RAG0.7460.644
Sufficient Context (API)0.7840.676
Selected rows from Table 4. Chance pairwise accuracy is 0.5. The API row uses a frontier model; RINSE and Self-RAG run locally. Table 4 in the paper

Citation

@misc{chen2026relevance,
  title={Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG},
  author={Suting Chen and Peichun Hua and Yunming Xiao},
  year={2026},
  eprint={2609.37469},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2609.37469}
}