arXiv preprint · 2026
Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG
* Equal contribution.
Overview
RINSE asks whether retrieved passages contain enough evidence to answer a question before a generator runs. It targets cases where passages are relevant to the topic but omit the facts needed for a supported answer.
Method
The paper builds paired sufficient and insufficient evidence cases through passage substitution, deletion, and question swaps. RINSE combines question-part coverage, answer-span evidence, and a small language model's cross-passage sufficiency signal. It corrects scores for evidence-set size and sums three standardized signals with equal weights.
Evaluation
Across six datasets, RINSE reaches 0.837 average pairwise accuracy, compared with 0.746 for the strongest prior method and 0.784 for a frontier model queried through an API. On 3,149 pairs formed from real retriever outputs, it averages 0.813 versus 0.674 for the strongest prior method. Median latency is 36.5 ms per record on one NVIDIA A100 GPU.
| Method | Average pairwise accuracy | Lowest dataset score |
|---|---|---|
| RINSE | 0.837 | 0.684 |
| Self-RAG | 0.746 | 0.644 |
| Sufficient Context (API) | 0.784 | 0.676 |
- retrieval-augmented generation
- evidence sufficiency
- answer abstention
- information retrieval
Citation
@misc{chen2026relevance,
title={Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG},
author={Suting Chen and Peichun Hua and Yunming Xiao},
year={2026},
eprint={2609.37469},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.37469}
}