Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.086343
Special Issues
Table of Content

Open Access

ARTICLE

Do LLMs Know When Evidence is Insufficient? An Evidence Sufficiency Benchmark for Answer-Abstention Calibration in Retrieval-Augmented Generation

Hantian Zhang1, Wentai Wu2,*
1 Department of Mathematics, Jinan University, Guangzhou, China
2 Department of Computer Science, Jinan University, Guangzhou, China
* Corresponding Author: Wentai Wu. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.086343

Received 28 May 2026; Accepted 29 June 2026; Published online 31 July 2026

Abstract

Large language models (LLMs) are increasingly used in retrieval-augmented generation (RAG) systems, where they are expected to answer questions based on retrieved evidence. In many cases, however, the right behavior is not to answer. A model should abstain when the evidence is insufficient, irrelevant, or contradictory. Existing evaluations mainly focus on final-answer accuracy, and they often pay less attention to whether models can recognize evidence quality before responding. To study this problem, we propose the Evidence Sufficiency Benchmark, a five-level benchmark for evaluating answer-abstention calibration. The benchmark covers evidence conditions from L1 Full Support to L5 Conflicting Evidence, including fully supportive, partially supportive, irrelevant, absent, and conflicting evidence. We evaluate seven LLMs from five families on the full L1–L5 gradient under three prompting strategies. The results show that current LLMs still have clear limitations in evidence-based abstention. Under L5 conflicting evidence, all evaluated models show high over-answer rates, ranging from 65% to 91%. The evidence sufficiency curves show that models reduce their answer rates as evidence quality decreases, but their abstention behavior remains unreliable. Chain-of-thought prompting improves abstention for some models, although the effect is not consistent across model families. Human validation on 200 samples further supports the reliability of the automatic evaluation. Overall, our findings suggest that current LLMs still struggle to recognize when evidence is insufficient in RAG settings.

Keywords

Large language models; retrieval-augmented generation; evidence sufficiency; abstention calibration; over-answering; benchmark evaluation
  • 368

    View

  • 17

    Download

  • 0

    Like

Share Link