TY - EJOU AU - Moon, Jihoon TI - Accountable NLP for Evidence-Grounded Decision Briefings: A Critical Review and Evaluation Framework T2 - Computers, Materials \& Continua PY - VL - IS - SN - 1546-2226 AB - Large language models and retrieval-augmented generation (RAG) systems are increasingly employed to transform evidence into decision-facing briefings, alerts, and recommendations. In these settings, explainability cannot be evaluated merely by fluency, readability, or factual correctness. A briefing may be factually correct while still being unsafe if it cites sources that do not substantiate the claim, suppresses uncertainty, converts correlational evidence into causal language, recommends an unauthorized action, or leaves no auditable path for human review. This review synthesizes 104 sources spanning explainable natural language processing (NLP), faithful explanation, hallucination and factuality evaluation, RAG, citation faithfulness, uncertainty communication, causal language, human–AI interaction, engineering and regulatory decision support, and institutional accountability. It makes four contributions. First, it defines evidence-grounded decision briefings as a distinct NLP setting characterized by identifiable evidence inputs, constrained decision-facing outputs, and minimum accountability requirements. Second, it proposes a five-layer taxonomy encompassing evidence representation, explanation generation, retrieval and source grounding, verification and evaluation, and human accountability. Third, it develops an operational evaluation framework for claims, citations, uncertainty statements, causal wording, action labels, and complete briefing episodes. Fourth, it complements the conceptual synthesis with source-level trend analyses, targeted quantitative comparisons, and representative use cases drawn from generic, engineering, and regulatory contexts. The synthesis reveals that existing surveys provide critical foundations but do not jointly address five interdependent requirements: citation-to-claim entailment (whether the cited evidence supports the exact claim), causal-language discipline, uncertainty preservation, action appropriateness, and human accountability in decision-facing generated text. This review concludes with open challenges for claim segmentation, retrieval adequacy, citation-to-claim entailment, causal test suites, uncertainty preservation, accountability logging, and preference-bias-aware human evaluation. KW - Explainable natural language processing (NLP); accountable NLP; evidence-grounded generation; retrieval-augmented generation (RAG); faithful explanation; source entailment; hallucination; causal language; action appropriateness; human accountability DO - 10.32604/cmc.2026.089115