arXiv AI

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence

arXiv:2606. 26437v1 Announce Type: cross Abstract: Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist.

arXiv AI
Sep 2

CoVer: Conflict-Aware Claim Verification

arXiv:2609.00508v1 Announce Type: new Abstract: Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative ne...

By Shuning Zhang, Dai Shi, Bohao Chu, Hui Wang, Yuwei Chuai, Yifan Wang, Jingruo Chen, Simin Li, Xin Yi, Hewu Li
arXiv AI
Sep 10

How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

The paper introduces a taxonomy of six user challenge types and a four-layer framework to analyze how large language models respond to user disagreement. Using a dataset of 2,310 challenge scenarios and 32,340 responses from 14 models, the study finds that models often validate users (85%) while still maintaining their original claim (65%). It also reports that models frequently apologize (33%) and transfer authority in advice contexts, with significant variation across model types and task domains.

By Riyadh Alnasser, Yusuf M\"ucahit \c{C}etinkaya, Sumin Zhao, Tu\u{g}rulcan Elmas
arXiv Computation and Language
Sep 4

Large Language Models in Resolving Contextual Knowledge Conflicts

The paper introduces a taxonomy of six types of contextual knowledge conflicts—factual, inferential, temporal, granularity, perspective, and ambiguity—and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine large language models reveal that current models struggle to resolve these conflicts, exhibit a bias toward earlier evidence, and show latent awareness of conflicts in their internal representations. The authors propose a training‑free, label‑free steering method that adjusts activations to better incorporate evidence, consistently improving reasoning accuracy and producing higher‑quality, balanced summaries on the dataset.

By Xinye Yang, Zhenyang Liu, Ruisi Li, Yuanyuan Lei
arXiv Computation and Language
Sep 15

Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

arXiv:2609.15561v1 Announce Type: new Abstract: Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy?...

By Sarra Gharsallah, Adele Robaldo, Mariia Tokareva, Giovanni Gatti Pinheiro, Ilyana Guendouz, Rapha\"el Troncy, Paolo Papotti, Pietro Michiardi
Hugging Face Trending Papers
Sep 2

Large Language Models in Resolving Contextual Knowledge Conflicts

The paper examines how large language models resolve conflicts that arise within contextual knowledge, rather than between internal knowledge and external context. It introduces a taxonomy of six contextual conflict types and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine LLMs reveal persistent shortcomings in conflict resolution, uncover a bias toward earlier evidence, and propose a training‑free steering method that improves accuracy and summary quality.