CoVer: Conflict-Aware Claim Verification
arXiv:2609.00508v1 Announce Type: new Abstract: Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative ne...
arXiv:2606. 26437v1 Announce Type: cross Abstract: Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist.
arXiv:2609.00508v1 Announce Type: new Abstract: Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative ne...
Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative news sources. To capture this challenge and support...
The paper introduces a taxonomy of six user challenge types and a four-layer framework to analyze how large language models respond to user disagreement. Using a dataset of 2,310 challenge scenarios and 32,340 responses from 14 models, the study finds that models often validate users (85%) while still maintaining their original claim (65%). It also reports that models frequently apologize (33%) and transfer authority in advice contexts, with significant variation across model types and task domains.
arXiv:2607. 01251v1 Announce Type: cross Abstract: Debate, where AI agents argue opposing positions, has emerged as a key approach to scalable oversight.
Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contrad...
arXiv:2508. 01273v3 Announce Type: replace Abstract: Explicit knowledge conflicts, occurring when retrieved contexts contain contradictory information, pose a fundamental challenge for Large Language Models (LLMs) as they integrate increasingly diverse data sources.
arXiv:2605. 17301v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) systems implicitly assume mutual consistency among retrieved documents -- an assumption that frequently fails in practice.
The paper introduces a taxonomy of six types of contextual knowledge conflicts—factual, inferential, temporal, granularity, perspective, and ambiguity—and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine large language models reveal that current models struggle to resolve these conflicts, exhibit a bias toward earlier evidence, and show latent awareness of conflicts in their internal representations. The authors propose a training‑free, label‑free steering method that adjusts activations to better incorporate evidence, consistently improving reasoning accuracy and producing higher‑quality, balanced summaries on the dataset.
arXiv:2609.15561v1 Announce Type: new Abstract: Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy?...
The paper examines how large language models resolve conflicts that arise within contextual knowledge, rather than between internal knowledge and external context. It introduces a taxonomy of six contextual conflict types and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine LLMs reveal persistent shortcomings in conflict resolution, uncover a bias toward earlier evidence, and propose a training‑free steering method that improves accuracy and summary quality.
arXiv:2606. 23989v1 Announce Type: cross Abstract: End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify.
arXiv:2609.24028v1 Announce Type: new Abstract: Generating coherent meta-reviews from multiple peer reviews is challenging when reviewer evidence conflicts and varies in reliability. Existing approac...