arXiv Computation and Language
Sep 4

Large Language Models in Resolving Contextual Knowledge Conflicts

The paper introduces a taxonomy of six types of contextual knowledge conflicts—factual, inferential, temporal, granularity, perspective, and ambiguity—and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine large language models reveal that current models struggle to resolve these conflicts, exhibit a bias toward earlier evidence, and show latent awareness of conflicts in their internal representations. The authors propose a training‑free, label‑free steering method that adjusts activations to better incorporate evidence, consistently improving reasoning accuracy and producing higher‑quality, balanced summaries on the dataset.

By Xinye Yang, Zhenyang Liu, Ruisi Li, Yuanyuan Lei
Hugging Face Trending Papers
Sep 2

Large Language Models in Resolving Contextual Knowledge Conflicts

The paper examines how large language models resolve conflicts that arise within contextual knowledge, rather than between internal knowledge and external context. It introduces a taxonomy of six contextual conflict types and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine LLMs reveal persistent shortcomings in conflict resolution, uncover a bias toward earlier evidence, and propose a training‑free steering method that improves accuracy and summary quality.

arXiv Computation and Language
Sep 10

Who Argues What? Joint Argument-Entity Detection and Classification in Political Debates

The paper introduces DNE‑ElecDeb, an enriched version of the USElecDeb dataset that annotates Debate Named Entities (DNEs) in both argumentative and non‑argumentative spans, and defines Debate Named Entity Recognition (DNER) as a new task. It proposes Joint Argument and Entity Tagging (JAET), a generative framework that fine‑tunes decoder‑only LLMs to insert inline argument and entity tags into debate turns while preserving the original transcript. JAET achieves significant improvements in joint AM+DNER performance (+27.3% relative F1 in the untyped setting and +41.9% in the typed setting) over sequential pipelines, and these gains generalize to Persuasive Essays (+26.6% and +52.7%).

By Lucio La Cava, Stefano Francesco Monea, Sergio Greco