arXiv Computation and Language

Large Language Models in Resolving Contextual Knowledge Conflicts

The paper introduces a taxonomy of six types of contextual knowledge conflicts—factual, inferential, temporal, granularity, perspective, and ambiguity—and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine large language models reveal that current models struggle to resolve these conflicts, exhibit a bias toward earlier evidence, and show latent awareness of conflicts in their internal representations. The authors propose a training‑free, label‑free steering method that adjusts activations to better incorporate evidence, consistently improving reasoning accuracy and producing higher‑quality, balanced summaries on the dataset.

Hugging Face Trending Papers
Sep 2

Large Language Models in Resolving Contextual Knowledge Conflicts

The paper examines how large language models resolve conflicts that arise within contextual knowledge, rather than between internal knowledge and external context. It introduces a taxonomy of six contextual conflict types and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine LLMs reveal persistent shortcomings in conflict resolution, uncover a bias toward earlier evidence, and propose a training‑free steering method that improves accuracy and summary quality.

arXiv AI
Jun 19

Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference

arXiv:2606. 20245v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive parametric knowledge and in-context learning ability, enabling them to incorporate external information provided in the input prompt.

By Huang Peng, Jiuyang Tang, Weixin Zeng, Hao Xu, Xiang Zhao
arXiv Computation and Language
Sep 4

PACE: Towards Surfacing Hidden Conflicts in User Requests

The paper introduces PACE, a dataset that tests whether personalized assistants can detect hidden conflicts in user requests by integrating implicit, egocentric knowledge from a knowledge base. PACE pairs user requests with persona-specific facts, requiring models to retrieve contextual evidence before deciding if a request is inappropriate. The authors also propose PaceMaker, a multi-agent system that improves conflict detection by coordinating query reformulation, multi-hop graph traversal, and conflict-aware filtering, outperforming existing methods on the PACE benchmark.

By Yoojin Kim, Jihyoung Jang, Hyounghun Kim
arXiv AI
Sep 10

We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation

The study investigates how a single politically charged framing term can influence large language models (LLMs) during translation tasks. By prompting eight models from Western, Chinese, and European origins to translate culturally attributed recipes across 17 languages under four framing conditions, the authors find that models resolve ambiguity rather than decline, with distinct behavior patterns tied to model families. Sensitivity to framing terms is consistent, showing that even subtle variations can modulate LLM behavior, raising concerns about implicit political judgments in translation contexts.

By Svetlana Gorovaia, Angelica Henestrosa, Ivan P. Yamshchikov
arXiv Computation and Language
Sep 18

Towards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue

The paper introduces UC-Bench, a human‑annotated benchmark for detecting user‑side implicit conflicts in Human‑LLM dialogue, a problem largely overlooked compared to LLM‑side conflicts. Experiments show current LLMs struggle with these conflicts, especially when they stem from implicit incompatibilities in dialogue history. To address this, the authors propose SynUC, a constraint‑guided data synthesis method that generates a new training set, UC‑Data, which improves performance of lightweight LLMs on UC‑Bench compared to larger general‑purpose models and existing synthesis approaches.

By Jinqiang Wang, Tao Zhu, Huansheng Ning