The paper examines how large language models resolve conflicts that arise within contextual knowledge, rather than between internal knowledge and external context. It introduces a taxonomy of six contextual conflict types and presents the ContextConflict dataset with 5,781 samples covering reasoning and summarization tasks. Experiments on nine LLMs reveal persistent shortcomings in conflict resolution, uncover a bias toward earlier evidence, and propose a training‑free steering method that improves accuracy and summary quality.
arXiv:2606. 20245v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive parametric knowledge and in-context learning ability, enabling them to incorporate external information provided in the input prompt.
By Huang Peng, Jiuyang Tang, Weixin Zeng, Hao Xu, Xiang Zhao
arXiv:2508. 01273v3 Announce Type: replace Abstract: Explicit knowledge conflicts, occurring when retrieved contexts contain contradictory information, pose a fundamental challenge for Large Language Models (LLMs) as they integrate increasingly diverse data sources.
By Xianda Zheng, Zijian Huang, Meng-Fen Chiang, Jiamou Liu, Yuan Fang, Michael Witbrock, Kaiqi Zhao
arXiv:2601. 09445v2 Announce Type: replace-cross Abstract: In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the model's parametric knowledge.
By Minh Vu Pham, Hsuvas Borkakoty, Yufang Hou
arXiv:2609.38799v1 Announce Type: new
Abstract: Understanding multi-perspective alternative narratives requires identifying how their information agrees, conflicts, or differs across sources. Existin...
By Eftekhar Hossain, Santu Karmaker
arXiv:2608. 13921v1 Announce Type: new Abstract: LLM agents increasingly maintain personal memory across sessions, but it can conflict.
By Lu Yang, Shusheng Xu, Zhuoran Li, Tongkai Yang, Longbo Huang
arXiv:2602. 20094v2 Announce Type: replace Abstract: As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to ground their reasoning in causality rather than spurious correlations.
By Yuzhe Wang, Yaochen Zhu, Jundong Li
arXiv:2606. 26437v1 Announce Type: cross Abstract: Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist.
By Siyi Liu, Aaron Halfaker, Dan Roth, Patrick Xia
LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or s...
The paper introduces PACE, a dataset that tests whether personalized assistants can detect hidden conflicts in user requests by integrating implicit, egocentric knowledge from a knowledge base. PACE pairs user requests with persona-specific facts, requiring models to retrieve contextual evidence before deciding if a request is inappropriate. The authors also propose PaceMaker, a multi-agent system that improves conflict detection by coordinating query reformulation, multi-hop graph traversal, and conflict-aware filtering, outperforming existing methods on the PACE benchmark.
By Yoojin Kim, Jihyoung Jang, Hyounghun Kim
The study investigates how a single politically charged framing term can influence large language models (LLMs) during translation tasks. By prompting eight models from Western, Chinese, and European origins to translate culturally attributed recipes across 17 languages under four framing conditions, the authors find that models resolve ambiguity rather than decline, with distinct behavior patterns tied to model families. Sensitivity to framing terms is consistent, showing that even subtle variations can modulate LLM behavior, raising concerns about implicit political judgments in translation contexts.
By Svetlana Gorovaia, Angelica Henestrosa, Ivan P. Yamshchikov
The paper introduces UC-Bench, a human‑annotated benchmark for detecting user‑side implicit conflicts in Human‑LLM dialogue, a problem largely overlooked compared to LLM‑side conflicts. Experiments show current LLMs struggle with these conflicts, especially when they stem from implicit incompatibilities in dialogue history. To address this, the authors propose SynUC, a constraint‑guided data synthesis method that generates a new training set, UC‑Data, which improves performance of lightweight LLMs on UC‑Bench compared to larger general‑purpose models and existing synthesis approaches.
By Jinqiang Wang, Tao Zhu, Huansheng Ning