arXiv:2608. 07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures.
By Shibo Chu, Yuze Liu, Tiehua Zhang, Zhishu Shen, Lianghua He, Haofen Wang, Zhijun Ding
arXiv:2608. 12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging.
By Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang
arXiv:2609.14528v1 Announce Type: cross
Abstract: Multi-Hop Knowledge Graph Question Answering (KGQA) tasks require models to assemble relational evidence along paths in a KG to answer natural-langua...
By Eduin E. Hernandez, Luis F. Garcia, Nurassyl Askar, Sergio A. Diaz, Stefano Rini
The paper introduces Causal-Counterfactual RAG, a new framework that augments Retrieval-Augmented Generation with explicit causal graphs and counterfactual reasoning. By incorporating cause‑effect relationships into retrieval and evaluating both direct causal evidence and counterfactual scenarios, the approach aims to produce more robust, accurate, and interpretable answers. This method seeks to maintain contextual coherence, reduce hallucinations, and improve reasoning fidelity compared to traditional RAG systems.
By Harshad Khadilkar, Abhay Gupta
Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when evidence is sparse, noisy, or in conflict with parametric knowledge. Existing grounding methods either attach citations after generation or encourage evidence retrieval inside the trace, but they often do not ensure that cited content is sufficient for the local inference and final answer.
arXiv:2608. 05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer.
By Purbesh Mitra, Sennur Ulukus
arXiv:2605. 12213v2 Announce Type: replace Abstract: LLM-based conversational AI agents struggle to maintain coherent behavior over long horizons due to limited context.
By Jiazhou Liang, Armin Toroghi, Yifan Simon Liu, Faeze Moradi Kalarde, Liam Gallagher, Scott Sanner
arXiv:2603.16654v3 Announce Type: replace-cross
Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, espe...
By Xiaojie Gu, Sherry T. Tong, Aosong Feng, Sophia Simeng Han, Jinghui Lu, Yingjian Chen, Yusuke Iwasawa, Yutaka Matsuo, Chanjun Park, Rex Ying, Irene Li
arXiv:2606. 20245v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive parametric knowledge and in-context learning ability, enabling them to incorporate external information provided in the input prompt.
By Huang Peng, Jiuyang Tang, Weixin Zeng, Hao Xu, Xiang Zhao
Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization.
REALHOP introduces a behavioral auditing framework to assess multi‑hop reasoning by measuring the Behavioral Necessity Rate (BNR), which quantifies how often removing targeted evidence prevents correct answers. Across five benchmarks, the framework reveals a wide gap between annotated reasoning chains and actual evidence dependence, with panel‑mean BNR ranging from 16.6% to 48.9%. By re‑binding entities, factorizing relations, adding competing paths, and placing evidence at traceable locations, REALHOP raises BNR dramatically—from 27.4% to 94.4% on MuSiQue questions—while maintaining high overall accuracy and improving performance on long‑context tasks.
By Jiawen Tao, Xiaokun Yuan, Yaoming Li, Chenxu Liu, Mengzhou Wu, Tong Yang, Maxm Pan
arXiv:2604. 01993v2 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through invalid intermediate reasoning.
By Daeyong Kwon, Soyoung Yoon, Seung-won Hwang