arXiv:2509. 23071v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories.
By Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King
The paper introduces CoT-Interpretability Alignment (CIA), a metric that quantifies how well a large language model’s chain-of-thought (CoT) explanations match its internal reasoning processes. Evaluated on two-hop question answering, hint intervention, and integer multiplication across three LLMs, the study finds limited alignment (44.8–75.9%) and demonstrates that post‑training with a reward combining task accuracy and parametric faithfulness can substantially improve CoT faithfulness without sacrificing accuracy. The authors provide a framework for auditing CoT faithfulness and a pathway to making explicit reasoning more trustworthy, with code and data publicly available.
By Yihuai Hong, Shauli Ravfogel, Chen Zhao, Eunsol Choi
PRO-Step introduces a step‑level process reward optimization framework for Retrieval‑Augmented Generation (RAG) that evaluates both logical validity and evidential grounding at each reasoning step. By training a generative Preference‑Based Reward Model (PRM) and using PRM‑guided value tree search to create preference pairs, the method optimizes the policy through step‑level Direct Preference Optimization. Experiments on single and multi‑hop QA benchmarks show that PRO‑STEP achieves the best average EM and F1 scores across five datasets.
By MinKeon Kim, Namjun Lee, Jaekwang Kim
arXiv:2609.21492v1 Announce Type: new
Abstract: Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely r...
By Jingyu Hu, Shu Yang, Weiru Liu, Di Wang
arXiv:2606. 29090v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems retrieve a fixed number of passages for every question regardless of its difficulty.
By Ansh Kamthan
DynaKRAG is a unified framework that learns a state‑conditioned policy to control evidence acquisition in multi‑hop retrieval‑augmented generation. It uses a deterministic validity layer to build an action set, a learned continuation gate to decide between generating an answer or gathering more evidence, and an advantage scorer to rank evidence operations by predicted gain. Across HotpotQA, 2Wiki, and MuSiQue with various backbone models, DynaKRAG achieves top EM and F1 scores, improves token and retrieval efficiency, and enables terminal evidence compression that reduces context size while boosting answer quality.
By Chenyu Zhou, Yaqi Wu, Xiaolei Guo, Jiaqi Huang, Xianfa Zhang, Junxu Zhang, Zhuo Yu, Zhubo Shi, Jianghao Lin, Dongdong Ge