arXiv:2608.30426v1 Announce Type: new
Abstract: Current dialogue systems struggle with dynamic information retrieval, often leading to hallucinations and lower response accuracy. We address this by a...
By Markel Ferro, Oier Lopez de Lacalle
arXiv:2609.07093v2 Announce Type: replace
Abstract: Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing external knowledge and has been widely a...
By Yifan Wang, Xinkui Lin, Yongxiu Xu, Shen Gao, Ruochen Yang, Kun Huang, Yubin Wang, Jie Wu, Wei Liu, Jian Luan, Hongbo Xu, Shuo Shang
arXiv:2607. 19345v1 Announce Type: cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier.
By Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang
arXiv:2608. 05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer.
By Purbesh Mitra, Sennur Ulukus
arXiv:2605. 12213v2 Announce Type: replace Abstract: LLM-based conversational AI agents struggle to maintain coherent behavior over long horizons due to limited context.
By Jiazhou Liang, Armin Toroghi, Yifan Simon Liu, Faeze Moradi Kalarde, Liam Gallagher, Scott Sanner
The paper introduces DATPO, a Difficulty‑Adaptive Sentence‑entropy‑guided Tree‑structured Policy Optimization method designed to improve reasoning coverage in Reinforcement Learning with Verifiable Rewards (RLVR). It builds on three design principles: adaptive difficulty rollouts, tree‑based rollouts, and sentence‑entropy‑guided forking to enhance semantic diversity. Experiments on mathematical reasoning benchmarks show that DATPO outperforms existing baselines, particularly in pass@k, leading to better test‑time scaling performance.
By Youngjun Yu, Sanghwan Jang, Hwanjo Yu