The paper investigates when it is better to return an existing draft answer or revise it using retrieved evidence in retrieval‑augmented QA systems. By grading both the draft and its candidate revision with the same correctness judge, the authors define a paired effect called recoverability and train policies to predict it before revision. Experiments on 25,870 open‑domain questions show that a recoverability‑based scorer outperforms a draft‑correctness scorer across multiple Llama setups, improving accuracy–revision trade‑offs and closing a significant portion of the oracle gap, though it still applies harmful revisions in a substantial fraction of cases.
By Nicholas Kashani Motlagh, Tim Anderson, Jeremy Gwinnup, Grant Erdmann
arXiv:2607. 28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers.
By Yefan Tao, Gerald Friedland, Madhusudhanan Chandrasekaran, Luyang Kong
The paper introduces a fixed‑budget revision protocol that uses deterministic verifiers to expose all remaining violations across exact‑length, lexical, and compositional constraints, thereby isolating model‑side revision behavior. Experiments on 19 open‑ and closed‑source LLMs show wide variability in controller‑level success, with some models achieving up to 99.8% success while others remain below 20%. Controlled studies reveal that post‑training and scale affect model responses to exact feedback, but do not consistently improve exact correction, and that recurrence of earlier outputs is linked to lower recoverability.
By Haitong Jiang, Chunlin Liu, Yile Wang, Yuhong Feng
arXiv:2606. 01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers.
By Jiaming Qu, Lucheng fu, Yibo Hu
The paper identifies a specific issue in supervised fine‑tuning (SFT) of large language models called factual access failure, where models can recognize correct facts under constrained tests but fail to generate them in open‑ended settings. It demonstrates that SFT can cause both genuine wrong answers and expression‑level errors such as verbosity or formatting mismatches. To mitigate this, the authors propose Recall‑Anchored Distillation (RAD), a self‑distillation method that aligns the fine‑tuned model with the base model’s soft output distribution on unlabeled out‑of‑distribution text, thereby recovering lost factual recall without needing labeled data.
By Haodong Chen, Yadong Wang, Shengtao Wen, Dong Liang, Xiang Chen
arXiv:2609.36587v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using...
By Yupeng Chang, Wenxuan Zhang, Yuan Wu