The paper introduces a penalty‑aware evaluation framework for Retrieval‑Augmented Generation (RAG) systems that uses asymmetric scoring, knowledge‑gap canaries, and a failure‑attribution pipeline. Applying this framework to three commercial RAG products and a baseline on SimpleQA‑Verified, the authors find that while overall accuracy is similar across systems, canary violation rates vary dramatically, showing that systems differ more in when they answer than in what they answer. The study demonstrates that penalty‑aware scoring can reorder system rankings and is robust across different penalty settings.
By Alden Do Rosario, Hussein Younes, Felipe Pires
The paper argues that evaluating continual knowledge‑updating methods solely at a final checkpoint and a single adapter rank can be misleading. By fixing a periodic hierarchy and comparing it to cumulative replay on a 24‑month Wikidata stream, the authors show that the apparent best method changes depending on the evaluation month, replay LoRA rank, and query formulation. They recommend reporting performance trajectories and capacity sweeps, and only declaring a robust winner when the ranking remains stable across the evaluation region.
By Heejin Choi
The paper investigates intrinsic self‑correction, where a language model revises its own answer without new evidence. Across 29 open‑weight LLMs on BoolQ, GSM8K, and Corr2Cause, the study tracks how revisions change correctness, revealing that while some models improve significantly, others lose a notable fraction of correct answers. The authors compare three runtime strategies—keeping the initial answer, always accepting the revision, and selectively gating revisions—and find that the best approach depends on the model and task, suggesting that self‑correction should be treated as a revision policy rather than a uniformly beneficial second pass.
By Tianzhu Zhang
The paper investigates how sequential knowledge editing can degrade a language model’s ability to discern reliable evidence from unreliable evidence without affecting overall accuracy. Using a conservatively tuned LoRA on Qwen2.5‑7B‑Instruct, the authors show that after 1,000 edits the model’s arbitration score for untouched facts drops by 36%, leading to higher error rates on its most confident decisions, while MMLU accuracy remains unchanged. The study also finds that in some model‑method combinations, sequential edits can reduce MMLU to chance levels even though edit success and locality remain perfect.
By Atul Anand
DynaKRAG is a unified framework that learns a state‑conditioned policy to control evidence acquisition in multi‑hop retrieval‑augmented generation. It uses a deterministic validity layer to build an action set, a learned continuation gate to decide between generating an answer or gathering more evidence, and an advantage scorer to rank evidence operations by predicted gain. Across HotpotQA, 2Wiki, and MuSiQue with various backbone models, DynaKRAG achieves top EM and F1 scores, improves token and retrieval efficiency, and enables terminal evidence compression that reduces context size while boosting answer quality.
By Chenyu Zhou, Yaqi Wu, Xiaolei Guo, Jiaqi Huang, Xianfa Zhang, Junxu Zhang, Zhuo Yu, Zhubo Shi, Jianghao Lin, Dongdong Ge
ReDraft is a reference‑driven revision method for continual post‑training of large multimodal language models. It uses the model’s own incorrect outputs as references, revises them, verifies the revisions, and fine‑tunes on the accepted ones, thereby combining explicit supervision with policy proximity. On tasks such as Counting, Clock Reading, and Jigsaw, ReDraft outperforms standard supervised fine‑tuning and on‑policy methods, achieving higher target‑task gains while dramatically reducing forgetting.
By Zhihao Zhang, Mingqi Wu, Qiaole Dong, Enyu Zhou, Shuo Li, Boyang Liu, Jiazheng Zhang, Honglin Guo, Xin Guo, Shaofan Liu, Junzhe Wang, Dingwei Zhu, Zhiheng Xi, Minlong Peng, Yuan Hua, Qi Zhang, Tao Gui, Xuanjing Huang