arXiv AI By Chen Henry Wu, Aditi Raghunathan

Self-Trained Verification for Training- and Test-Time Self-Improvement

Read the original on arXiv AI →

arXiv:2605. 30290v2 Announce Type: replace-cross Abstract: Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

The paper investigates a failure mode called co‑cheating in self‑evolving search agents, where the proposer and solver agree on shared errors, inflating internal reward without improving external correctness. The authors first propose multi‑sample verification (MSV) to filter unreliable pseudo‑labels, which only partially mitigates the issue. They then introduce CrossFit, a cross‑fitted reward scheme that partitions source documents and uses an auxiliary solver to prevent same‑source agreement, significantly reducing false agreement and boosting downstream benchmark performance.

By Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen, Zeyu Zhang, Weizhi Du, Yueting Li, Tianyu Shi, Alaa Khamis