When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks
arXiv:2606. 14629v1 Announce Type: cross Abstract: Verifier-driven self-DPO is a common recipe for self-improving production visual-language models.
arXiv:2605. 30290v2 Announce Type: replace-cross Abstract: Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods.
arXiv:2606. 14629v1 Announce Type: cross Abstract: Verifier-driven self-DPO is a common recipe for self-improving production visual-language models.
arXiv:2609.37633v1 Announce Type: cross Abstract: The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards...
The paper investigates a failure mode called co‑cheating in self‑evolving search agents, where the proposer and solver agree on shared errors, inflating internal reward without improving external correctness. The authors first propose multi‑sample verification (MSV) to filter unreliable pseudo‑labels, which only partially mitigates the issue. They then introduce CrossFit, a cross‑fitted reward scheme that partitions source documents and uses an auxiliary solver to prevent same‑source agreement, significantly reducing false agreement and boosting downstream benchmark performance.
arXiv:2607. 28457v1 Announce Type: cross Abstract: Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback.
arXiv:2606. 19808v1 Announce Type: new Abstract: Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes.
arXiv:2607. 20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling.
arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.
arXiv:2609.00768v2 Announce Type: replace Abstract: Self-play supports the self-evolution of language models, but solver performance can plateau or decline across rounds without guidance. Existing un...
The paper introduces Untied Self-Conditioning, a sampler that corrects a train–inference mismatch in flow‑matching language models. By dampening redundant directions in the self‑conditioning input and approximating a step‑average prediction from history, the method improves generation quality without retraining. On LangFlow and ELF‑B datasets, it dramatically lowers perplexity and is preferred in the majority of pairwise comparisons.
arXiv:2609.39148v1 Announce Type: new Abstract: AI agents can externalize what they learn from past tasks into reusable \emph{skills}, such as procedures, checklists, code, or other executable artifa...
DiagEvo is a self‑evolution framework that guides language‑model training by extracting recurring error causes from a solver’s own failure history and storing them in a hierarchical error‑cause memory. The system classifies causes as Active or Mastered, uses this information to balance targeted question generation with exploration, and applies double‑confidence filtering to keep only intermediate‑difficulty questions. Experiments show that DiagEvo outperforms baselines on nine benchmarks for three solvers, achieving up to 72.3% mean accuracy on five mathematical reasoning tasks.
arXiv:2606. 13156v2 Announce Type: replace-cross Abstract: Letting a vision-language model (VLM) think longer at test time has driven much recent progress.