arXiv:2605. 17609v2 Announce Type: replace Abstract: Many inference-time language-model pipelines combine a cheap reward signal with an expensive verifier, such as exact answer checking in mathematical reasoning or hidden-test execution in code generation.
By Shaddin Dughmi, Mahdi Haghifam, Yusuf Hakan Kalayci
arXiv:2508.14313v4 Announce Type: replace-cross
Abstract: Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement learning with sparse outcome rewards or sea...
By Can Jin, Yang Zhou, Qixin Zhang, Hongwu Peng, Di Zhang, Zihan Dong, Marco Pavone, Ligong Han, Zhang-Wei Hong, Tong Che, Dimitris N. Metaxas
The paper introduces Test‑Time Policy Optimization (TTPO), a method that enables large language models to improve mathematical reasoning during inference without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels to guide an asymmetric objective: agreeing rollouts are distilled via On‑Policy Self‑Distillation, while disagreeing rollouts are penalized with Grouped Reinforcement Learning, with token‑level selection refining both branches. Experiments show that TTPO matches label‑supervised OPSD on five benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, demonstrating strong cross‑task generalization.
arXiv:2608. 03119v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability.
By Yongshi Ye, Liang Zhang, Yidong Chen, Xiaodong Shi, Biao Fu
arXiv:2606. 09635v1 Announce Type: cross Abstract: Ensuring the reliability of Large Language Models (LLMs) under distribution drift requires inference-time adaptation.
By Hankun Lin, Ruqi Zhang
arXiv:2604. 01476v2 Announce Type: replace Abstract: Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intended task.
By Rui Wu, Ruixiang Tang
The paper introduces a test‑time reinforcement learning framework for anomalous video understanding, addressing challenges such as unreliable pseudo‑labels, inadequate reward design, and collapsed group‑relative advantages. It proposes dual‑query consistency filtering, an entropy‑aware consensus reward, and a virtual negative anchor mechanism to improve sample reliability, reward quality, and policy‑gradient signals. Experiments on VAU‑Bench demonstrate significant performance gains, especially on the ECVA subset where accuracy rises from 75.81% to 90.00%.
By Huining Li, Yuxiang Duan, Jiyang Tan, Qian Li, MingCai Chen, Jian Zhang, Xingdong Sheng, Yuntao Du
arXiv:2606. 27369v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown.
By Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang, Kun Zhou, Tongtong Liang, Zhewei Yao, Yi-An Ma, Yuxiong He
The paper introduces label‑free bias‑only test‑time reinforcement learning (TTRL), which uses majority‑vote pseudo‑labels as rewards and optimizes only about 100 K bias parameters while keeping the pretrained backbone frozen. On the MATH‑500 benchmark it achieves 76.67 % accuracy, slightly better than a labeled bias‑steering baseline, and improves performance on several vision‑language and audio reasoning tasks. The authors also show that the learned steering vectors transfer to 4,500 held‑out MATH problems and analyze why such a highly restricted adaptation works, linking majority‑vote reliability to rollout consensus and gradient energy in bias subspaces.
By Naveen Vakada, Mingyuan Li, Shaoxiong Ji
arXiv:2602.13551v3 Announce Type: replace
Abstract: Reward models (RMs) play a central role throughout the language model (LM) pipeline, particularly in non-verifiable domains. However, the dominant...
By Yike Wang, Faeze Brahman, Shangbin Feng, Teng Xiao, Hannaneh Hajishirzi, Yulia Tsvetkov
arXiv:2606. 16236v1 Announce Type: new Abstract: Reinforcement learning (RL) often suffers from performance degradation when deployed in environments that differ from those encountered during training.
By Ekasit Usaratniwart, Xilin Gao, Marc Ong, Youhei Akimoto
The paper introduces Test‑Time Policy Optimization (TTPO), an approach that enables large language models to improve mathematical reasoning without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels and an asymmetric objective: it distills rollouts that agree with the pseudo‑label via On‑Policy Self‑Distillation and penalizes disagreeing rollouts with Grouped Reinforcement Learning. Token‑level selection further refines the process, down‑weighting already‑converged positions during distillation and penalizing only confident errors during RL. Experiments show that TTPO matches label‑supervised OPSD on five competition‑level benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, while also generalizing well across tasks.
By Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen