arXiv AI

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

The paper introduces label‑free bias‑only test‑time reinforcement learning (TTRL), which uses majority‑vote pseudo‑labels as rewards and optimizes only about 100 K bias parameters while keeping the pretrained backbone frozen. On the MATH‑500 benchmark it achieves 76.67 % accuracy, slightly better than a labeled bias‑steering baseline, and improves performance on several vision‑language and audio reasoning tasks. The authors also show that the learned steering vectors transfer to 4,500 held‑out MATH problems and analyze why such a highly restricted adaptation works, linking majority‑vote reliability to rollout consensus and gradient energy in bias subspaces.

arXiv AI
Jun 3

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification

arXiv:2606. 03608v1 Announce Type: cross Abstract: Test-time reinforcement learning has emerged as a promising paradigm for enhancing the complex reasoning abilities of large language models in a completely label-free manner.

By Jiahui Li, Jianfeng Shan, Wenpei Chen, Shunyu Wu, Jian Lou, Wenjie Feng, Dan Li, See-Kiong Ng
Hugging Face Trending Papers
Aug 27

TTPO: Test-Time Policy Optimization

The paper introduces Test‑Time Policy Optimization (TTPO), a method that enables large language models to improve mathematical reasoning during inference without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels to guide an asymmetric objective: agreeing rollouts are distilled via On‑Policy Self‑Distillation, while disagreeing rollouts are penalized with Grouped Reinforcement Learning, with token‑level selection refining both branches. Experiments show that TTPO matches label‑supervised OPSD on five benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, demonstrating strong cross‑task generalization.

arXiv Machine Learning
Aug 27

TTSR: Test-Time Self-Evolving via Reflection

TTSR (Test-Time Self-Reflection) is a framework that enables large language models to adapt during inference by alternating between a Student role that solves test questions and a Teacher role that analyzes failures and generates targeted variant questions. The method incorporates a weakness memory and a strategy note to guide exploration, reducing reliance on noisy pseudo-labels and inefficient rollouts. Experiments on mathematical reasoning benchmarks demonstrate consistent test-time improvements, strong cross-backbone generalization, and transfer to general-domain reasoning tasks.

By Haoyang He, Zihua Rong, Yunjia Zhao, Lan Yang, Jian Chang, Honggang Zhang
arXiv Computation and Language
Sep 16

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

arXiv:2609.16660v1 Announce Type: new Abstract: Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathema...

By Kailong Fan, Anqi Pu, Yichen Wu, Wanhua Li, Yicong Li, Hanspeter Pfister, Huafeng Liu, Xiang Li, Quanzheng Li, Ning Guo
Hugging Face Trending Papers
Jul 20

MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models

Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models ($\leq 4 \, \mathrm{B}$ parameters) trained under limited budgets. We introduce MADA-RL, a post-training framework that specializes compact models into generator and critic roles and trains them with a debate-aware learning signal, fine-tuning only a small subset of parameters via LoRA adapters.

arXiv Computation and Language
Aug 28

TTPO: Test-Time Policy Optimization

The paper introduces Test‑Time Policy Optimization (TTPO), an approach that enables large language models to improve mathematical reasoning without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels and an asymmetric objective: it distills rollouts that agree with the pseudo‑label via On‑Policy Self‑Distillation and penalizes disagreeing rollouts with Grouped Reinforcement Learning. Token‑level selection further refines the process, down‑weighting already‑converged positions during distillation and penalizing only confident errors during RL. Experiments show that TTPO matches label‑supervised OPSD on five competition‑level benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, while also generalizing well across tasks.

By Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen