arXiv:2606. 03608v1 Announce Type: cross Abstract: Test-time reinforcement learning has emerged as a promising paradigm for enhancing the complex reasoning abilities of large language models in a completely label-free manner.
By Jiahui Li, Jianfeng Shan, Wenpei Chen, Shunyu Wu, Jian Lou, Wenjie Feng, Dan Li, See-Kiong Ng
arXiv:2609.37119v1 Announce Type: cross
Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
By Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry
arXiv:2604. 00830v3 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time.
By Zhanzhi Lou, Hui Chen, Yibo Li, Qian Wang, Bryan Hooi
The paper introduces Test‑Time Policy Optimization (TTPO), a method that enables large language models to improve mathematical reasoning during inference without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels to guide an asymmetric objective: agreeing rollouts are distilled via On‑Policy Self‑Distillation, while disagreeing rollouts are penalized with Grouped Reinforcement Learning, with token‑level selection refining both branches. Experiments show that TTPO matches label‑supervised OPSD on five benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, demonstrating strong cross‑task generalization.
TTSR (Test-Time Self-Reflection) is a framework that enables large language models to adapt during inference by alternating between a Student role that solves test questions and a Teacher role that analyzes failures and generates targeted variant questions. The method incorporates a weakness memory and a strategy note to guide exploration, reducing reliance on noisy pseudo-labels and inefficient rollouts. Experiments on mathematical reasoning benchmarks demonstrate consistent test-time improvements, strong cross-backbone generalization, and transfer to general-domain reasoning tasks.
By Haoyang He, Zihua Rong, Yunjia Zhao, Lan Yang, Jian Chang, Honggang Zhang
arXiv:2508. 10123v3 Announce Type: replace-cross Abstract: Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT).
By Maxime Heuillet, Yufei Cui, Boxing Chen, Audrey Durand, Prasanna Parthasarathi