arXiv Computation and Language By Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen

TTPO: Test-Time Policy Optimization

Read the original on arXiv Computation and Language →

The paper introduces Test‑Time Policy Optimization (TTPO), an approach that enables large language models to improve mathematical reasoning without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels and an asymmetric objective: it distills rollouts that agree with the pseudo‑label via On‑Policy Self‑Distillation and penalizes disagreeing rollouts with Grouped Reinforcement Learning. Token‑level selection further refines the process, down‑weighting already‑converged positions during distillation and penalizing only confident errors during RL. Experiments show that TTPO matches label‑supervised OPSD on five competition‑level benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, while also generalizing well across tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Aug 27

TTPO: Test-Time Policy Optimization

The paper introduces Test‑Time Policy Optimization (TTPO), a method that enables large language models to improve mathematical reasoning during inference without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels to guide an asymmetric objective: agreeing rollouts are distilled via On‑Policy Self‑Distillation, while disagreeing rollouts are penalized with Grouped Reinforcement Learning, with token‑level selection refining both branches. Experiments show that TTPO matches label‑supervised OPSD on five benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, demonstrating strong cross‑task generalization.

arXiv Computation and Language
Aug 27

Tool Verification for Test-Time Reinforcement Learning

The paper introduces T$^3$RL, a tool‑verification framework for test‑time reinforcement learning (TTRL) that mitigates the false‑popular failure mode by using external tool evidence to upweight verified rollouts during voting. By grounding pseudo‑label construction in verified evidence, T$^3$RL produces more reliable pseudo‑labels and improves performance over standard TTRL on math benchmarks such as MATH‑500, AMC, and AIME 2024. The approach positions T$^3$RL as a verified online data synthesizer, highlighting the importance of tool verification for reliable online adaptation and demonstrating extensibility to other verifiable domains.

By Ruotong Liao, Nikolai R\"ohrich, Xiaohan Wang, Yuhui Zhang, Yasaman Samadzadeh, Volker Tresp, Serena Yeung-Levy
arXiv Computation and Language
Sep 16

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

arXiv:2609.16660v1 Announce Type: new Abstract: Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathema...

By Kailong Fan, Anqi Pu, Yichen Wu, Wanhua Li, Yicong Li, Hanspeter Pfister, Huafeng Liu, Xiang Li, Quanzheng Li, Ning Guo
arXiv AI
1d ago

SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

SIPO (Self‑Instructing Policy Optimization) unifies reinforcement learning with on‑policy self‑distillation by using a contrastive self‑teacher to generate token‑level credit signals. The method samples multiple rollouts per prompt, pairs each with a reference answer and its mistakes, and uses the difference in teacher log‑probabilities to provide dense feedback while still respecting the overall task reward. Experiments on reasoning and code‑generation benchmarks show that SIPO outperforms both RLVR and OPSD baselines without requiring an external teacher or extra generation steps.

By Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang