arXiv Computation and Language

From Rollouts to Recipes: Self-Contained Post-Training for LLMs

The paper introduces Self‑Routing, a post‑training framework that tailors optimization for each sample based on its rollout correctness and confidence. Instead of applying a single recipe to all data, samples are routed to different strategies—GRPO, on‑policy self‑distillation, regularization, or skipped—allowing training to adapt without external teachers or extra annotations. Experiments on Qwen3 and Qwen3.5 show consistent improvements over uniform methods and reveal that the routing distribution evolves during training, reducing unnecessary updates on low‑signal or already stable samples.

arXiv AI
Sep 25

CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment

CounterRoute is an online reinforcement‑learning framework that jointly learns how to route a language model’s reasoning between a ‘think’ and a ‘direct answer’ mode, using a single shared policy derived from a dual‑mode checkpoint. It employs counterfactual rollouts to credit routing decisions and a curriculum that starts with forced dual‑mode rollouts before shifting to self‑routed updates, achieving better accuracy‑efficiency trade‑offs across nine benchmarks. The method reduces generated tokens by up to 51% on Qwen3‑8B while improving macro‑average accuracy, and its routing strategy generalizes to unseen coding, science, and commonsense tasks.

By Ruochen Jiao, Besnik Fetahu, Zhenyu Shi, Priyanka Nigam
arXiv AI
Aug 6

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy

arXiv:2608. 02087v2 Announce Type: replace Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration.

By Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
arXiv Computation and Language
Aug 28

TTPO: Test-Time Policy Optimization

The paper introduces Test‑Time Policy Optimization (TTPO), an approach that enables large language models to improve mathematical reasoning without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels and an asymmetric objective: it distills rollouts that agree with the pseudo‑label via On‑Policy Self‑Distillation and penalizes disagreeing rollouts with Grouped Reinforcement Learning. Token‑level selection further refines the process, down‑weighting already‑converged positions during distillation and penalizing only confident errors during RL. Experiments show that TTPO matches label‑supervised OPSD on five competition‑level benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, while also generalizing well across tasks.

By Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen