Compute Aligned Training: Optimizing for Test Time Inference
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces Test‑Time Policy Optimization (TTPO), a method that enables large language models to improve mathematical reasoning during inference without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels to guide an asymmetric objective: agreeing rollouts are distilled via On‑Policy Self‑Distillation, while disagreeing rollouts are penalized with Grouped Reinforcement Learning, with token‑level selection refining both branches. Experiments show that TTPO matches label‑supervised OPSD on five benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, demonstrating strong cross‑task generalization.
arXiv:2510. 10541v2 Announce Type: replace-cross Abstract: Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).
The paper introduces Test‑Time Policy Optimization (TTPO), an approach that enables large language models to improve mathematical reasoning without relying on ground‑truth labels. TTPO uses majority‑vote pseudo‑labels and an asymmetric objective: it distills rollouts that agree with the pseudo‑label via On‑Policy Self‑Distillation and penalizes disagreeing rollouts with Grouped Reinforcement Learning. Token‑level selection further refines the process, down‑weighting already‑converged positions during distillation and penalizing only confident errors during RL. Experiments show that TTPO matches label‑supervised OPSD on five competition‑level benchmarks, boosts Qwen3‑1.7B from 38.0 % to 45.2 % in test‑time training, and achieves significant gains without explicit reasoning steps, while also generalizing well across tasks.
arXiv:2508.14313v4 Announce Type: replace-cross Abstract: Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement learning with sparse outcome rewards or sea...
Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training data. To resolve this, Importance sampling (IS) is proposed, while the token-level ratios compound over long sequences, causing severe variance exploded.
arXiv:2606. 03608v1 Announce Type: cross Abstract: Test-time reinforcement learning has emerged as a promising paradigm for enhancing the complex reasoning abilities of large language models in a completely label-free manner.