arXiv AI By Ruochen Jiao, Besnik Fetahu, Zhenyu Shi, Priyanka Nigam

CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment

Read the original on arXiv AI →

CounterRoute is an online reinforcement‑learning framework that jointly learns how to route a language model’s reasoning between a ‘think’ and a ‘direct answer’ mode, using a single shared policy derived from a dual‑mode checkpoint. It employs counterfactual rollouts to credit routing decisions and a curriculum that starts with forced dual‑mode rollouts before shifting to self‑routed updates, achieving better accuracy‑efficiency trade‑offs across nine benchmarks. The method reduces generated tokens by up to 51% on Qwen3‑8B while improving macro‑average accuracy, and its routing strategy generalizes to unseen coding, science, and commonsense tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

arXiv:2606. 30345v1 Announce Type: cross Abstract: Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks.

By Haisen Luo, Yiwei Liu, Haoning Wang, Dan Liu, Junxi Yin, Haotian Wang, Lei Zhang, Xiaoyu Tian, Shuaiting Chen, Yuansheng Song, Baoyan Guo, Xiongfei Yan, Bolan Yang, Chengwei Liu, Ming Cui, Jiong Chen
arXiv AI
4d ago

SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

SIPO (Self‑Instructing Policy Optimization) unifies reinforcement learning with on‑policy self‑distillation by using a contrastive self‑teacher to generate token‑level credit signals. The method samples multiple rollouts per prompt, pairs each with a reference answer and its mistakes, and uses the difference in teacher log‑probabilities to provide dense feedback while still respecting the overall task reward. Experiments on reasoning and code‑generation benchmarks show that SIPO outperforms both RLVR and OPSD baselines without requiring an external teacher or extra generation steps.

By Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang