The paper introduces OSOL, a method for mitigating higher‑order interference in multi‑domain reinforcement learning. OSOL selects a focus domain each iteration, uses token‑level footprints from the previous checkpoint to rank rebound risk, and applies an adaptively scaled correction to the GRPO update. Experiments on Qwen3‑30B‑A3B show a 5.7% improvement over the best baseline without higher‑order differentiation.
By Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Guojun Yin, Wei Lin, Ran He
arXiv:2608.20873v1 Announce Type: new
Abstract: Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint,...
By Zifeng Liu, Zhiyong Du, Yaxin Lu, Yiming Mao, Zhenhe Wang, Wenqi Shi, Zhengkun Jing
arXiv:2605. 28860v2 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities.
By Jeanmely Rojas Nunez, Viraj Sawant, Nathan Allen, Nomgondalai Amgalanbaatar, Yannis Zongo, Vasu Sharma, Maheep Chaudhary
arXiv:2607. 01232v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers.
By Zijian Zhang, Rizhen Hu, Athanasios Glentis, Dawei Li, Chung-Yiu Yau, Hongzhou Lin, Mingyi Hong
arXiv:2601. 18699v2 Announce Type: replace Abstract: Sequential fine-tuning of Large Language Models (LLMs) adaptation to target tasks often triggers catastrophic forgetting, where the acquisition of novel target skills degrades ancestral capabilities.
By Gustav Olaf Yunus Laitinen-Fredriksson Lundstrom-Imanov
arXiv:2606. 08446v1 Announce Type: cross Abstract: Despite being powerful, reinforcement learning with verifiable rewards (RLVR) induces extremely long COT, making it computationally expensive.
By Yang Zhou, Ranajoy Sadhukhan, Zhaofeng Sun, Zhuoming Chen, Souvik Kundu, Saket Dingliwal, Sai Muralidhar Jayanthi, Aram Galstyan, Haizhong Zheng, Beidi Chen
arXiv:2606. 15455v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a key approach for enhancing the reasoning abilities of large language models.
By Suqin Yuan, Jinkun Chen, Jiyang Zheng, Muyang Li, Lei Feng, Dadong Wang, Tao Xiang, Tongliang Liu, Bo An
arXiv:2512. 02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations arise over time.
By Naveen George, Naoki Murata, Yuhta Takida, Konda Reddy Mopuri, Yuki Mitsufuji
arXiv:2609.36813v1 Announce Type: new
Abstract: Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. How...
By Chuanpu Liu, Miao Yu, Yikai Cai, Yuanhe Zhang, Zhenhong Zhou, Li Sun, Zuming Jiang, Yufei Guo
The paper investigates how to apply Reinforcement Learning with Verifiable Rewards (RLVR) to large language models across multiple domains. It compares two training paradigms—mixed multi-task RLVR and separate RLVR followed by model merging—using tasks such as math, coding, science, instruction following, and agent. Experiments show that RLVR across domains causes minimal interference and that reasoning-intensive domains can synergize, with insights drawn from information constraints, prediction behavior, and self-verification.
By Haoqing Wang, Xiang Long, Ziheng Li, Yilong Xu, Tingguang Li, Yehui Tang
arXiv:2609.33149v2 Announce Type: replace
Abstract: A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how...
By Hongbo Chen, Guohua Lu, Ting Dang, Hong Jia
arXiv:2608.21811v1 Announce Type: new
Abstract: Reinforcement learning for vision-language math reasoning starves under sparse reward: on a pool of 20,830 visual-math problems where Qwen2-VL-2B answe...
By Qiqian Fu