arXiv Computation and Language By Yipeng Wang, Ziwei Zhang, Jiahui Zhang, Qi Gan, Kai Sheng

KuaiRP Series Role-playing Models Technical Report

Read the original on arXiv Computation and Language →

The paper presents the KuaiRP series of role‑playing models, detailing a multi‑stage training pipeline that balances deep domain knowledge injection with the preservation of general agent capabilities. The approach includes a standardized character template, a supervised fine‑tuning (SFT) data pipeline, a rule‑based reward function for reinforcement learning, and a novel two‑stage on‑policy distillation (OPD) with cumulative‑divergence decay (CDD) to recover general skills. Experimental results show that the models achieve state‑of‑the‑art role‑playing fidelity in target domains while maintaining low deployment costs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 19

Reinforcement-aware Knowledge Distillation for LLM Reasoning

arXiv:2602. 22495v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such models motivates distillation into smaller students.

By Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo Yang, Zhuowen Tu, Wei Xia, Stefano Soatto
arXiv AI
Jul 28

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

arXiv:2607. 24280v1 Announce Type: new Abstract: Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision.

By Junlin Liu, Jiangwang Chen, Zixin Song, Shuaiyu Zhou, Chunji Lv, Hank Wu, Kailin Jiang, Jinyang Wu, Bohan Yu, Chenxi Zhou
arXiv Machine Learning
Aug 28

Learning Generalizable Behaviors for Terminal Agents

The paper introduces the Agentic Compositional Generalization hypothesis, suggesting that reinforcement learning (RL) primarily refines high‑level decision‑making behaviors that orchestrate pre‑trained low‑level skills, rather than teaching new domain‑specific skills from scratch. It proposes River, a training recipe that enhances reward quality by filtering low‑quality synthetic environments and adding process‑level behavior regularization. Using River, RL‑trained agents outperform other open‑source 8B models on four terminal‑agent benchmarks, achieving significant gains with fewer than 30% of the training environments.

By Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao, Shafiq Joty, Semih Yavuz