arXiv Machine Learning

Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning

The paper introduces Multi-Objective In-context Knowledge Editing (MO‑IKE), a reinforcement learning framework that treats prompt construction for knowledge editing as a constrained Markov decision process. MO‑IKE jointly optimizes three competing objectives—reliability, generality, and specificity—by training a dynamic retriever to balance these goals and produce globally coherent prompts. Experiments on Llama‑3.2 show that MO‑IKE raises edit success from 85.0 % to 92.0 %, improves paraphrase consistency from 77 % to 79 %, and boosts retention rate by 23 % compared to earlier RL‑based methods.

arXiv Machine Learning
1d ago

SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization

SkillSpec is a two‑phase framework for evolving natural‑language skills in large language model agents. The first phase, consensus‑gated evolution, generates candidate skills from complementary editing intents and commits updates only when paired evaluations reach consensus on overall improvement and non‑negative aggregate gain. The second phase, representation specialization, uses signals from the optimization trajectory to choose an appropriate flat, graph, or hybrid structure for the skill, improving success rates by an average of 6.89% over SkillOpt across six benchmarks and three target language models.

By Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu
arXiv AI
Aug 25

DeepRefine: Agentic Knowledge Refinement via Reinforcement Learning

DeepRefine is a reinforcement learning framework that improves the quality of pre‑constructed structured knowledge bases—such as knowledge graphs or LLM‑Wikis—by engaging in multi‑turn interactions with the base. It performs abductive diagnosis to locate defects, then applies targeted refinement actions to incrementally update the knowledge base. The system uses a Gain‑Beyond‑Draft reward to train its refinement policy end‑to‑end, achieving consistent downstream performance gains over strong baselines.

By Haoyu Huang, Jiaxin Bai, Shujie Liu, Yang Wei, Huihao Jing, Hong Ting Tsang, Yisen Gao, Zhongwei Xie, Yufei Li, Yangqiu Song
arXiv Computation and Language
Aug 31

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot is a proactive context‑management framework designed to improve long‑horizon agentic reasoning with large language models. It expands the toolset to include planning, long‑term memory, and soft context offloading, and introduces a reinforcement‑learning strategy that focuses on critical editing decisions and assigns action‑level advantages. Experiments on long‑context QA and deep search tasks demonstrate that ContextPilot achieves stronger performance with a more compact working context, outperforming existing baselines across various base models and benchmarks.

By Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun
arXiv AI
Sep 15

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

The paper introduces an exploration-guided prompt scaffolding framework for multimodal large language models, dynamically adjusting the prompt distribution during reinforcement learning post-training. It uses an Exploration Potential Score (EPS) derived from KL-regularized policy improvement to assess prompt utility without extra overhead, and a teacher model rewrites low-utility prompts to preserve intent while improving informativeness. Experiments on Geo3K, MMK12, MathVision, and MMMU-Pro show consistent performance gains, up to 9.7% in-domain and over 11% on out-of-distribution benchmarks.

By Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang, Lei Shen, Jun Huang
arXiv AI
4d ago

Generalizable Lifelong Model Editing via Preference Optimization

The paper introduces GLIME, a lifelong model editing framework that integrates knowledge editing with preference optimization to handle continual updates in large language models. GLIME employs replay-based editing and a gradient constraint to prevent overfitting to target prompts and preserve previously edited knowledge. Experiments demonstrate that GLIME enhances knowledge generalization while maintaining editing performance and overall model capabilities.

By Dahyun Jung, Suhyune Son, Heuiseok Lim
arXiv AI
Aug 13

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

arXiv:2608. 11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world.

By Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang, Shiyang Wang, Zichen Miao, Linjun Zhang, Haoyu Wang, Jing Gao
arXiv AI
Jun 19

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.

By Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei