arXiv Machine Learning By Xuzhong Wang, Maiqi Jiang, Tejal Nair, Girija Bhusal, Yanfu Zhang, Haipeng Chen

Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper introduces Multi-Objective In-context Knowledge Editing (MO‑IKE), a reinforcement learning framework that treats prompt construction for knowledge editing as a constrained Markov decision process. MO‑IKE jointly optimizes three competing objectives—reliability, generality, and specificity—by training a dynamic retriever to balance these goals and produce globally coherent prompts. Experiments on Llama‑3.2 show that MO‑IKE raises edit success from 85.0 % to 92.0 %, improves paraphrase consistency from 77 % to 79 %, and boosts retention rate by 23 % compared to earlier RL‑based methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization

SkillSpec is a two‑phase framework for evolving natural‑language skills in large language model agents. The first phase, consensus‑gated evolution, generates candidate skills from complementary editing intents and commits updates only when paired evaluations reach consensus on overall improvement and non‑negative aggregate gain. The second phase, representation specialization, uses signals from the optimization trajectory to choose an appropriate flat, graph, or hybrid structure for the skill, improving success rates by an average of 6.89% over SkillOpt across six benchmarks and three target language models.

By Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu
arXiv AI
Aug 25

DeepRefine: Agentic Knowledge Refinement via Reinforcement Learning

DeepRefine is a reinforcement learning framework that improves the quality of pre‑constructed structured knowledge bases—such as knowledge graphs or LLM‑Wikis—by engaging in multi‑turn interactions with the base. It performs abductive diagnosis to locate defects, then applies targeted refinement actions to incrementally update the knowledge base. The system uses a Gain‑Beyond‑Draft reward to train its refinement policy end‑to‑end, achieving consistent downstream performance gains over strong baselines.

By Haoyu Huang, Jiaxin Bai, Shujie Liu, Yang Wei, Huihao Jing, Hong Ting Tsang, Yisen Gao, Zhongwei Xie, Yufei Li, Yangqiu Song
arXiv Computation and Language
Aug 31

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot is a proactive context‑management framework designed to improve long‑horizon agentic reasoning with large language models. It expands the toolset to include planning, long‑term memory, and soft context offloading, and introduces a reinforcement‑learning strategy that focuses on critical editing decisions and assigns action‑level advantages. Experiments on long‑context QA and deep search tasks demonstrate that ContextPilot achieves stronger performance with a more compact working context, outperforming existing baselines across various base models and benchmarks.

By Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun
arXiv AI
Sep 15

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

The paper introduces an exploration-guided prompt scaffolding framework for multimodal large language models, dynamically adjusting the prompt distribution during reinforcement learning post-training. It uses an Exploration Potential Score (EPS) derived from KL-regularized policy improvement to assess prompt utility without extra overhead, and a teacher model rewrites low-utility prompts to preserve intent while improving informativeness. Experiments on Geo3K, MMK12, MathVision, and MMMU-Pro show consistent performance gains, up to 9.7% in-domain and over 11% on out-of-distribution benchmarks.

By Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang, Lei Shen, Jun Huang