arXiv AI

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

arXiv:2607. 21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected data must be removed after training.

arXiv AI
Jul 14

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

arXiv:2607. 09773v1 Announce Type: new Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments.

By Mianqiu Huang, Taofeng Xue, Chong Peng, Jinrui Ding, Sicheng Fan, Jiale Hong, Yufei Gao, Xiaocheng Zhang, Linsen Guo, Xin Yang, Dengchang Zhao, Yuchen Xie, Peng Pei, Xunliang Xie, Xipeng Qiu
arXiv AI
Sep 24

DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents

The paper introduces Direct Relational Set‑Risk Pruning (DRSR), a method for compressing the history of long‑horizon language‑model agents by selecting deletion sets based on risk constraints rather than independent unit scores. DRSR builds counterfactual supervision offline, then uses a lightweight scorer to predict set‑level harm during deployment, removing the largest safe set while respecting recency, protocol, and budget limits. Experiments on WorkBuddyBench Full260 and Eval40 show that DRSR improves mean reward from 0.699 to 0.802 and reduces token usage by over 20%, with further analyses highlighting the importance of decision‑conditioned relations, retained context, pair interactions, and abstention.

By Mingxuan Wang, Bo Wang, Fei Luo, Guorun Yao, Chao Ning, Yinglong Guo, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv AI
Jun 2

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

arXiv:2605. 24202v2 Announce Type: replace Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that are poorly understood.

By Yifan Zeng, Yiran Wu, Yaolun Zhang, Wentian Zhao, Kun Wan, Qingyun Wu, Huazheng Wang
arXiv AI
Jun 10

Deployment-Time Memorization in Foundation-Model Agents

arXiv:2606. 10062v1 Announce Type: new Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time function rather than solely a property of model weights.

By Lei (Rachel), Chen, Guilin Zhang, Kai Zhao, Dalmo Cirne, Andy Olsen, Xu Chu, Zeke Miller, Alet Blanken, Amine Anoun, Jerry Ting
arXiv Machine Learning
Sep 14

PEARL: Structural Privacy-Utility Control in Human-Centric CPS via Personalized Early-Exit Deep Reinforcement Learning

PEARL is a framework for human‑centric cyber‑physical systems that uses a dual‑path Early‑Exit Deep Q‑Network to control the trade‑off between privacy and utility. By training per‑branch binary labels—Utility Confidence Labels (UCL) and Privacy Confidence Labels (PCL)—based on mutual information between private states and observable actions, PEARL selects the shallowest exit that satisfies both privacy and utility constraints, avoiding noise injection. The system includes an MI‑based feedback loop to detect behavioral drift and trigger retraining, and experiments on a smart‑home HVAC system and a VR smart classroom show a 25.67% reduction in adversarial state‑inference accuracy with only a 10‑16% utility cost.

By Mojtaba Taherisadr, Salma Elmalaki