arXiv AI By Chaofan Pan, Lingfei Ren, Xiangyu Jiang, Yanhua Li, Xuemei Cao, Xiangkun Wang, Hao Yu, Wei Wei, Xin Yang

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

Read the original on arXiv AI →

arXiv:2607. 21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected data must be removed after training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

arXiv:2607. 09773v1 Announce Type: new Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments.

By Mianqiu Huang, Taofeng Xue, Chong Peng, Jinrui Ding, Sicheng Fan, Jiale Hong, Yufei Gao, Xiaocheng Zhang, Linsen Guo, Xin Yang, Dengchang Zhao, Yuchen Xie, Peng Pei, Xunliang Xie, Xipeng Qiu
arXiv AI
Sep 24

DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents

The paper introduces Direct Relational Set‑Risk Pruning (DRSR), a method for compressing the history of long‑horizon language‑model agents by selecting deletion sets based on risk constraints rather than independent unit scores. DRSR builds counterfactual supervision offline, then uses a lightweight scorer to predict set‑level harm during deployment, removing the largest safe set while respecting recency, protocol, and budget limits. Experiments on WorkBuddyBench Full260 and Eval40 show that DRSR improves mean reward from 0.699 to 0.802 and reduces token usage by over 20%, with further analyses highlighting the importance of decision‑conditioned relations, retained context, pair interactions, and abstention.

By Mingxuan Wang, Bo Wang, Fei Luo, Guorun Yao, Chao Ning, Yinglong Guo, Hongyue Chen, Yanbiao Ma, Jungong Han