arXiv:2605.13054v2 Announce Type: replace-cross
Abstract: Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. Whe...
By Minung Kim, Jeongmo Kim, Gwanwoo Choi, Seungyul Han
arXiv:2607. 11720v1 Announce Type: cross Abstract: Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction.
By Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous.
arXiv:2607. 24720v1 Announce Type: cross Abstract: Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear.
By Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao
The paper introduces LP‑BTS, a learning‑guided planning framework for mobile charging in large, dynamic action spaces. It uses a graph proposal policy to narrow candidate stops, a value critic to evaluate leaf nodes, and edge‑budgeted PUCT to compare short simulated futures before action selection. Experiments on a 30‑scenario battery‑life benchmark show LP‑BTS achieving the highest survival and alive‑AUC, outperforming domain‑engineered baselines and heuristic policies.
By Liang-Ching Tao, Pi-Chung Wang
arXiv:2608. 02305v1 Announce Type: new Abstract: Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget.
By Jiaorong Feng, Qian Li, Ying Li