arXiv:2607. 11906v1 Announce Type: new Abstract: The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest in in-context reinforcement learning (ICRL): the ability of a pretrained or fine-tuned decision model to infer latent task rules and improve future behavior from interaction context, without test-time parameter updates.
By A Run, Ziluo Ding
The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.
By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv:2606. 24160v1 Announce Type: new Abstract: Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.
By Elias Bareinboim, Junzhe Zhang, Sanghack Lee
Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i. e.
arXiv:2606. 07890v1 Announce Type: new Abstract: Performative prediction studies feedback loops that arise when predictive models are deployed in consequential domains.
By Jaewook Lee, Tijana Zrnic
arXiv:2608.20909v1 Announce Type: new
Abstract: Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled l...
By Xuyao Lin, Yixiang Shan, Jinru Duan, Tao Yang, Xinyu Zhao, Runyu Lei, Yiming Zhao, Jiaxin Fan, Zongbao Feng, Peng Jia
arXiv:2609.23753v1 Announce Type: cross
Abstract: Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeli...
By Yikun Miao, Fangqi Zhu, Quanxin Shou, Xiaoyi Pang, Zhengyang Yan, Junhao Li, Haodong Wang, Zicong Hong, Song Guo
arXiv:2507. 10142v2 Announce Type: replace Abstract: Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the assumptions under which algorithms are designed and evaluated.
By Siyi Hu, Mohamad A Hady, Jianglin Qiao, Jimmy Cao, Mahardhika Pratama, Ryszard Kowalczyk
arXiv:2602. 23545v2 Announce Type: replace Abstract: In the real world, planning is often challenged by distribution shifts.
By Matteo Ceriscioli, Karthika Mohan
arXiv:2607. 16090v1 Announce Type: cross Abstract: Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains.
By Hanyang Chen, Anirudh Satheesh, Longchao Da, Hua Wei
The paper challenges the common assumption that the successor measure in reinforcement learning is approximately low-rank, showing instead that a low-rank structure emerges in a shifted successor measure that ignores initial transitions. It provides finite-sample guarantees for estimating this low-rank approximation, introduces Type II Poincaré inequalities to bound spectral recoverability, and links the necessary shift to the decay of high-order singular values and local mixing properties. Experiments confirm that shifting the successor measure improves goal-conditioned RL performance.
By Bastien Dubail, Stefan Stojanovic, Alexandre Prouti\`ere
The paper introduces a causal framework for concept drift, using Structural Causal Models to classify drift events by their causal origin—exogenous variables, endogenous mechanisms, confounders, and target-generating processes. It presents an SCM-based data stream generator that simulates controlled mechanism-level drift, and empirically shows that different causal origins produce distinct distribution shifts and predictive behaviors. By integrating causal discovery, the authors create realistic data streams that improve downstream performance and provide a foundation for causally-aware evaluation in non‑stationary settings.
By Eduardo V. L. Barboza, Jean Paul Barddal, Robert Sabourin, Rafael M. O. Cruz