arXiv AI

In-Context Reinforcement Learning via Communicative World Models

arXiv:2508. 06659v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) agents often struggle to generalize to new tasks and contexts without updating their parameters, mainly because their learned representations and policies are overfit to the specifics of their training environments.

arXiv AI
Jun 4

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring

arXiv:2604. 12645v2 Announce Type: replace-cross Abstract: Although autonomous underwater vehicles promise the capability of marine ecosystem monitoring, their deployment is fundamentally limited by the difficulty of controlling vehicles under highly uncertain and non-stationary underwater dynamics.

By Melvin Laux, Yi-Ling Liu, Rina Alo, S\"oren T\"opper, Mariela De Lucas Alvarez, Frank Kirchner, Rebecca Adam
arXiv AI
Jul 15

In-Context Reinforcement Learning under Non-Stationarity: A Survey

arXiv:2607. 11906v1 Announce Type: new Abstract: The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest in in-context reinforcement learning (ICRL): the ability of a pretrained or fine-tuned decision model to infer latent task rules and improve future behavior from interaction context, without test-time parameter updates.

By A Run, Ziluo Ding
arXiv Computation and Language
Sep 15

Learning to Coach for Experiential Learning

arXiv:2609.15851v1 Announce Type: new Abstract: Language models can learn from experience, but raw solution trajectories are often too long and noisy to provide effective guidance. In this work, we p...

By Guanheng Chen, Tianzhu Ye, Li Dong, Xun Wu, Shaohan Huang, Furu Wei
arXiv Machine Learning
Sep 14

Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

The paper introduces CLAW, a method that uses a hypernetwork to generate low‑rank adapters for world models during test time, enabling efficient adaptation to new environments with only a few episodes of interaction. By jointly pretraining the hypernetwork and base model on simulated adaptations, CLAW balances computational efficiency and expressivity, outperforming both in‑context learning and gradient‑based adaptation in locomotion and manipulation tasks. The approach also mitigates overfitting in data‑scarce regimes and demonstrates that the benefit stems from expressive adapters rather than context conditioning.

By Fernando Palafox, David Fridovich-Keil
Hugging Face Trending Papers
Aug 3

Quo Vadis, World Modeling?

Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions.

arXiv AI
Aug 5

Quo Vadis, World Modeling?

arXiv:2608. 02713v1 Announce Type: cross Abstract: Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize.

By Yu Yang, Xuemeng Yang, Licheng Wen, Lingdong Kong, Xiaobin Hu, Dongyue Lu, Wei Chow, Xiyan Huang, Yuxiang Feng, Yue Liao, Jianbiao Mei, Daocheng Fu, Rong Wu, Pinlong Cai, Ran Yi, Ying Tai, Jiangning Zhang, Botian Shi, Yong Liu, Shuicheng Yan
Hugging Face Trending Papers
Jun 18

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously exploring the environment, learning from its own experiences, and iteratively self-updating its context about the environment, thereby achieving progressively better performance on future tasks conditioned on the updated context. Major components of the CoD framework include: (1) algorithm design and infrastructure for end-to-end reinforcement learning (RL) with long rollout sequences interleaving solve-task and update-context episodes; (2) tasks and environments for incentivizing and eliciting the targeted meta-capability in LLMs during training, as well as for faithfully measuring progress during evaluation.