arXiv AI By Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi

From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

Read the original on arXiv AI →

arXiv:2607. 18830v1 Announce Type: cross Abstract: Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 19

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

arXiv:2606. 20002v1 Announce Type: cross Abstract: This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously exploring the environment, learning from its own experiences, and iteratively self-updating its context about the environment, thereby achieving progressively better performance on future tasks conditioned on the updated context.

By Yanxi Chen, Weijie Shi, Yuexiang Xie, Boyi Hu, Yaliang Li, Bolin Ding, Jingren Zhou
arXiv Machine Learning
Sep 2

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.

By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz
arXiv Machine Learning
2d ago

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

The paper introduces COPC, a Coupled Off-Policy Correction method for asynchronous reinforcement learning of large language models. COPC coordinates policy-side and advantage-side corrections by combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals, addressing both policy mismatch and advantage staleness. Experiments show COPC outperforms existing asynchronous baselines on tool-integrated mathematical reasoning and search tasks, while maintaining training stability and minimal overhead.

By Zicheng Hu, Zhijian Zhou, Xuan Zhang, Yuchen Liu, Cheng Chen, Yuan Li, Qi Gu, Yan Feng, Hongyan Hao, Chao Qu