← Back to all news
arXiv Statistics ML October 2, 2026 By Shashank Gupta, Pilhwa Lee

Meta-reinforcement learning with minimum attention

Read the original on arXiv Statistics ML →

The Flow has not summarised this story yet — read it at arXiv Statistics ML.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 22

From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

arXiv:2607. 18830v1 Announce Type: cross Abstract: Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks.

By Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi
ragreinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jun 2

Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks

arXiv:2605. 22305v2 Announce Type: replace Abstract: We analytically solve the Mountain Car problem, a canonical benchmark in RL, and derive an optimal control solution, closing a gap after 36 years.

By Stefan Huber, Hannes Unger, Georg Sch\"afer, Jakob Rehrl
agentsreinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jul 16

Pretraining in Actor-Critic Reinforcement Learning for Locomotion

arXiv:2510. 12363v4 Announce Type: replace-cross Abstract: The pretraining-finetuning paradigm has facilitated numerous transformative advancements in artificial intelligence research in recent years.

By Jiale Fan, Andrei Cramariuc, Tifanny Portela, Marco Hutter
reinforcement-learningroboticsfine-tuningbenchmarks
More like this →
arXiv Machine Learning
Sep 21

REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement

arXiv:2609.21108v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies...

By Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs
agentsreinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jul 30

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

arXiv:2607. 26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models.

By Yunpeng Chu
llmsreinforcement-learningbenchmarkssafety
More like this →
arXiv AI
Aug 7

Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination

arXiv:2605. 04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning.

By Jonathan Spieler, Sven Behnke
reinforcement-learningbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea