Hugging Face Trending Papers

Finding the Time to Think: Learning Planning Budgets in Real-Time RL

Deliberating takes time. In real-time settings, that time is not free.

arXiv AI
Jul 23

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.

By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu
Hugging Face Trending Papers
Aug 19

Reinforced Planning with Latent World Models

Reinforced Planning with Latent World Models introduces RP1, a neural planner that learns to evaluate imagined outcomes via a critic and improve multi‑step plans through an optimizer trained offline on world‑model roll‑outs. Unlike existing planners that are hand‑designed or only inform policies, RP1 fully learns to refine plans and can be attached to any pretrained latent world model. In experiments on visual navigation, arm reaching, and robotic manipulation, RP1 outperforms hand‑designed search algorithms, achieving near‑perfect success while using 1,000× fewer roll‑outs and up to 67× faster inference.

arXiv Machine Learning
Aug 20

Reinforced Planning with Latent World Models

Reinforced Planning with Latent World Models (RP1) is a novel method that learns to evaluate imagined outcomes via a critic and to improve multi‑step plans through an optimizer trained offline on world‑model roll‑outs. It is the first approach to fully learn plan improvement and can be attached to any pretrained latent world model. In experiments on visual navigation, arm reaching, and robotic manipulation, RP1 outperforms hand‑designed search algorithms, achieving near‑perfect success while using far fewer roll‑outs and running up to 67× faster than the strongest alternative.

By Armin Sommer, Jannik Schilling
arXiv Machine Learning
Aug 26

Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

The paper introduces ARLI, a latency‑aware framework that enables reinforcement learning fine‑tuning of large generalist robot policies despite inference delays. ARLI combines asynchronous inference with state augmentations—incorporating committed actions and mid‑inference observations—to restore near‑Markovian dynamics and maintain reactivity. Experiments on simulated and real‑world manipulation tasks show that ARLI allows effective policy improvement under latency, outperforming standard RL even in no‑latency scenarios.

By Brian Zhu (Siemens), Momen Khalil (Siemens), E Harrison (UC Berkeley), Emanuele Poggi (Siemens), Philipp Schmitt (Siemens), Bernd Kast (Siemens), Philine Meister (Siemens), Pranav Atreya (UC Berkeley), Qiyang Li (UC Berkeley), Finn Ferchau (Siemens), Cesar Colmenero (Siemens), Yash Shahapurkar (Siemens), Gokul Narayanan (Siemens), Melih Erdogan (Siemens), Kai Wurm (Siemens), Georg von Wichert (Siemens), Oier Mees (Microsoft, ETH Zurich, UC Berkeley), Eugen Solowjow (Siemens), Andrew Wagenmaker (UC Berkeley), Sergey Levine (UC Berkeley)
arXiv AI
Aug 14

Exploiting Symbolic Heuristics for the Synthesis of Domain-Specific Temporal Planning Guidance using Reinforcement Learning

arXiv:2505. 13372v2 Announce Type: replace Abstract: Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of temporal planners when a domain is fixed and a set of training problems (not plans) is given.

By Irene Brugnara, Alessandro Valentini, Andrea Micheli
arXiv Machine Learning
Sep 21

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

GEM-MPC is a reinforcement learning method that blends MPPI planning with policy learning to balance exploration and exploitation in high-dimensional continuous control tasks. It trains a policy to clone the planner while also maintaining a KL-regularized policy that explores around the planner’s suggestions, thereby improving the synergy between planning and learning. The approach introduces Gated Prior Distillation, which selectively updates policies from stored planning distributions only when they offer better targets, reducing the influence of stale data without costly reanalysis. Across continuous-control benchmarks, GEM-MPC outperforms existing planning-based baselines while using lower computational budgets.

By Alvaro Serra-Gomez, Thomas Moerland