Hugging Face Trending Papers

DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

Read the original on Hugging Face Trending Papers →

DeliveryGym is a 3D reinforcement learning environment that simulates continuous courier shifts, integrating multimodal tool interaction and persistent world dynamics to compute trajectory rewards based on simulator events. It provides feedback on resource consumption—time, energy, and money—across entire delivery trajectories, enabling agents to learn planning that balances immediate task success with long‑term constraints. The environment also adapts future training shifts to a policy’s weaknesses while keeping evaluation fixed, demonstrating that both learning from complete shifts and targeted training improve agent performance on complex delivery tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 18

DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

DeliveryGym is a 3D reinforcement learning environment that simulates continuous courier shifts, integrating multimodal tool interaction, persistent world dynamics, and trajectory‑based rewards derived from simulator events. It allows agents to learn how their decisions affect time, energy, and money across an entire shift, and it adapts future training shifts to the policy’s weaknesses while keeping evaluation fixed. Experiments on six models and 13 city maps show a significant gap between task execution and optimal sequencing, with RL improving Qwen3‑VL‑4B’s net income by 54.3% and adaptive training boosting test income by 16.5% over uniform sampling.

By Haoqiang Kang, Yiming Zhang, Yiyang Guo, Chuying Li, Jianzhi Shen, Tianruo Rose Xu, Xiaokang Ye, Lianhui Qin
arXiv AI
Jul 15

DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents

arXiv:2509. 21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generation, ensuring an enjoyable user experience.

By Yansong Ning, Rui Liu, Jun Wang, Kai Chen, Wei Li, Jun Fang, Kan Zheng, Naiqiang Tan, Hao Liu
arXiv AI
Sep 24

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

The paper introduces VHD-Play, a pipeline that first samples and solves a mathematical model before generating agentic reinforcement learning environments, ensuring that dynamics and evaluation are aligned from the outset. This approach yields 3,300 diverse environments at a low cost and significantly improves the performance of a large language‑model agent (Qwen3.6‑35B‑A3B) across multiple diagnostic families and external benchmarks. The study demonstrates that stateful interaction is a key factor in learning gains and that scaling the training substrate can further enhance performance.

By Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
arXiv AI
Aug 28

Decoupling Planning and Control for Instructable Agents

The paper introduces Instruct-to-Act, a system that decouples planning and control by combining a vision‑language model (VLM) planner with a world‑model controller. The VLM generates sparse, high‑level text instructions, while the controller executes them at high frequency, trained via relabeling rollouts with synthetic instructions and joint optimization of behavior cloning, reward, and world‑model objectives. Across seven embodied environments—including multi‑agent settings—this approach outperforms controller‑only and direct VLM action methods, maintains fast control, and allows swapping pretrained VLM planners without fine‑tuning, achieving competitive results with strong baselines on most tasks.

By Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr