arXiv Computation and Language By Jianshu Zhang, Keliang Wu, Chengxuan Qian, Xiyuan Yang, Ce Zhang, Ariel Tian, Anbang Liu, Haoran Lu, Han Liu

ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

Read the original on arXiv Computation and Language →

The paper introduces ProgressCompass, a framework that enhances Embodied Progress Reward Models (PRMs) by providing the necessary contextual information for accurate progress estimation in long manipulation tasks. It presents ContextProgress-Bench, a benchmark with 24 tasks that tests PRMs under three context-dependent scenarios—State Recall, Sequence Tracking, and Recurrence Disambiguation—showing that even history-aware PRMs struggle without proper context. By integrating a context-aware loop that leverages general-purpose vision‑language models, ProgressCompass reduces PRM progress error by up to 82% and improves rank agreement by 76%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 22

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

The paper introduces Trace, a framework that transforms sparse-reward trajectories into executable walkthroughs by identifying progress anchors, propagating credit, and estimating action prerequisites. Trace compiles noisy trajectories into state‑conditioned, verifiable procedures that remove loops and detours, enabling reuse, intermediate‑state resumption, and programmatic verification. Experiments on J‑TTL, WebShop, and ScienceWorld with three open‑source LLMs show that Trace outperforms eight baselines, improving average AUC and Final‑$3$ by 30.0% and 40.5% while using fewer inference tokens.

By Kaijie Chen, Chenyu Fang, Liang Yan, Bo Li, Bo Zhang, Peng Ye