Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
The paper introduces ProgressCompass, a framework that enhances Embodied Progress Reward Models (PRMs) by providing the necessary contextual information for accurate progress estimation in long manipulation tasks. It presents ContextProgress-Bench, a benchmark with 24 tasks that tests PRMs under three context-dependent scenarios—State Recall, Sequence Tracking, and Recurrence Disambiguation—showing that even history-aware PRMs struggle without proper context. By integrating a context-aware loop that leverages general-purpose vision‑language models, ProgressCompass reduces PRM progress error by up to 82% and improves rank agreement by 76%.
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
The paper introduces Trace, a framework that transforms sparse-reward trajectories into executable walkthroughs by identifying progress anchors, propagating credit, and estimating action prerequisites. Trace compiles noisy trajectories into state‑conditioned, verifiable procedures that remove loops and detours, enabling reuse, intermediate‑state resumption, and programmatic verification. Experiments on J‑TTL, WebShop, and ScienceWorld with three open‑source LLMs show that Trace outperforms eight baselines, improving average AUC and Final‑$3$ by 30.0% and 40.5% while using fewer inference tokens.
arXiv:2606. 28529v1 Announce Type: cross Abstract: Embodied foundation models have recently been widely used to improve robot generalization and task success rates.
arXiv:2608.29537v1 Announce Type: cross Abstract: Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so t...
arXiv:2610.00604v1 Announce Type: cross Abstract: Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappe...
arXiv:2610.00982v1 Announce Type: cross Abstract: Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the acti...
EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.
arXiv:2606. 04970v1 Announce Type: cross Abstract: We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \textit{how} to coach.
We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \textit{how} to coach. However, progress is limited by the absence of large-scale, cross-domain benchmarks that reflect realistic conditions, particularly the common case in which users deviate from the expected step sequence.
arXiv:2605. 17877v2 Announce Type: replace Abstract: A significant hurdle for current LLMs is the execution of complex, multi-stage tasks.
arXiv:2603. 04910v2 Announce Type: replace-cross Abstract: Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condition on single-step observations or short-context histories, making them struggle with non-Markovian tasks that require long-term memory.
arXiv:2608.30378v1 Announce Type: cross Abstract: Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: th...