InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 07746v1 Announce Type: new Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making.
arXiv:2603. 13707v3 Announce Type: replace-cross Abstract: Humanoid loco-manipulation requires coordinated task-space motion planning with stable loco-manipulation command tracking under complex robot-environment dynamics and long-horizon tasks.
SynthDemo‑RL introduces a teacher‑student framework that uses an automated teacher to generate successful manipulation trajectories from simulator‑privileged state, which are then distilled into a Vision‑Language‑Action (VLA) student via supervised fine‑tuning. The student is further refined with PPO using binary task‑success rewards. On the LIBERO‑PRO benchmark, SynthDemo‑RL rescues all 27 previously unsolvable tasks and achieves near‑perfect success rates, matching performance that would otherwise require human demonstrations.
arXiv:2606. 11891v1 Announce Type: cross Abstract: Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy.
arXiv:2608.31167v1 Announce Type: cross Abstract: Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objecti...
arXiv:2607. 00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures.