arXiv AI By Xue Qin, Simin Luan, Cong Yang, Zhijun Li

Evolving Skill Modules under a Fixed Planner: Versioning, Rollback, and Runtime Governance for Long-Lived Robot Systems

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 18

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

The paper introduces HALTER, a graph-based system that automates the reset and evaluation of long-horizon robot manipulation tasks. HALTER constructs a spatial scene graph from point clouds and vision models, uses an LLM to score rollouts, plan resets, and verify success, all without labeled success images. In experiments on a Franka arm, HALTER restores scenes in 76% of episodes, improves skill completion estimation, and reduces operator time by 72% compared to manual reset.

By Jing Jiang, Yue Yang, Xinkai Jiang, Gedas Bertasius, Daniel J. Szafir, Rudolf Lioutikov
arXiv AI
3d ago

Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

The paper introduces URAI, a Universal Robot‑Agent Interface that separates robot control into two roles: a programming agent that writes reusable, task‑specific tools from intent, and an execution agent that calls these tools in a feedback loop. This design keeps high‑level decision making in the model while delegating low‑level motion to code, allowing tool revisions to persist across episodes without retraining the foundation model. Experiments on RoboDojo and AgileX tasks show significant gains in success rate, speed, and token efficiency compared to direct fingertip control and pre‑written programs.

By Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao
arXiv Machine Learning
Sep 25

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy

The paper investigates task collapse—a failure mode where online RL fine‑tuning of a pretrained flow‑matching vision‑language‑action policy erodes performance on individual tasks—using a 450M‑parameter SmolVLA policy on LIBERO‑10. Three exploration‑noise strategies are compared: a fixed noise scale, a learned noise network, and an uncertainty‑gated controller that reallocates exploration based on novelty and competence signals without task labels. The uncertainty‑gated controller prevents task collapse across all tested seeds, whereas the other two approaches consistently cause collapse, demonstrating its effectiveness in preserving task performance during fine‑tuning.

By Mehmet Turan Yard{\i}mc{\i}, Yunus Emre \c{C}o\u{g}urcu