arXiv AI

Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

Triplet2Track (TTS) is a closed‑loop long‑horizon imitation learning system that uses human videos to reduce robot‑collected data. It represents high‑level subgoals as instance‑grounded triplets, converts them into continuous track priors for execution, and monitors task progress from observations for online replanning. In diverse real‑world long‑horizon tasks, TTS achieves a 74.8% average success rate and supports object‑level and compositional generalization.

arXiv AI
Sep 18

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

HIL-UMI is a policy-guided Universal Manipulation Interface that enables robot‑free, human‑in‑the‑loop post‑training of vision‑language‑action models. By querying the current policy during handheld demonstrations and using an Energy Score to detect out‑of‑distribution states, it selectively collects new data and refines a progress‑based advantage estimator. The updated estimator then drives advantage‑conditioned behavioral cloning, improving performance on long‑horizon and precise manipulation tasks while reducing per‑frame collection time compared to HG‑DAgger.

By Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao, Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang, Ruodai Li, Hui Shen, Hao Dong
arXiv AI
Jul 1

Stage-Transition Dense Reward Modeling for Reinforcement Learning

arXiv:2606. 31377v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations.

By Yang Yang, Bingjie Chen, Zihan Wang, Yizhe Li, Guoping Pan, Yi Cheng, Houde Liu
arXiv AI
Aug 25

RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

The paper introduces RARM, a Reference‑Anchored Reward Model that uses a single successful demonstration to generate dense, progress‑aware rewards for reinforcement learning in robot manipulation. RARM is trained on general‑purpose videos with a contrastive temporal objective, requiring no task‑specific data or reward engineering. During deployment it matches rollout clips to reference clips and rewards only confident forward progress, reducing false positives. Experiments on nine simulated tasks and four real‑world tasks show that RARM achieves the best overall success rates, especially on long‑horizon tasks like cloth folding.

By Pengzhi Yang, Xinyu Wang, Pengyu Jing, Kehan Wen, Yiduo Qu, Zhenhao Huang, Minghao Fu, Xin Liu, Yaheng Shen, Fan Shi
arXiv AI
Jun 30

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.

By Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem
arXiv AI
Sep 23

SAIL: Test-Time Scaling for In-Context Imitation Learning with VLM

SAIL is a framework that transforms robot imitation learning into an iterative refinement problem, enabling test-time scaling of trajectory generation. It employs Monte Carlo Tree Search where each node represents a full trajectory and edges denote refinements, guided by an archive of successful trajectories, a vision‑language model for scoring, and step‑level feedback. Experiments on six manipulation tasks in simulation and real‑world settings show that higher test‑time compute consistently raises success rates, reaching up to 95% on complex tasks.

By Makoto Sato, Yusuke Iwasawa, Yujin Tang, So Kuroki
arXiv Computer Vision
6d ago

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

The paper introduces an interaction‑centric framework that unifies representations for two‑finger gripper manipulation across different robot embodiments. By using a parameterized universal gripper abstraction and a canonical gripper‑frame representation, the system infers sub‑tasks from language and RGB‑D inputs, grounds interaction triplets, and employs hybrid features and a Flow‑Matching Transformer to generate smooth 7‑DoF action sequences. Experiments in both simulation and real‑world settings show that this approach achieves competitive benchmark performance while enabling extreme cross‑embodiment and cross‑viewpoint zero‑shot sim‑to‑real transfer to heterogeneous robot platforms.

By Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang
arXiv Computer Vision
Aug 25

TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation

arXiv:2608.22296v1 Announce Type: cross Abstract: Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact th...

By Haoran Lin, Mingyu Yang, Pengfei Qi, Kehan Chen, Qiang Diao, Liangji Zeng, Wenrui Chen, Yaonan Wang, Kailun Yang
arXiv AI
Jul 16

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset

arXiv:2603. 11811v2 Announce Type: replace-cross Abstract: The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the prohibitive cost and scalability limits of human-in-the-loop collection paradigms.

By Yongzhong Wang, Keyu Zhu, Yong Zhong, Liqiong Wang, Jinyu Yang, Feng Zheng