FACT: Failure-Aware Causal Training for World-Action Models
arXiv:2608. 10232v1 Announce Type: cross Abstract: Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation.
The paper introduces TRACC, a pipeline that learns humanoid skills from a single failed human video by first imitating the usable portion of the motion trajectory and then completing the task based on the inferred outcome. It treats the motion prefix before failure as prior knowledge and uses a task-completion reward to guide learning toward the intended goal without needing a successful demonstration. The method is evaluated on six failed tasks from the Oops! dataset, showing its effectiveness in learning from failures.
arXiv:2608. 10232v1 Announce Type: cross Abstract: Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation.
BeyondRetarget is an end‑to‑end framework that learns to generate executable humanoid robot motions directly from monocular RGB videos, bypassing the need for an explicit human motion representation. By learning robot‑oriented implicit representations and incorporating a contact‑aware motion optimization mechanism, the method captures cross‑morphology motion structures and improves temporal consistency and physical plausibility. Experiments demonstrate that BeyondRetarget achieves higher execution success rates, lower latency, and greater accuracy and robustness in both simulation and real humanoid robots.
Zero-WAM introduces a causal video-action model that enables robots to perform unseen manipulation tasks by following in-context human video guidance. The authors create HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks, and propose an in-context future chunk prediction objective to prevent shortcut learning. In simulation, Zero-WAM attains a 47.0% success rate on seven unseen tasks, outperforming the best video-action baseline by 29.5 percentage points, and demonstrates real‑world generalization to complex, long‑horizon, and fine‑grained tasks.
arXiv:2606. 17011v1 Announce Type: cross Abstract: Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models.
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task ca...
arXiv:2602. 19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.
arXiv:2506. 20668v3 Announce Type: replace-cross Abstract: We propose DemoDiffusion, a simple method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring task-specific training or paired human-robot data.
RoboTok is an internet‑scale data engine that retrieves human manipulation videos from the web to train dexterous robot policies. It learns a latent motion space from 3D hand trajectories in actor‑centered reference frames, allowing manipulation behaviors to be compared across different viewpoints, scenes, and occlusions while remaining compact for efficient search. Experiments show RoboTok retrieves more relevant demonstrations and improves downstream robot task success compared to existing retrieval methods.
KnowDemo is a framework that generates diverse robot demonstrations from human videos by leveraging structured manipulation knowledge. It uses a vision‑language model to extract task requirements and permissible execution variations, then resolves these against target‑scene entities to guide candidate generation and screening before motion planning. The resulting demonstrations feature multimodal behavior, alternative contact strategies, and valid subtask orders, and have been shown to improve planning success and enable sim‑to‑real policy transfer across three tasks.
arXiv:2609.38057v1 Announce Type: new Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...
Triplet2Track (TTS) is a closed‑loop long‑horizon imitation learning system that uses human videos to reduce robot‑collected data. It represents high‑level subgoals as instance‑grounded triplets, converts them into continuous track priors for execution, and monitors task progress from observations for online replanning. In diverse real‑world long‑horizon tasks, TTS achieves a 74.8% average success rate and supports object‑level and compositional generalization.
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effe...