Hugging Face Trending Papers

Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model

arXiv AI
Sep 24

Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model

The paper investigates why latent‑flow world models that use a frozen self‑supervised latent space lose the ability to manipulate motion. It shows that the pretrained flow does not move the manipulated object and that training with latent‑only losses only produces stillness or teleport‑like motion. The authors introduce Decode‑Augmented Rollout Training (DART), which keeps the representation frozen but retrains the flow using decode‑path supervision, restoring temporal motion structure and improving prediction quality, even closing much of the gap to an oracle‑informed reference. The study also notes that pixel error alone can favor frozen predictions.

By Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu, Houde Liu
arXiv AI
3d ago

Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases

The paper introduces Action Forcing, a method that transforms ordinary unlabeled video into action‑supervised training data by extracting egomotion bases through principal component analysis of pixel displacements. This approach yields grounded throttle–yaw control signals without requiring instrumented platforms or manual annotation, and it trains a high‑capacity video model while preventing pixel‑level overfitting via an online latent critic. The authors also critique standard video generation metrics and propose a reference‑free evaluation that measures controllability, plausibility, conjuring, and geometric integrity, showing that their model can reverse, scale, and compose actions despite limited reverse‑action data.

By Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang
arXiv AI
Aug 3

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

arXiv:2607. 29235v1 Announce Type: cross Abstract: Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout.

By Peize Li, Ruimeng Zhang, Ru Zhang, Cong Huang, Kai Chen, Shanghang Zhang
arXiv Computer Vision
Sep 7

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.

By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
arXiv Computer Vision
6d ago

RotVLA: Rotational Latent Action for Vision-Language-Action Model

RotVLA introduces a Vision‑Language‑Action framework that replaces discrete latent action encoding with a continuous rotational latent action representation on the group SO(n). This design provides continuity, compositionality, and structured geometry that better capture real‑world action dynamics, and a triplet frame learning scheme enforces meaningful temporal dynamics while preventing degeneration. Trained with 1.7 B parameters on large cross‑embodiment datasets, RotVLA achieves state‑of‑the‑art performance on LIBERO and RoboTwin2.0 benchmarks and shows strong real‑world manipulation results.

By Qiwei Li, Xicheng Gong, Xinghang Li, Peiyan Li, Quanyun Zhou, Hangjun Ye, Jiahuan Zhou, Yadong Mu