PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
arXiv:2606. 28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation.
EVEWorld introduces a physical evolution-supervision framework for embodied world models, addressing the issue of Model Laziness by focusing on physical consistency rather than visual fidelity. The framework comprises Instance-Guided Restoration (IGR) to enforce instance consistency and Temporal Instance Alignment (TIA) to align target instances across adjacent frames. Experiments on DreamGenBench, EWMBench, and PBench show an 87.5% reduction in the Model Laziness Rate (MLR) compared to GigaWorld-0, and the model ranks 6th in JEPA Similarity on the WorldArena 2.0 Track 1 leaderboard.
arXiv:2606. 28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation.
arXiv:2607. 26903v1 Announce Type: new Abstract: The key bottleneck in embodied AI is not model architecture but data.
arXiv:2609.09210v1 Announce Type: cross Abstract: Teleoperated demonstrations are often multimodal even when the underlying dynamics are nearly deterministic given the executed action. We argue that...
arXiv:2609.25627v1 Announce Type: cross Abstract: General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate prec...
arXiv:2607. 11643v1 Announce Type: cross Abstract: Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints.
RIFAR is a new continual learning method for robots that uses reliability screening and drift-aware replay to mitigate forgetting while keeping storage low. It reconstructs past trajectories from short demonstration prefixes and employs a frozen inverse-dynamics model to verify action‑visual consistency. In experiments on LIBERO suites and real‑world tasks, RIFAR outperforms previous generative replay approaches, achieving high performance with only a small fraction of stored steps.
PhysBrain 1.5 is a unified vision‑language model that learns to understand physical environments, generate actions, and predict future states by encoding language, end‑effector motion, and dense visual targets as discrete sequences and training them with autoregressive next‑token prediction. The model is pre‑trained on human interaction videos and fine‑tuned on human demonstrations, robot trajectories, and simulated experience, achieving an average score of 72.5 across 28 embodied understanding benchmarks and outperforming other open‑source models on 14 of them. It also demonstrates the ability to produce end‑effector trajectories and predict future scenes with spatially aligned RGB, depth, and robot‑mask outputs.
arXiv:2601.13247v2 Announce Type: replace-cross Abstract: Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural ground...
ZimaBlue is a scalable framework that learns generalizable World Action Models (WAMs) from large-scale egocentric videos. It follows a three-stage curriculum: causal video pre‑training, video‑action mid‑training with a unified action representation, and final specialization to a target robot. The system employs an asynchronous Slow‑Fast architecture to enable real‑time 30 Hz action prediction, achieving a jump in real‑robot zero‑shot success from 36.1% to 77.8% when leveraging over 120,000 hours of embodied video.
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved: (1) reasoning can effectively improve the performance of the policy, but its scaling mechanism has seldom been studied; (2) historical information is essential, as embodied tasks are inherently long-horizon and sequential, making sole reliance on current observations for action scaling inadequate due to the lack of historical context utilization.
arXiv:2609.28236v1 Announce Type: new Abstract: Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounte...