Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 02753v1 Announce Type: cross Abstract: Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective.
arXiv:2603. 02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent interaction.
arXiv:2608.29910v1 Announce Type: new Abstract: Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling appli...
arXiv:2512. 02473v2 Announce Type: replace-cross Abstract: Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions.
PhysWAM is a unified world-action model for autonomous driving that jointly denoises multiview video, metric depth, and ego motion using a flow‑matching transformer. It introduces Coupled Point Projection (CPP), a geometric constraint that aligns generated depth points with LiDAR data after applying the predicted SE(3) ego motion, thereby enforcing physical consistency. At inference, trajectory selection uses a simple label‑free consensus rule, and the model demonstrates strong planning performance, zero‑shot transfer to unseen environments, and accurate, temporally coherent depth and video predictions.
arXiv:2608. 07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions.