arXiv Computer Vision By Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie, Wang sihan

ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

Read the original on arXiv Computer Vision →

ReWorld-Track introduces a recursive event world model for language‑guided multi‑camera tracking that explicitly carries association uncertainty into future predictions. By treating candidate matches and waiting as alternative target states, the model updates a persistent recurrent belief that preserves uncertainty across successive observations. This approach improves identity continuity and next‑camera accuracy, achieving HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, and reducing median arrival‑time error from 0.78 s to 0.71 s.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
3d ago

Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds

The paper introduces EvolvingNav, a system that builds a time‑indexed belief about moving targets in dynamic environments by combining timestamped 3D object histories with a persistence‑relocation model. It uses an event‑driven filter to update beliefs over time, incorporates RGB‑D evidence, and applies a zero‑shot vision‑language controller for action selection. The authors also present EvoWorld‑Bench, a large benchmark of human‑activity‑based scenes, and demonstrate that EvolvingNav outperforms baselines in both simulation and real‑robot experiments, especially when temporal patterns are learnable.

By Mingjian Gao, Zhaocheng Li, Haoyang Huang, Wenqiao Zhang, Yingjie Niu, Hao Zhou, Chao Li, Juncheng Li, Siliang Tang, Yueting Zhuang
arXiv Computer Vision
4d ago

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

CST‑WM is a causally structured world model designed for embodied visual tracking, where a robot must keep a moving target visible and recover it after occlusion or drift. The model separates state into target‑evidence, robot, and observation branches, removing direct action‑to‑target‑evidence edges to prevent causal hallucination and instead letting actions influence evidence through robot motion and resulting views. Evaluated on EVT‑Bench, Habitat 3.0, and real‑world trials with a Unitree Go2 quadruped, CST‑WM outperforms reactive trackers and other world‑model baselines in following, distance control, safety, and re‑acquisition, achieving 20 of 30 successful real‑world recoveries versus 14 for TrackVLA.

By Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang
arXiv AI
Sep 15

LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models

LPA-CWM introduces a Learned Physical Adjudicator (LPA) to improve counterfactual world models (CWM) for motion reasoning by learning to weight candidate responses based on visual context and response structure. The 3.0M‑parameter LPA is trained on dense MOVi‑F trajectories while keeping the CWM predictor and intervention generator frozen. A new Completeness‑aware Motion Correspondence (CMC) protocol evaluates localization, trajectory completeness, visibility, and continuity, and LPA‑CWM achieves significant gains on DAVIS and Kinetics subsets.

By Kunwei Wu, Xiang Liu, Guocai Yao, Junming Chen, Zhikang Chen, Min Zhang, Pengwei Wang, Sen Cui
arXiv Computer Vision
Sep 21

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.

By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang