arXiv:2606. 07233v1 Announce Type: cross Abstract: LiDAR-based 3D Multi-Object Tracking (MOT) typically relies solely on geometric information, which is often insufficient to distinguish between targets during prolonged occlusions or in crowded human-populated environments.
By Eduardo Borges, Lu\'is Garrote, Urbano J. Nunes
LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.
By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
Segmentation-based multi-object tracking (MOT) with foundation video models such as SAM2 offers strong localization quality, yet remains fragile in crowded, real-world scenes. In detector-prompted SAM...
arXiv:2607. 17157v1 Announce Type: cross Abstract: Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time.
By Yanrong Qin, Xiaoyan Cao, Yao Yao
arXiv:2609.17427v1 Announce Type: cross
Abstract: Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where targets temporarily or completely disappear f...
By Mais Mohammed, Sharifa Mohammed, Hanan Awadh, Haneen Bamaas, Raghad Bawazeer, Elham Alghamdi
The paper introduces MovingDroneCrowd++, a large-scale video dataset for dense crowd counting and tracking from moving drones, featuring varied flight altitudes, camera angles, and lighting. It presents two new methods: GD3A for Video Individual Counting and GIA-Track for Multi-Object Tracking, both leveraging group-wise density assignment and identity association to handle aerial challenges. Experiments demonstrate significant improvements, reducing counting error by 47.4% and boosting tracking accuracy by 64.6%.
By Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.
By Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan
arXiv:2608.22064v1 Announce Type: new
Abstract: We present our solution for the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. The challenge evaluates r...
By Mingqi Gao, Sijie Li, Jungong Han
arXiv:2606. 23604v2 Announce Type: replace-cross Abstract: The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation.
By Mohamed Nagy, Naoufel Werghi, Jorge Dias, Majid Khonji
Long-shot windsurfing video combines small targets, large camera pans, prolonged overlaps, and rapidly changing backgrounds. The desired output is not a generic MOT trace but a separate, stable rider-...
ESMTrack is a fully end‑to‑end self‑supervised RGB‑T tracking framework that eliminates the need for costly modality‑aligned bounding boxes or offline pseudo‑label generation. It learns discriminative, temporally consistent representations using a grounding triplet loss on the initial annotated frame and a cross‑frame temporal triplet loss on unlabeled search frames, with reliable samples selected via forward‑backward consistency. A three‑branch architecture (fusion, RGB, thermal) and a modality decoupling mechanism mitigate modality dominance bias, enabling competitive state‑of‑the‑art performance, strong cross‑dataset generalization, and real‑time inference on five RGB‑T benchmarks.
By Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou, Javen Qinfeng Shi, Xinyu Zhang
RA‑SOD is a new RGB‑Thermal salient object detection framework that explicitly models the reliability of each modality. It introduces a reliability‑conditioned representation, an uncertainty‑guided dual‑stream refinement, and a pixel‑wise modality competition mechanism to adaptively compensate degraded features and suppress unreliable evidence. Experiments on four benchmarks show that RA‑SOD achieves state‑of‑the‑art performance and remains robust under severe modality degradation.
By Hongbo Gao, Zhengyu Li, Xueru Nie, Dihao Zhu, Lijun Zhao, Yunke Wang, Chang Xu