arXiv:2608.27971v1 Announce Type: new
Abstract: Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scen...
By Jingpu Yang, Debin Tang, Yilin Sun, Fengxian Ji, Jiahua Zhu, Wenrui Ding, Yufeng Wang
arXiv:2511.17681v2 Announce Type: replace
Abstract: Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal...
By Weiyi Lv, Ning Zhang, Hanyang Sun, Haoran Jiang, Kai Zhao, Yixiao Gu, Jing Xiao, Dan Zeng
arXiv:2607. 17157v1 Announce Type: cross Abstract: Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time.
By Yanrong Qin, Xiaoyan Cao, Yao Yao
OptiSAR-Net++ introduces a new cross‑domain remote sensing visual grounding task (CD‑RSVG) and the first large‑scale benchmark dataset, OptSAR‑RSVG. The framework replaces Transformer decoding with a CLIP‑based contrastive approach, employing a patch‑level Low‑Rank Adaptation Mixture of Experts for efficient cross‑domain feature decoupling and a text‑guided dual‑gate fusion module for improved semantic‑visual alignment. Experiments show state‑of‑the‑art performance on OptSAR‑RSVG and DIOR‑RSVG, with notable gains in localization accuracy and computational efficiency.
By Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan
arXiv:2609.12261v1 Announce Type: new
Abstract: Multi-object tracking (MOT) is dominated by the tracking-by-detection paradigm, whose methods typically rely on a small set of hyperparameters that are...
By Momir Ad\v{z}emovi\'c
ESMTrack is a fully end‑to‑end self‑supervised RGB‑T tracking framework that eliminates the need for costly modality‑aligned bounding boxes or offline pseudo‑label generation. It learns discriminative, temporally consistent representations using a grounding triplet loss on the initial annotated frame and a cross‑frame temporal triplet loss on unlabeled search frames, with reliable samples selected via forward‑backward consistency. A three‑branch architecture (fusion, RGB, thermal) and a modality decoupling mechanism mitigate modality dominance bias, enabling competitive state‑of‑the‑art performance, strong cross‑dataset generalization, and real‑time inference on five RGB‑T benchmarks.
By Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou, Javen Qinfeng Shi, Xinyu Zhang