arXiv AI

Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking

The paper introduces TSDA-Track, a Template-Search Domain Adaptation framework designed to reduce modality gaps in cross‑modal visual object tracking. Two variants are explored: Pre‑AFA TSDA‑Track uses adversarial alignment before transformer interaction, while Enc‑CFA TSDA‑Track applies contrastive alignment after interaction to strengthen cross‑modal correspondence. Experiments on datasets such as LasHeR, RGBT234, GTOT, and Anti‑UAV‑024 show that both variants outperform state‑of‑the‑art trackers, with Pre‑AFA achieving an SR/PR of 43.2/56.0 on RGBT234 under the modality‑switch protocol.

arXiv Computer Vision
Sep 25

OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding

OptiSAR-Net++ introduces a new cross‑domain remote sensing visual grounding task (CD‑RSVG) and the first large‑scale benchmark dataset, OptSAR‑RSVG. The framework replaces Transformer decoding with a CLIP‑based contrastive approach, employing a patch‑level Low‑Rank Adaptation Mixture of Experts for efficient cross‑domain feature decoupling and a text‑guided dual‑gate fusion module for improved semantic‑visual alignment. Experiments show state‑of‑the‑art performance on OptSAR‑RSVG and DIOR‑RSVG, with notable gains in localization accuracy and computational efficiency.

By Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan
arXiv Computer Vision
3d ago

End-to-End Self-Supervised RGB-T Tracking without Modality Misleading

ESMTrack is a fully end‑to‑end self‑supervised RGB‑T tracking framework that eliminates the need for costly modality‑aligned bounding boxes or offline pseudo‑label generation. It learns discriminative, temporally consistent representations using a grounding triplet loss on the initial annotated frame and a cross‑frame temporal triplet loss on unlabeled search frames, with reliable samples selected via forward‑backward consistency. A three‑branch architecture (fusion, RGB, thermal) and a modality decoupling mechanism mitigate modality dominance bias, enabling competitive state‑of‑the‑art performance, strong cross‑dataset generalization, and real‑time inference on five RGB‑T benchmarks.

By Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou, Javen Qinfeng Shi, Xinyu Zhang
arXiv Computer Vision
Sep 16

MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking

MAETrack introduces a lightweight framework to adapt pretrained masked autoencoder (MAE) representations for 3D single object tracking (SOT). It uses Layer‑Selective Initialization (LSI) to keep shallow geometric layers from the pre‑training while re‑initializing deeper layers, and Geometric Residual Gating (GRG) to emphasize salient regions in BEV features before template‑search fusion. Experiments on standard 3D SOT benchmarks show consistent improvements over vanilla fine‑tuning with minimal computational cost.

By Sifan Zhou, Qiwei Wang, Linyue Tan, Ziyu Liu, Ziyu Zhao, Xiaobo Lu