arXiv Computer Vision By Fabian G\"ulhan, Emil Mededovic, Yuli Wu, Johannes Stegmaier

SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors

Read the original on arXiv Computer Vision →

SelfMOTR proposes a detector‑free approach to multi‑object tracking that decouples proposal discovery from association by generating internal detection priors. The method builds on end‑to‑end transformer trackers, showing that joint detection‑association decoding retains hidden detection capacity and can be leveraged without external detectors. Experiments demonstrate competitive results, achieving 69.2 HOTA on DanceTrack and 71.1 HOTA on Bird Flock Tracking.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 4

ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object Tracking

The paper introduces ORMOT, a new task that extends Referring Multi‑Object Tracking to omnidirectional 360° imagery, ensuring full scene context for language‑guided tracking. It presents ORSet, a dataset of 27 omnidirectional scenes with 848 language descriptions and 3,401 annotated objects, and introduces ORTrack, an LVLM‑driven framework that performs zero‑shot detection and robust cross‑frame association. Experiments on ORSet show that ORTrack achieves state‑of‑the‑art performance, establishing a strong baseline for future research.

By Zihan Zhou, Sijia Chen, Yanqiu Yu, En Yu, Wenbing Tao
arXiv Computer Vision
Sep 7

PuTR-CouT: Counting-by-Tracking in Camera-Trap Image Sequences

PuTR-CouT is a transformer‑based counting‑by‑tracking framework designed for camera‑trap image sequences. It generates synthetic training data using structural priors to create pseudo‑tracking labels, enabling the tracker to associate detections across frames and estimate per‑species counts. The method improves upon the MaxBoxCount baseline on the iWildCam 2021 benchmark, offering competitive counting results along with multi‑species predictions and track‑level verification.

By Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos
arXiv Computer Vision
Aug 31

TQD-Track: Temporal Query Denoising for 3D Multi-Object Tracking

TQD-Track introduces Temporal Query Denoising (TQD) for 3D Multi‑Object Tracking, extending query denoising beyond single frames by initializing denoising queries from previous‑frame ground truths and propagating them as independent association candidates. The method enriches track queries with temporal context and instance‑specific features, and incorporates diverse noise types to emulate real‑world tracking challenges. Experiments on nuScenes and Argoverse 2 show consistent improvements across multiple MOT baselines with only training‑process modifications.

By Yutong Yang, Shuxiao Ding, Mohammed Amine Bencheikh Lehocine, Julian Wiederer, Markus Braun, Peizheng Li, Juergen Gall, Bin Yang