arXiv Computer Vision

Socialized UAV Cross-Task Learning: Towards Cross-Granularity Collaboration through Hierarchical Interaction

The paper introduces CrossUAV, a benchmark for joint object detection and instance segmentation in UAV imagery, and proposes Cross-Granularity Socialized Collaboration (CGSC), a framework that regulates hierarchical interactions between tasks. CGSC progressively activates cross-task exchanges and adaptively adjusts interaction strength based on task contribution, aiming to reduce interference and exploit complementary coarse- and fine-grained knowledge. Experiments show consistent improvements on both detection and segmentation tasks, supporting the effectiveness of hierarchical dynamic interaction for cross-granularity collaboration.

arXiv AI
Aug 20

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

GrabVG is a visual grounding framework for UAV imagery that tackles the challenges of small, densely packed, and visually similar objects. It splits the task into preattentive hypothesis search and graph‑attentive feature binding, using distillation‑guided proposals and a sparse graph to capture intra‑ and inter‑instance relationships. Experiments on AerialVG and AerialSense show that GrabVG achieves higher accuracy and speed, outperforming baselines by significant margins.

By Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu, Jiaxu Tian, Yuheng Li, Guangqian Guo, Shan Gao
Hugging Face Trending Papers
Aug 19

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

GrabVG is a visual grounding framework for UAV imagery that tackles the challenges of small, densely packed, and visually similar objects by separating the task into preattentive hypothesis search and graph-attentive feature binding. It first generates a compact set of reliable object hypotheses using distillation-guided proposal induction and text-aware filtering, then constructs a sparse graph where language-guided visual cues and inter-instance topological relationships are jointly bound and propagated via graph attention. Experiments on AerialVG and AerialSense demonstrate that GrabVG achieves a strong accuracy–speed trade‑off, reaching 67.31% and 80.34% Acc@0.5 and outperforming baselines by 10.55 and 8.76 percentage points.

arXiv Computer Vision
Sep 22

Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT

This paper conducts a systematic empirical study of multi‑object tracking (MOT) algorithms, focusing on how detection and association components affect overall performance. By evaluating state‑of‑the‑art methods on benchmarks such as MOT16/17/20, SportsMOT, DanceTrack, and CrowdTrack, the authors find that detection quality has a far greater impact than association strategies, and that transformer‑based end‑to‑end models are more robust to detection variations but computationally expensive. The study provides a unified pipeline diagram and practical guidance for researchers and practitioners in selecting and designing MOT systems.

By Linh Van Ma, Juhua Hu, Wei Cheng, Unse Fatima, Moongu Jeon
arXiv Computer Vision
Sep 3

Towards Zero-Shot Transfer Across Embodiments For Driving VLAs

The paper investigates how Vision‑Language‑Action (VLA) models can generalise across different driving environments and camera setups. It introduces a multi‑dataset training strategy and an auxiliary objective called BEV‑Forcing, which injects bird‑eye‑view spatial information into the VLA backbone to improve both in‑distribution and out‑of‑distribution performance on a limited number of camera rigs. The authors observe that while BEV‑Forcing helps when training data is scarce, its advantage diminishes as the number of training embodiments grows, suggesting that scaling diversity may reduce the impact of such auxiliary tasks.

By Caio Azevedo, Stefano Sabatini, Sascha Hornauer, Fabien Moutarde
arXiv Computer Vision
Sep 2

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

The paper investigates how visual understanding and generation objectives interact within unified multimodal models (UMMs). At the representation level, each objective enriches the other, but forcing them through the same computation path can cause one to dominate; a task‑decoupled architecture mitigates this. At the task and system levels, the authors demonstrate bidirectional transfer between shared knowledge and superior performance of an end‑to‑end UMM over a planner–executor pipeline on complex tasks.

By Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu