Context-Guided Semantic Alignment for Feature Fusion Networks
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 02092v2 Announce Type: replace Abstract: Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics.
The paper introduces an Attention-Driven Complementarity Resampling framework to enhance cross-modality object detection. It employs a shared channel spatial attention mechanism that exchanges semantic masks between modalities, encouraging the backbone to learn generalized features. Additionally, a learnable channel competition module samples and aggregates features channel‑wise, improving robustness and achieving competitive results on multiple datasets.
arXiv:2604. 00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks.
arXiv:2606. 08156v1 Announce Type: cross Abstract: Vision Transformers (ViTs) achieve strong performance but suffer from high computational costs due to quadratic self-attention complexity.
arXiv:2511. 12810v2 Announce Type: replace-cross Abstract: Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blend seamlessly into their environments due to high similarity in color, texture, and size.
arXiv:2606. 29136v1 Announce Type: cross Abstract: Event cameras capture sparse brightness changes with high temporal resolution and high dynamic range, compensating for the deficiencies of the conventional RGB frames.