arXiv Computer Vision

S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection

The paper introduces S2A, a semantic-to-spatial alignment framework designed for alignment‑free RGB‑T salient object detection. It employs a global‑guided hierarchical fusion module to refine intra‑modal features, an alignment‑free cross‑modal channel attention module to exchange semantic information, and a spatial deformable cross‑attention module to recover local spatial correspondence. These components collectively reduce misalignment‑induced feature contamination and achieve competitive performance on public benchmarks without additional bells and whistles.

arXiv Computer Vision
Aug 27

Context-Guided Semantic Alignment for Feature Fusion Networks

The paper introduces Feature Interaction Network (FINE), a lightweight semantic alignment module for feature fusion networks in object detectors. FINE refines low‑level features using high‑level contextual guidance through cross‑level attention, and employs Alignment‑Aware Token Sampling to reduce attention complexity. The resulting spatial‑channel modulation map selectively enhances semantically relevant pixels while preserving sub‑pixel localization, leading to improved detection accuracy with minimal computational overhead.

By Hyungseop Lee, Jiho Lee, Woochul Kang
arXiv Computer Vision
Sep 14

RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation

RA‑SOD is a new RGB‑Thermal salient object detection framework that explicitly models the reliability of each modality. It introduces a reliability‑conditioned representation, an uncertainty‑guided dual‑stream refinement, and a pixel‑wise modality competition mechanism to adaptively compensate degraded features and suppress unreliable evidence. Experiments on four benchmarks show that RA‑SOD achieves state‑of‑the‑art performance and remains robust under severe modality degradation.

By Hongbo Gao, Zhengyu Li, Xueru Nie, Dihao Zhu, Lijun Zhao, Yunke Wang, Chang Xu
Hugging Face Trending Papers
Aug 4

SGFormer: Structure-Guided Transformer for Robust Local Feature Matching

Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction, stereo mapping, and visual localization. While recent detector-free matching methods, like LoFTR, have advanced the field, the global features obtained by leveraging the global-range modeling capacity of the unconstrained attention mechanism compromise the model's attention to the salient structures in certain scenarios.

arXiv Computer Vision
Aug 27

TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection

TDFNet introduces a Tri-projection Deformable Fusion Network that uses equirectangular, cube map, and tangent projections to mitigate geometric distortions in panoramic salient object detection. It incorporates a cross-projection deformable attention module for geometry-aware sampling and a latitude-guided fusion module that balances ERP and CMP features using spherical latitude priors. The network’s three-branch encoding preserves global continuity, local detail, and boundary precision, improving detection performance over existing projection-based methods.

By Qiangqiang Zhou, Jiacong Yu, Jiawei Xu, Yong Chen, Xin Huang, Ping Li
arXiv AI
Sep 2

SARTM: Segment Any RGB Thermal Model with Language aided Distillation

The paper introduces SARTM, a framework that adapts the Segment Anything Model (SAM) for RGB‑thermal (RGB‑T) semantic segmentation. It fine‑tunes SAM with LoRA layers, incorporates language guidance, and employs a Cross‑Modal Knowledge Distillation module to bridge modality gaps. The approach also modifies the segmentation head and adds an auxiliary semantic head, achieving superior performance on MFNET, PST900, and FMB benchmarks.

By Dong Xing, Jinhe Zhang, Hang Yang, Yuqing Wang
arXiv Machine Learning
Aug 26

NAIMA: Semantics Aware RGB Guided Depth Super-Resolution

The paper introduces NAIMA, a guided depth super‑resolution framework that leverages global contextual semantic priors from pretrained vision transformer token embeddings. Its Guided Token Attention (GTA) module uses depth encodings as queries to attend over semantic tokens, with a zero‑initialized gate controlling the influence of semantic evidence. NAIMA achieves competitive in‑distribution performance while delivering superior cross‑dataset generalization without relying on decoded priors or auxiliary objectives.

By Tayyab Nasir, Daochang Liu, Ajmal Mian