arXiv Computer Vision
Sep 3

Evidential Deep Learning for Multi-Modal Anti-UAV Detection

The paper investigates the use of evidential deep learning (EDL) for multi‑modal anti‑UAV detection, comparing it with sigmoid baselines, Dempster‑Shafer evidence fusion, and uncertainty‑driven temporal sensor gating across three benchmarks (thermal tracking, RGB‑audio‑RF classification, and RGB‑IR tracking). EDL improves accuracy by up to 5.9 percentage points and better ranks classification errors, while the other components (DS fusion, Dirichlet vacuity, temporal gating) do not provide the expected benefits. The study concludes that the primary advantage of EDL stems from its training objective rather than its uncertainty estimates.

By Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
arXiv AI
Sep 10

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

CUSP (Collective Uncertainty through Semantic Opinion Pooling) is a training‑free framework that aggregates responses from multiple vision‑language models into a shared semantic space, producing a pooled opinion and two system‑level uncertainty signals: collective uncertainty (dispersion) and Jensen‑Shannon divergence (model conflict). It decomposes collective entropy into the mean of individual semantic entropies plus JSD, enabling reliable uncertainty estimation without token logits or calibration labels. In static ensembles, collective uncertainty outperforms baseline methods for error detection and abstention, while JSD excels in commercial settings, and the pooled prediction consistently improves accuracy over individual models. "whyItMatters":"CUSP provides a practical, model‑agnostic way to quantify system‑level reliability and improve decision‑making in multimodal reasoning tasks."

By Chung-En Johnny Yu, David Garcia, Brian Jalaian, Nathaniel D. Bastian
arXiv Computer Vision
Sep 4

Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection

The paper introduces FlexibleFusion, a method for infrared-visible object detection that adapts to both complete and missing-modality scenarios. It employs a Modality-Aware Experts Collaboration mechanism to selectively fuse cross-modal or intra-modal pathways, and a Residual Self-Paced Entropic Optimal Transport module to align heterogeneous feature distributions without heavy optimization. Experiments demonstrate consistent performance across various modality configurations.

By Yue Zhao, Hua Yu, Yukun Zhao, Yuzhi Zhang, Maoguo Gong, Xin Mei, Zhuping Hu, Yanchi Li, A. K. Qin
arXiv Computer Vision
Sep 14

RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation

RA‑SOD is a new RGB‑Thermal salient object detection framework that explicitly models the reliability of each modality. It introduces a reliability‑conditioned representation, an uncertainty‑guided dual‑stream refinement, and a pixel‑wise modality competition mechanism to adaptively compensate degraded features and suppress unreliable evidence. Experiments on four benchmarks show that RA‑SOD achieves state‑of‑the‑art performance and remains robust under severe modality degradation.

By Hongbo Gao, Zhengyu Li, Xueru Nie, Dihao Zhu, Lijun Zhao, Yunke Wang, Chang Xu