arXiv Computer Vision

Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging

arXiv Machine Learning
Aug 28

MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework

MM‑Spectrum is a sparse Mixture‑of‑Experts framework designed to infer molecular structures from multimodal spectroscopic data. It introduces a modality‑aware routing mechanism that exposes spectral identity to the router, along with shared and interaction experts of heterogeneous capacities to capture both modality‑unique and cross‑modal synergistic information while reducing noise interference. Experiments across full‑modality, bimodal, and missing‑modality scenarios show consistent and substantial performance gains, supported by ablation studies and interpretability analyses.

By Hai-tao Yu, Nan Min, Zheng Fang, Hongyu Zhan, Yusen Tan, Yuhan Wang, Jun Xia
arXiv AI
Aug 25

Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework

arXiv:2604.08884v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evide...

By Xinyu Zhang, Zurong Mai, Qingmei Li, Xiaoya Fan, Zjin Liao, Haoyuan Liang, Yibin Wen, Yuhang Chen, Chan Tsz Ho, Bi Tianyuan, Ruifeng Su, Zihao Qiang, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
arXiv Computer Vision
4d ago

HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

HyperVision introduces the first ground‑based hyperspectral pre‑trained backbone, addressing challenges of varying spectral configurations, limited annotations, and dataset diversity. It employs a channel‑adaptive dynamic embedding to unify heterogeneous inputs, a multi‑source pseudo‑labeling strategy combining SAM2 spatial cues with HyperFree spectral details, and cross‑modal knowledge distillation from a pre‑trained RGB vision model. Trained on 15k images from 26 datasets, HyperVision achieves significant improvements—up to 16.3% relative gain in hyperspectral semantic segmentation, 2.1% in object tracking AUC, and 35.5% reduction in salient object detection MAE—while requiring only head‑only adaptation.

By Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Xiangyu Liu, Fengchao Xiong, Jianfeng Lu, Chengrong Chen, Jun Zhou
Hugging Face Trending Papers
Jun 3

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existing representation alignment metrics are symmetric, collapsing both modalities into a single score and hiding which modality drives cross-modal degradation.

arXiv Machine Learning
3d ago

Minerals in the Wild: A Hyperspectral-XRF Dataset for Elemental Composition Estimation

arXiv:2608.30537v1 Announce Type: cross Abstract: Rapid mineral characterization is essential for applications ranging from mineral exploration to industrial ore processing. To this end, Hyperspectra...

By Eleftheria Tetoula-Tsonga (Institute of Communication and Computer Systems, Athens, Greece), George Arvanitakis (Geonova, Athens, Greece), Theodoros Giannakas (Institute of Communication and Computer Systems, Athens, Greece)