arXiv Computer Vision

Agentic Multimodal Models for Environmental Hyperspectral Unmixing

arXiv Computer Vision
4d ago

HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

HyperVision introduces the first ground‑based hyperspectral pre‑trained backbone, addressing challenges of varying spectral configurations, limited annotations, and dataset diversity. It employs a channel‑adaptive dynamic embedding to unify heterogeneous inputs, a multi‑source pseudo‑labeling strategy combining SAM2 spatial cues with HyperFree spectral details, and cross‑modal knowledge distillation from a pre‑trained RGB vision model. Trained on 15k images from 26 datasets, HyperVision achieves significant improvements—up to 16.3% relative gain in hyperspectral semantic segmentation, 2.1% in object tracking AUC, and 35.5% reduction in salient object detection MAE—while requiring only head‑only adaptation.

By Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Xiangyu Liu, Fengchao Xiong, Jianfeng Lu, Chengrong Chen, Jun Zhou
arXiv AI
Aug 25

Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework

arXiv:2604.08884v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evide...

By Xinyu Zhang, Zurong Mai, Qingmei Li, Xiaoya Fan, Zjin Liao, Haoyuan Liang, Yibin Wen, Yuhang Chen, Chan Tsz Ho, Bi Tianyuan, Ruifeng Su, Zihao Qiang, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
arXiv Machine Learning
3d ago

MWIR-4-Plastic: The Identification of Complex End-of-Life Industrial Plastic using Mid-wave Infrared Hyperspectral Imaging and Machine Learning

The paper presents the first publicly available hyperspectral imaging (HSI) dataset of shredded black plastics from end‑of‑life vehicle waste, covering four industrial polymers across RGB, VNIR, SWIR, and MWIR modalities. It introduces a multi‑modal spectral‑spatial framework that combines foreground isolation, pixel‑wise classification, and object‑level majority voting, leveraging hyperspectral transformers and chemometric band selection to accurately classify complex black plastics. The study benchmarks nine processing methods—including chemometric, machine learning, and deep learning architectures—providing a reproducible, comprehensive benchmark for industrial hyperspectral object analysis.

By Elias Arbash, Andr\'ea de Lima Ribeiro, Filipa Sim\~oes, Ahmed Jamal Afifi, Aldino Rizaldy, Yuleika Madriz, Samuel Thiele, Sandra Lorenz, Margret Fuchs, Pedram Ghamisi, Paul Scheunders, Richard Gloaguen
arXiv Machine Learning
3d ago

Minerals in the Wild: A Hyperspectral-XRF Dataset for Elemental Composition Estimation

arXiv:2608.30537v1 Announce Type: cross Abstract: Rapid mineral characterization is essential for applications ranging from mineral exploration to industrial ore processing. To this end, Hyperspectra...

By Eleftheria Tetoula-Tsonga (Institute of Communication and Computer Systems, Athens, Greece), George Arvanitakis (Geonova, Athens, Greece), Theodoros Giannakas (Institute of Communication and Computer Systems, Athens, Greece)
arXiv AI
Jun 18

LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis

arXiv:2606. 18661v1 Announce Type: cross Abstract: Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios.

By Chengfu Liu, Dongyang Hou, Junwu Xiang, Cheng Yang, Xuezhi Cui, Zeyuan Wang, Liangtian Liu, Zelang Miao
arXiv AI
Aug 19

Multi-Scale Spectral Attention Module-based Hyperspectral Segmentation in Autonomous Driving Scenarios

The paper investigates a Multi-Scale Spectral Attention Module (MSAM) for hyperspectral image segmentation in autonomous driving. MSAM uses three parallel 1D convolutions with different kernel sizes (1–11) and adaptive feature aggregation, integrated into UNet’s skip connections. Experiments on urban driving datasets show that MSAM improves mIoU by 2.32% and mF1 by 2.88% over baseline UNet-SC while keeping GPU performance competitive, with optimal kernel combinations varying by dataset.

By Imad Ali Shah, Jiarong Li, Tim Brophy, Martin Glavin, Edward Jones, Enda Ward, Brian Deegan