The paper investigates a Multi-Scale Spectral Attention Module (MSAM) for hyperspectral image segmentation in autonomous driving. MSAM uses three parallel 1D convolutions with different kernel sizes (1–11) and adaptive feature aggregation, integrated into UNet’s skip connections. Experiments on urban driving datasets show that MSAM improves mIoU by 2.32% and mF1 by 2.88% over baseline UNet-SC while keeping GPU performance competitive, with optimal kernel combinations varying by dataset.
By Imad Ali Shah, Jiarong Li, Tim Brophy, Martin Glavin, Edward Jones, Enda Ward, Brian Deegan
arXiv:2607. 16338v1 Announce Type: cross Abstract: This article presents DMFNet, a dual-backbone multiscale feature fusion framework with residual feature propagation and spatial attention for remote sensing scene classification.
By Anamitra Ghosh, Abhiroop Chatterjee, Susmita Ghosh
arXiv:2608.06205v2 Announce Type: replace
Abstract: Multispectral object detection combines visible and thermal imagery to improve perception under challenging illumination and environmental conditio...
By Nima Hatami, Karim Faez, Saeed Sharifian, Hamidreza Amindavar
AGSA-Net is a hyperspectral image classification framework that incorporates spectral unmixing priors through an abundance-guided self‑attention network. It first estimates physically meaningful subpixel abundance maps with non‑negativity and sum‑to‑one constraints, then uses these abundances to build an affinity prior that directs a spectral transformer to focus on class‑discriminative interactions. The transformer features are fused with compact abundance descriptors for final classification, and experiments on Indian Pines, Augsburg, and Berlin datasets show improved performance, especially in heterogeneous urban scenes.
By Nafisa Anjum, Satavisa Dey Borno, Ananna Saha, Mir Faiyaz Hossain, Sifat Momen, Nabeel Mohammed, Shafin Rahman
arXiv:2609.21522v1 Announce Type: new
Abstract: Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representation...
By Hang Cheng, Yan Chen, Mingyu Fan, Long Zeng
arXiv:2510.10471v3 Announce Type: replace-cross
Abstract: Environmental perception systems are crucial for high-precision mapping and autonomous navigation, with LiDAR serving as a core sensor provid...
By Chuang Chen, Yi Lin, Bo Wang, Jing Hu, Xi Wu, Wenyi Ge
Hyperspectral remote sensing has advanced across diverse deep learning paradigms, including spectral spatial CNNs, Vision Transformers, Mamba, graph neural networks, Kolmogorov Arnold networks, and se...
HyperVision introduces the first ground‑based hyperspectral pre‑trained backbone, addressing challenges of varying spectral configurations, limited annotations, and dataset diversity. It employs a channel‑adaptive dynamic embedding to unify heterogeneous inputs, a multi‑source pseudo‑labeling strategy combining SAM2 spatial cues with HyperFree spectral details, and cross‑modal knowledge distillation from a pre‑trained RGB vision model. Trained on 15k images from 26 datasets, HyperVision achieves significant improvements—up to 16.3% relative gain in hyperspectral semantic segmentation, 2.1% in object tracking AUC, and 35.5% reduction in salient object detection MAE—while requiring only head‑only adaptation.
By Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Xiangyu Liu, Fengchao Xiong, Jianfeng Lu, Chengrong Chen, Jun Zhou
The technical report introduces Hyperspectral Image Models, a modular framework that unifies 55 deep‑learning models across six paradigms for hyperspectral remote sensing. It standardizes tensor conventions, evaluation protocols, and dataset handling, integrating 24 benchmark scenes from various sensors and providing tools to avoid train‑test overlap. Experiments across 1,320 model‑scene combinations show that scene difficulty outweighs architecture, with no single paradigm dominating and small models achieving performance comparable to much larger ones.
By Tanishq Rachamalla, Aryan Das, Srishti Kaushik, Swalpa Kumar Roy
The paper compares classical machine learning algorithms—such as Logistic Regression, SVM, Random Forest, XGBoost, and CatBoost—with Tabular Deep Learning models (TabNet, FT-Transformer, TabTransformer, TabSeq, and 1D CNNs) for urban land cover classification using a UCI dataset derived from high‑resolution aerial imagery. It evaluates performance across nine land cover classes, addressing challenges like high dimensionality, heterogeneous features, and class imbalance by applying weighted cross‑entropy loss for deep models and measuring accuracy, macro‑precision, macro‑recall, macro‑F1, AUC‑ROC, and confusion matrices. Results indicate that while tree ensembles remain strong baselines, Tabular Deep Learning can match or surpass them when non‑linear interactions are prominent and imbalance handling is effective.
By Muntasir Tabasum, Tanpia Tasnim, Md. Ekramul Islam, Al Zadid Sultan Bin Habib
arXiv:2608. 16241v1 Announce Type: cross Abstract: Feature extraction for hyperspectral image classification is conventionally addressed using rigid tensor decompositions that fail to capture complex spatio-spectral interdependencies, or heavily parameterized convolutional neural networks that are computationally expensive.
By S\"uha Tuna, \"Ulker Ba\c{s}ar
Token Clustering and Semantic Sequence Mamba (STMamba) is a new approach for hyperspectral image classification that organizes sparse tokens into semantically coherent sequences. It uses a hierarchical encoder-decoder with a Token Clustering Module (TCM) to select semantic tokens and a Cross-scale Neighborhood Attention (CNA) Upsampler to restore dense features. At the micro level, density-aware clustering and a quadtree-based dynamic selection keep sparse, spatially distributed tokens, while Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture long-range spatial and spectral dependencies within homogeneous semantic token sequences. Experiments on three large-scale benchmark datasets show that STMamba outperforms state‑of‑the‑art methods in both quantitative and qualitative metrics.
By Yimin Zhu, Mahmood Elahi, Lincoln Linlin Xu