Token Clustering and Semantic Sequence Mamba (STMamba) is a new approach for hyperspectral image classification that organizes sparse tokens into semantically coherent sequences. It uses a hierarchical encoder-decoder with a Token Clustering Module (TCM) to select semantic tokens and a Cross-scale Neighborhood Attention (CNA) Upsampler to restore dense features. At the micro level, density-aware clustering and a quadtree-based dynamic selection keep sparse, spatially distributed tokens, while Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture long-range spatial and spectral dependencies within homogeneous semantic token sequences. Experiments on three large-scale benchmark datasets show that STMamba outperforms state‑of‑the‑art methods in both quantitative and qualitative metrics.
By Yimin Zhu, Mahmood Elahi, Lincoln Linlin Xu
Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones. However, MambaOut demonstrates that a Gated CNN block can match or exceed VMamba on image classification, questioning the necessity of SSMs for vision.
HyperVision introduces the first ground‑based hyperspectral pre‑trained backbone, addressing challenges of varying spectral configurations, limited annotations, and dataset diversity. It employs a channel‑adaptive dynamic embedding to unify heterogeneous inputs, a multi‑source pseudo‑labeling strategy combining SAM2 spatial cues with HyperFree spectral details, and cross‑modal knowledge distillation from a pre‑trained RGB vision model. Trained on 15k images from 26 datasets, HyperVision achieves significant improvements—up to 16.3% relative gain in hyperspectral semantic segmentation, 2.1% in object tracking AUC, and 35.5% reduction in salient object detection MAE—while requiring only head‑only adaptation.
By Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Xiangyu Liu, Fengchao Xiong, Jianfeng Lu, Chengrong Chen, Jun Zhou
arXiv:2607. 18625v1 Announce Type: cross Abstract: Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones.
By Jin Yu, Juyoun Park
The paper investigates a Multi-Scale Spectral Attention Module (MSAM) for hyperspectral image segmentation in autonomous driving. MSAM uses three parallel 1D convolutions with different kernel sizes (1–11) and adaptive feature aggregation, integrated into UNet’s skip connections. Experiments on urban driving datasets show that MSAM improves mIoU by 2.32% and mF1 by 2.88% over baseline UNet-SC while keeping GPU performance competitive, with optimal kernel combinations varying by dataset.
By Imad Ali Shah, Jiarong Li, Tim Brophy, Martin Glavin, Edward Jones, Enda Ward, Brian Deegan
arXiv:2604.08884v2 Announce Type: replace-cross
Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evide...
By Xinyu Zhang, Zurong Mai, Qingmei Li, Xiaoya Fan, Zjin Liao, Haoyuan Liang, Yibin Wen, Yuhang Chen, Chan Tsz Ho, Bi Tianyuan, Ruifeng Su, Zihao Qiang, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
3D vision-language models (3D VLMs) enable spatial reasoning over multi-view scenes but suffer from substantial token redundancy due to duplicated observations and large uninformative regions, leading to high computational cost. Although visual token compression has shown promise in accelerating 2D VLMs, it fails to capture the structured nature of 3D scenes and leads to incomplete spatial coverage and loss of fine-grained details.
arXiv:2609.06959v1 Announce Type: cross
Abstract: 3D point cloud semantic segmentation is essential for real-world spatial understanding, yet the prohibitive cost of human annotations motivates unsup...
By Zhenghao Zhang, Xinjie Wang, Wei Wang, Jun Zhang, Hanyun Wang
AGSA-Net is a hyperspectral image classification framework that incorporates spectral unmixing priors through an abundance-guided self‑attention network. It first estimates physically meaningful subpixel abundance maps with non‑negativity and sum‑to‑one constraints, then uses these abundances to build an affinity prior that directs a spectral transformer to focus on class‑discriminative interactions. The transformer features are fused with compact abundance descriptors for final classification, and experiments on Indian Pines, Augsburg, and Berlin datasets show improved performance, especially in heterogeneous urban scenes.
By Nafisa Anjum, Satavisa Dey Borno, Ananna Saha, Mir Faiyaz Hossain, Sifat Momen, Nabeel Mohammed, Shafin Rahman
MambaMPD is a new segmentation framework that leverages Vision Mamba models for marine pollution detection in remote‑sensing imagery. It introduces two structural priors—Frequency‑Aware Augmentation (FAA) and multi‑scale Edge‑Guided Attention (EGA)—to better capture low‑contrast, fragmented pollution patterns and sharpen boundaries. Experiments on the MADOS and M4D datasets show that MambaMPD outperforms existing methods in mIoU while using far less computation than foundation‑model approaches.
By Shuaiyu Chen, Wei Han, Peng Ren, Chunbo Luo, Zeyu Fu
arXiv:2609.36648v1 Announce Type: new
Abstract: Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics...
By Yuanwei Hu, Bo Peng, Yuheng Jia, Xinting Hu, Yadan Luo, Wenjie Zhu
UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.
By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun