Hyperspectral Image Models: Technical Report
Hyperspectral remote sensing has advanced across diverse deep learning paradigms, including spectral spatial CNNs, Vision Transformers, Mamba, graph neural networks, Kolmogorov Arnold networks, and se...
The technical report introduces Hyperspectral Image Models, a modular framework that unifies 55 deep‑learning models across six paradigms for hyperspectral remote sensing. It standardizes tensor conventions, evaluation protocols, and dataset handling, integrating 24 benchmark scenes from various sensors and providing tools to avoid train‑test overlap. Experiments across 1,320 model‑scene combinations show that scene difficulty outweighs architecture, with no single paradigm dominating and small models achieving performance comparable to much larger ones.
Hyperspectral remote sensing has advanced across diverse deep learning paradigms, including spectral spatial CNNs, Vision Transformers, Mamba, graph neural networks, Kolmogorov Arnold networks, and se...
HyperSAM is a promptable foundation model for hyperspectral remote sensing that integrates a data‑centric synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). The model generates full‑spectrum hyperspectral cubes from high‑resolution multispectral imagery using a physics‑informed abundance‑transfer generator, and employs SAM3‑derived pseudo‑masks for object‑centric supervision. With a frozen SAM3 RGB branch, a trainable hyperspectral encoder, ControlNet‑style feature injection, and a mixture‑of‑experts mask refiner, HyperSAM demonstrates strong generalization across diverse hyperspectral tasks such as classification, anomaly detection, change detection, target detection, and airborne oil‑spill mapping.
HyperVision introduces the first ground‑based hyperspectral pre‑trained backbone, addressing challenges of varying spectral configurations, limited annotations, and dataset diversity. It employs a channel‑adaptive dynamic embedding to unify heterogeneous inputs, a multi‑source pseudo‑labeling strategy combining SAM2 spatial cues with HyperFree spectral details, and cross‑modal knowledge distillation from a pre‑trained RGB vision model. Trained on 15k images from 26 datasets, HyperVision achieves significant improvements—up to 16.3% relative gain in hyperspectral semantic segmentation, 2.1% in object tracking AUC, and 35.5% reduction in salient object detection MAE—while requiring only head‑only adaptation.
The paper critiques the common practice of random pixel splits in hyperspectral image classification, noting that such splits allow test pixels to be adjacent to training pixels, inflating accuracy. It proposes a leakage‑free evaluation protocol that enforces spatial separation based on the model’s receptive field and applies it to ten diverse architectures, finding a significant drop in Macro‑F1 (average 0.147) and substantial changes in model rankings. The study also shows that all ten models misclassify the same pixels, indicating a spectral ambiguity in the data that current methods cannot resolve.
arXiv:2607. 23024v1 Announce Type: cross Abstract: High-resolution satellite imagery is the backbone of good land-cover classification, and without that, environmental monitoring, urban planning, and sustainable resource management all fall short.
SIMPLER is a pre‑fine‑tuning method that reduces inference and deployment costs for Earth Observation foundation models by pruning redundant layers. It uses layer‑wise representation similarity on unlabeled task data to identify and remove up to 79% of parameters without requiring gradients, magnitude heuristics, or hyperparameter tuning. Experiments on Prithvi‑EO‑2, TerraMind, and ImageNet‑pretrained ViT‑MAE show that SIMPLER retains 94% of baseline performance while achieving 2.1× faster training and 2.6× faster inference.
arXiv:2608.29609v1 Announce Type: new Abstract: Semantic segmentation is a crucial task for understanding Mars, the most Earth-like planet in our solar system. However, it is challenging because the...
arXiv:2607. 03644v1 Announce Type: cross Abstract: Decades of orbital missions have produced multi-modal remote sensing data for the Moon, spanning optical imagery, spectroscopy, thermal emission, radar, gravity, and elemental composition.
The paper presents an agentic framework that uses a large vision‑language model to refine hyperspectral unmixing results from existing modular pipelines. By iteratively gathering spectral and spatial evidence through tools such as spectral‑library retrieval and abundance‑map visualization, the agent merges or discards endmembers and re‑estimates abundances. Experiments on HYDICE Urban, Jasper Ridge, and Stonewall Playa datasets show consistent improvements in endmember cardinality and overall decomposition quality across multiple pipelines, while remaining competitive with end‑to‑end methods.
Hyperspectral image (HSI) classification systems are increasingly deployed on platforms with strict computational budgets, such as UAVs and small spaceborne sensors. In these settings, accuracy alone is not enough; the model must also run within tight latency and memory constraints.
arXiv:2609.13332v1 Announce Type: new Abstract: Automated landslide segmentation on Mars is one of the important tasks for understanding its surface processes, and all will aid in future space explor...
AGSA-Net is a hyperspectral image classification framework that incorporates spectral unmixing priors through an abundance-guided self‑attention network. It first estimates physically meaningful subpixel abundance maps with non‑negativity and sum‑to‑one constraints, then uses these abundances to build an affinity prior that directs a spectral transformer to focus on class‑discriminative interactions. The transformer features are fused with compact abundance descriptors for final classification, and experiments on Indian Pines, Augsburg, and Berlin datasets show improved performance, especially in heterogeneous urban scenes.