The paper introduces TSR-ITNR, a two‑stage, self‑supervised framework for hyperspectral image super‑resolution that fuses high‑resolution multispectral and low‑resolution hyperspectral data. Stage 1 refines an implicit Tucker representation using a low‑rank spatial tensor and spectral basis, enhanced by a pretrained denoiser, to capture fine spatial details and spectral correlations. Stage 2 applies parameter‑free calibration to extract complementary corrections from both observations, preserving geometry and ensuring orthogonal complementarity, leading to superior reconstruction quality demonstrated on benchmark datasets and improved downstream segmentation performance.
By Liqian Yang, Xingchi Chen, Xinfeng Gui, Xiangyong Cao, Qianxin Yi
HyperVision introduces the first ground‑based hyperspectral pre‑trained backbone, addressing challenges of varying spectral configurations, limited annotations, and dataset diversity. It employs a channel‑adaptive dynamic embedding to unify heterogeneous inputs, a multi‑source pseudo‑labeling strategy combining SAM2 spatial cues with HyperFree spectral details, and cross‑modal knowledge distillation from a pre‑trained RGB vision model. Trained on 15k images from 26 datasets, HyperVision achieves significant improvements—up to 16.3% relative gain in hyperspectral semantic segmentation, 2.1% in object tracking AUC, and 35.5% reduction in salient object detection MAE—while requiring only head‑only adaptation.
By Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Xiangyu Liu, Fengchao Xiong, Jianfeng Lu, Chengrong Chen, Jun Zhou
arXiv:2609.01060v1 Announce Type: cross
Abstract: Compact snapshot hyperspectral cameras provide rich instantaneous spectral measurements for ground-level machine vision, but at lower spatial resolut...
By Mohamad Jouni, Aur\'elien Godet, Mauro Dalla Mura
The fusion of hyperspectral (HS) and Light Detection and Ranging (LiDAR) data plays a crucial role in enhancing land-cover classification by jointly exploiting spectral, spatial, and structural cues....
The paper presents M2Heat, a physics-inspired framework for fusing hyperspectral and LiDAR data in land‑cover classification. It introduces a visual heat conduction module (vHeat) and Frequency Value Embeddings (FVEs) to model anisotropic information flow, capturing global dependencies with sub‑quadratic complexity. Combined with a Cross‑Frequency Fusion (CFF) module, M2Heat delivers discriminative, robust features and achieves competitive performance on Trento, Houston2013, and Augsburg benchmarks while offering an interpretable heat‑conduction perspective.
By Kan Wei, Jiahui Cui, Jing Yao, Xinyu Zhao, Lei Wang, Pedram Ghamisi
arXiv:2609.13654v1 Announce Type: new
Abstract: Pretrained foundation models (FMs) have achieved remarkable success in computer vision, yet their high fine-tuning cost limits practical deployment. Pa...
By Han Luo, Ruoyu Yang, Yinhe Liu, Yanfei Zhong
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
The paper introduces SGFNet, a Semantic‑Guided Fusion Network for classifying multi‑source remote sensing images. It features a Semantic Mixing Convolution Block that generates semantic‑aware kernels based on contextual relationships, and a Frequency Modulated Fusion Block that fuses cross‑modal information in the frequency domain to mitigate spatial misalignment. Experiments on the Augsburg and Houston 2018 datasets show SGFNet consistently outperforms state‑of‑the‑art methods.
By Yuwei Zhao, Chuanzheng Gong, Baogui Huan, Feng Gao, Junyu Dong, Qian Du
arXiv:2510. 02308v2 Announce Type: replace Abstract: Estimating the tangent spaces of a data manifold is a fundamental problem in geometric data analysis.
By Dhruv Kohli, Sawyer J. Robertson, Gal Mishne, Alexander Cloninger
arXiv:2606. 04493v1 Announce Type: cross Abstract: Correspondence pruning aims to identify inliers from an initial set of correspondences.
By Zhihua Wang, Yanping Li, Yizhang Liu
RoES is a Rotational Equivariant Selective-frequency fusion network that dynamically separates low- and high-frequency components of infrared-visible images. It uses a trainable rotation-enhanced updater to decouple frequencies, then fuses them with a dual-branch module: a rotation-equivariant Mamba for low-frequency structural dependencies and a polar spectral attention Dual-Fourier block for high-frequency detail refinement. Experiments show RoES outperforms existing methods in fusion quality and downstream object detection, offering a robust multimodal fusion solution.
By Jiabao Wang, Wenjian Liu, Yaoming Cai, Gengyu Zhang, Boyan Zhao, Zijia Zhang, Yao Ding, Xiaobo Liu
arXiv:2608.29220v1 Announce Type: new
Abstract: Multimodal image fusion (MMIF) aims to integrate complementary sensor data into a single representation that preserves intrinsic scene reality while el...
By Haozhen Wei, Chengjun Jiang, Yutong Guo, Xinrui Ju, Xingyuan Li, Xiang Chen, Jinyuan Liu