arXiv Computer Vision

MambaX: Image Super-Resolution with State Predictive Control

arXiv Computer Vision
Sep 7

Learning Spatial-Spectral Refinement and Calibrating Complementary Observations for Hyperspectral Image Super-Resolution

The paper introduces TSR-ITNR, a two‑stage, self‑supervised framework for hyperspectral image super‑resolution that fuses high‑resolution multispectral and low‑resolution hyperspectral data. Stage 1 refines an implicit Tucker representation using a low‑rank spatial tensor and spectral basis, enhanced by a pretrained denoiser, to capture fine spatial details and spectral correlations. Stage 2 applies parameter‑free calibration to extract complementary corrections from both observations, preserving geometry and ensuring orthogonal complementarity, leading to superior reconstruction quality demonstrated on benchmark datasets and improved downstream segmentation performance.

By Liqian Yang, Xingchi Chen, Xinfeng Gui, Xiangyong Cao, Qianxin Yi
arXiv Machine Learning
Aug 11

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

arXiv:2608. 08553v1 Announce Type: cross Abstract: Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration.

By Rong Fu, Chunlei Meng, Yangchen Zeng, Xiaowen Ma, Yongtai Liu, Wangyu Wu, Shuo Yin, Zijian Zhang, Sicheng Li, Yingrui Ji, Chenhao Wang, Simon Fong
arXiv Computer Vision
2d ago

SP-MoMamba: Superpixel-driven Mixture of State Space Experts for Efficient Image Super-Resolution

SP-MoMamba introduces a superpixel-driven mixture of state space experts for efficient image super‑resolution. By grouping spatially coherent features into region‑level tokens, the Superpixel‑SSM performs global sequence modeling over compact representations, reducing redundant computation while enabling long‑range structural interaction. The Multi‑Scale Superpixel Mixture of State Space Experts further adapts to varying representation granularities, and a Local Spatial Modulation Expert refines local high‑frequency details, resulting in strong reconstruction performance with a favorable trade‑off among model size, computational cost, and inference efficiency.

By Wenbin Zou, Yawen Cui, Yi Wang, Lap-Pui Chau, Liang Chen, Jinshan Pan, Huiping Zhuang, Guanbin Li
arXiv Computer Vision
Sep 14

RoES: Rotational Equivariant Selective-frequency Fusion for Multimodal Images

RoES is a Rotational Equivariant Selective-frequency fusion network that dynamically separates low- and high-frequency components of infrared-visible images. It uses a trainable rotation-enhanced updater to decouple frequencies, then fuses them with a dual-branch module: a rotation-equivariant Mamba for low-frequency structural dependencies and a polar spectral attention Dual-Fourier block for high-frequency detail refinement. Experiments show RoES outperforms existing methods in fusion quality and downstream object detection, offering a robust multimodal fusion solution.

By Jiabao Wang, Wenjian Liu, Yaoming Cai, Gengyu Zhang, Boyan Zhao, Zijia Zhang, Yao Ding, Xiaobo Liu
Hugging Face Trending Papers
Aug 9

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence.

arXiv AI
Jun 2

Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration

arXiv:2605. 00310v2 Announce Type: replace-cross Abstract: Super-resolution (SR) techniques have made major advances in reconstructing high-resolution images from low-resolution inputs.

By Zhili Li, Kangyang Chai, Zhihao Wang, Xiaowei Jia, Yanhua Li, Gengchen Mai, Sergii Skakun, Dinesh Manocha, Yiqun Xie
arXiv Machine Learning
Jul 23

Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination

arXiv:2607. 19719v1 Announce Type: new Abstract: Latent world models improve sample efficiency in continuous control by optimizing policies over imagined latent trajectories, but common neural transitions offer limited direct control over modal persistence and error accumulation in long rollouts.

By Jiaqi Li, Xinglong Zhang, Haibin Xie, Yixing Lan, Wei Pan, Xin Xu