arXiv:2605.09687v2 Announce Type: replace
Abstract: Remote sensing single-image super-resolution aims to generate high-resolution imagery from low-resolution observations while preserving fine struct...
By Md Aminur Hossain, Ayush V. Patel, Sanjay K. Singh, Yogesh Jethani, Biplab Banerjee
The paper introduces SPARK, a lightweight input‑conditioned controller that modulates only a few dominant channels in frozen Diffusion Transformer (DiT) based super‑resolution models. By predicting bounded per‑channel affine transformations for selected channels, SPARK improves reconstruction fidelity and perceptual quality without fine‑tuning the backbone or adding adapters. Experiments on three DiT‑based SR backbones across DIV2K, RealSR, and DRealSR demonstrate consistent gains while modulating only eight channels per stream and block.
By Federico Putamorsi, Leonardo Zini, Marcella Cornia, Lorenzo Baraldi
arXiv:2505.16157v3 Announce Type: replace
Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Tran...
By Yuang Ai
arXiv:2608. 20263v1 Announce Type: new Abstract: We propose UHDformer++, a general Transformer-based framework to solve numerous Ultra-High-Definition (UHD) image restoration tasks.
By Cong Wang, Liyan Wang, Jinshan Pan, Wei Wang, Wenqi Ren, Jun Liu, Xiaochun Cao
arXiv:2601.17723v3 Announce Type: replace
Abstract: Implicit neural representation (INR) has become the standard approach for arbitrary-scale image super-resolution (ASSR). However, no systematic emp...
By Tayyab Nasir, Daochang Liu, Ajmal Mian
arXiv:2606. 19617v1 Announce Type: cross Abstract: We present GB-LSR (Global-Bandwidth Local Spectral Representation), a fixed-grid local spectral representation for continuous image reconstruction.
By Max Shad, Naeem Khoshnevis
We present GB-LSR (Global-Bandwidth Local Spectral Representation), a fixed-grid local spectral representation for continuous image reconstruction. The image domain is partitioned into non-overlapping square patches, each carrying coefficients for a truncated Fourier basis predicted from shared convolutional-encoder features.
arXiv:2608.30782v1 Announce Type: new
Abstract: Real-world image super-resolution (Real-ISR) aims to preserve structures supported by the degraded observation while reconstructing perceptually realis...
By Bingtian Qiao, Yue Shi, Yong Guo, Wenjun Zhang, Jiezhang Cao
arXiv:2606. 02092v1 Announce Type: cross Abstract: Semantic segmentation of remote sensing imagery requires models that capture both global context and local detail under tight computational budgets.
By \"Umit Mert \c{C}a\u{g}lar, Alptekin Temizel
arXiv:2607. 03612v1 Announce Type: cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success.
By Jianing Deng, Yuanzhe Li, Jialu Wang, Song Wang, Tianlong Chen, Huanrui Yang, Jingtong Hu
ProgResViT is an input‑adaptive Vision Transformer that processes images progressively across multiple rounds, starting with a low‑resolution image and a narrow subnetwork and refining the prediction with higher resolution and a wider subnetwork if needed. The method introduces Progress‑Conditioned Soft Gating (PSG) to share a single backbone across rounds while conditioning token fusion and layer outputs on the current round, block, and input resolution. Experiments on DeiT show improved accuracy‑compute trade‑offs compared to adaptive‑width, adaptive‑depth, and dynamic‑token baselines, and the design also benefits self‑supervised DINO representations and downstream semantic segmentation.
By Ali Hojjat, Janek Haberer, Olaf Landsiedel
Semantic segmentation of remote sensing imagery requires models that capture both global context and local detail under tight computational budgets. Prior work typically optimizes for one of these axes: attention for global context, convolution for local detail, or compactness for efficiency.