The paper introduces RS-OPSD, a reliable privileged on-policy self-distillation framework designed for ultra‑high‑resolution remote sensing visual question answering. It leverages a new dataset, GeoEvidence‑6K, and a human‑feedback guided skill refinement process to provide explicit question‑relevant evidence. By incorporating context‑preserving visual privilege and correctness‑aligned distillation, RS‑OPSD achieves state‑of‑the‑art performance on several benchmarks without requiring additional visual search or tool calls at inference time.
By Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li
arXiv:2509. 23876v3 Announce Type: replace-cross Abstract: Autoregressive (AR) models based on next-scale prediction have emerged as a powerful tool for image generation, but they face a critical weakness: information inconsistencies between patches across timesteps introduced by progressive resolution scaling.
By Ky Dan Nguyen, Hoang Lam Tran, Anh-Dung Dinh, Daochang Liu, Weidong Cai, Xiuying Wang, Chang Xu
The paper introduces SGMA, a structure‑guided masked autoencoding framework designed for ultra‑high‑resolution scientific images. SGMA combines a content‑adaptive quadtree tokenizer that reduces gigapixel images to a fixed‑length sequence with a structure‑conditioned masking process that focuses reconstruction on spatially informative regions. The method, enhanced by Damped Accumulation to stabilize multi‑scale signals, achieves superior performance over standard MAE baselines on electron microscopy, whole‑slide optical microscopy, and X‑ray CT datasets, delivering significant accuracy gains and up to a 24.8× inference speedup.
By Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir Koushyar Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Masaru Tezuka, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed Wahib
arXiv:2608.30129v1 Announce Type: new
Abstract: This work presents $\textbf{Lapis}$, a $\textbf{l}$inear-$\textbf{a}$ttention-based $\textbf{pi}$xel-$\textbf{s}$pace generative framework that achieve...
By Bingde Liu, Wu Ran, Jinglei Zhang, Huanhuan Yuan, Chao Ma
arXiv:2603.16932v2 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) typically process images at a native high-resolution, forcing a trade-off between accuracy and computational ef...
By Nimrod Shabtay, Moshe Kimhi, Artem Spector, Sivan Haray, Ehud Rivlin, Chaim Baskin, Raja Giryes, Eli Schwartz
arXiv:2609.36929v1 Announce Type: new
Abstract: Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However,...
By Thai Duy Nguyen, Addison Lin Wang
arXiv:2608. 09133v1 Announce Type: cross Abstract: Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging.
By Yu Shi, Yuyao Zhang, Yu-wing Tai
arXiv:2608. 16546v1 Announce Type: cross Abstract: Most super-resolution models learn from paired data by supervising only the final high-resolution output.
By Zikang Zhan
The paper introduces the Logit Refiner, a lightweight autoregressive module that restores intra‑scale dependencies in Visual Autoregressive Models (VAR) by sequentially sampling tokens conditioned on frozen backbone features. This refiner adds only about 10% more parameters and less than 5% of the base model’s training compute, and can be applied to any pretrained VAR checkpoint without retraining. Experiments on ImageNet 256×256 show that the refiner consistently improves generation quality across backbones ranging from 310 M to 2 B parameters, enabling a 1.1 B‑parameter model to outperform a model twice its size, and the method generalizes to text‑to‑image generation, demonstrating that the mean‑field bottleneck is effectively alleviated.
By Meimingwei Li, Stefan Andreas Baumann, Felix Krause, Bj\"orn Ommer
arXiv:2609.38968v1 Announce Type: new
Abstract: Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-mo...
By Zeyu Wang, Mingyu Ge, Haiyu Song, Haoran Duan
Real-world image super-resolution (Real-ISR) aims to reconstruct high-quality (HQ) images from low-quality (LQ) inputs subject to diverse real-world degradations. Recent advances have leveraged the LQ inputs and natural image priors learned by Stable Diffusion models to achieve impressive results.
arXiv:2609.24510v2 Announce Type: replace
Abstract: Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs...
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu