arXiv AI

OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution

arXiv Computer Vision
2d ago

RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

The paper introduces RS-OPSD, a reliable privileged on-policy self-distillation framework designed for ultra‑high‑resolution remote sensing visual question answering. It leverages a new dataset, GeoEvidence‑6K, and a human‑feedback guided skill refinement process to provide explicit question‑relevant evidence. By incorporating context‑preserving visual privilege and correctness‑aligned distillation, RS‑OPSD achieves state‑of‑the‑art performance on several benchmarks without requiring additional visual search or tool calls at inference time.

By Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li
arXiv AI
Jul 8

Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

arXiv:2509. 23876v3 Announce Type: replace-cross Abstract: Autoregressive (AR) models based on next-scale prediction have emerged as a powerful tool for image generation, but they face a critical weakness: information inconsistencies between patches across timesteps introduced by progressive resolution scaling.

By Ky Dan Nguyen, Hoang Lam Tran, Anh-Dung Dinh, Daochang Liu, Weidong Cai, Xiuying Wang, Chang Xu
arXiv Computer Vision
4d ago

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

The paper introduces SGMA, a structure‑guided masked autoencoding framework designed for ultra‑high‑resolution scientific images. SGMA combines a content‑adaptive quadtree tokenizer that reduces gigapixel images to a fixed‑length sequence with a structure‑conditioned masking process that focuses reconstruction on spatially informative regions. The method, enhanced by Damped Accumulation to stabilize multi‑scale signals, achieves superior performance over standard MAE baselines on electron microscopy, whole‑slide optical microscopy, and X‑ray CT datasets, delivering significant accuracy gains and up to a 24.8× inference speedup.

By Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir Koushyar Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Masaru Tezuka, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed Wahib
arXiv Machine Learning
Sep 11

Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

The paper introduces the Logit Refiner, a lightweight autoregressive module that restores intra‑scale dependencies in Visual Autoregressive Models (VAR) by sequentially sampling tokens conditioned on frozen backbone features. This refiner adds only about 10% more parameters and less than 5% of the base model’s training compute, and can be applied to any pretrained VAR checkpoint without retraining. Experiments on ImageNet 256×256 show that the refiner consistently improves generation quality across backbones ranging from 310 M to 2 B parameters, enabling a 1.1 B‑parameter model to outperform a model twice its size, and the method generalizes to text‑to‑image generation, demonstrating that the mean‑field bottleneck is effectively alleviated.

By Meimingwei Li, Stefan Andreas Baumann, Felix Krause, Bj\"orn Ommer