OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces RS-OPSD, a reliable privileged on-policy self-distillation framework designed for ultra‑high‑resolution remote sensing visual question answering. It leverages a new dataset, GeoEvidence‑6K, and a human‑feedback guided skill refinement process to provide explicit question‑relevant evidence. By incorporating context‑preserving visual privilege and correctness‑aligned distillation, RS‑OPSD achieves state‑of‑the‑art performance on several benchmarks without requiring additional visual search or tool calls at inference time.
arXiv:2509. 23876v3 Announce Type: replace-cross Abstract: Autoregressive (AR) models based on next-scale prediction have emerged as a powerful tool for image generation, but they face a critical weakness: information inconsistencies between patches across timesteps introduced by progressive resolution scaling.
The paper introduces SGMA, a structure‑guided masked autoencoding framework designed for ultra‑high‑resolution scientific images. SGMA combines a content‑adaptive quadtree tokenizer that reduces gigapixel images to a fixed‑length sequence with a structure‑conditioned masking process that focuses reconstruction on spatially informative regions. The method, enhanced by Damped Accumulation to stabilize multi‑scale signals, achieves superior performance over standard MAE baselines on electron microscopy, whole‑slide optical microscopy, and X‑ray CT datasets, delivering significant accuracy gains and up to a 24.8× inference speedup.
arXiv:2608.30129v1 Announce Type: new Abstract: This work presents $\textbf{Lapis}$, a $\textbf{l}$inear-$\textbf{a}$ttention-based $\textbf{pi}$xel-$\textbf{s}$pace generative framework that achieve...
arXiv:2603.16932v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) typically process images at a native high-resolution, forcing a trade-off between accuracy and computational ef...
arXiv:2609.36929v1 Announce Type: new Abstract: Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However,...