arXiv Computer Vision

RefGC-SR$^2$: Reference-guided Super-Resolution and Refinement of AI Generated Content

arXiv Computer Vision
6d ago

SR-Ground: Image Quality Grounding for Super-Resolved Content

SR‑Ground is a large‑scale dataset created to enable fine‑grained segmentation of visual artifacts in super‑resolved images. It contains 63,000 images processed by various state‑of‑the‑art SR models, each annotated at the pixel level for six distinct artifact types, validated through a crowdsourcing study with 1,062 participants. The dataset improves the training of image quality assessment models with grounding capabilities and supports a fine‑tuning pipeline that reduces perceptible artifacts in SR outputs, outperforming no‑reference methods on both benchmark and real‑world low‑resolution datasets.

By Artem Borisov, Evgeney Bogatyrev, Khaled Abud, Dmitriy Vatolin
arXiv Computer Vision
Sep 28

PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution

arXiv:2609.30988v1 Announce Type: new Abstract: Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while...

By Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao, Jiaming Guo, Renjing Pei, Xueyang Fu, Yang Cao, Zheng-Jun Zha
arXiv Computer Vision
Sep 3

SelfLift: Accelerating Few-Step Diffusion via Self-Recovering Resolution Transition

SelfLift is a progressive‑resolution framework that accelerates few‑step diffusion models by enabling late, self‑recovering transitions between low‑ and high‑resolution latents. It introduces a training‑free Artifact‑Aware Consistency Lift that uses disagreement between direct latent lifting and pixel‑VAE re‑encoding to detect and correct artifacts, and a self‑recovery policy that transfers high‑resolution guidance from an internal teacher. Experiments on FLUX.2‑Klein and Z‑Image‑Turbo show latency reductions of 41.5% and 44.1%, and overall speedups of 29.61× and 19.21× over 50‑step baselines while maintaining competitive generation quality.

By Tingyan Wen, Chenqian Yan, Xurui Peng, Xiazhang Fang, Shuai Wang, Xueqian Wang, Songwei Liu
arXiv AI
Jul 2

UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

arXiv:2511. 18050v1 Announce Type: cross Abstract: Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect ratios exposes a tightly coupled failure mode spanning positional encoding, VAE compression, and optimization.

By Tian Ye, Song Fei, Lei Zhu