arXiv Machine Learning By Meimingwei Li, Stefan Andreas Baumann, Felix Krause, Bj\"orn Ommer

Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

Read the original on arXiv Machine Learning →

The paper introduces the Logit Refiner, a lightweight autoregressive module that restores intra‑scale dependencies in Visual Autoregressive Models (VAR) by sequentially sampling tokens conditioned on frozen backbone features. This refiner adds only about 10% more parameters and less than 5% of the base model’s training compute, and can be applied to any pretrained VAR checkpoint without retraining. Experiments on ImageNet 256×256 show that the refiner consistently improves generation quality across backbones ranging from 310 M to 2 B parameters, enabling a 1.1 B‑parameter model to outperform a model twice its size, and the method generalizes to text‑to‑image generation, demonstrating that the mean‑field bottleneck is effectively alleviated.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 25

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.

By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
arXiv Computer Vision
2d ago

SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

arXiv:2609.37969v1 Announce Type: new Abstract: High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first genera...

By Haozhe Liu, Tian Ye, Shuchen Xue, Yitong Li, Junsong Chen, Haopeng Li, Jincheng Yu, Duomin Wang, Ruihua Zhang, Lei Zhu, Song Han, Enze Xie