arXiv AI

Infinite-Precision Autoregressive Modeling for Vector Graphics and Layouts

arXiv:2601. 05680v2 Announce Type: replace-cross Abstract: While Transformer-based autoregressive models excel in data generation, their token discretization strategy inherently limits their precision in continuous domains.

arXiv Machine Learning
Sep 3

A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization

The paper introduces a unified rate–distortion framework for discrete visual tokenization, encompassing vector, product, and scalar quantization. It shows that minimizing distortion, rather than maximizing codebook utilization, is the key objective for reconstruction fidelity and establishes fairness conditions for comparing quantizers. Under these conditions, the study confirms the distortion hierarchy VQ–PQ–SQ and demonstrates that modern VQ methods achieve the lowest distortion.

By Xianghong Fang, Wenlong Mou, Yuan Yuan, Dehan Kong, Tim G. J. Rudner
arXiv Machine Learning
Sep 11

Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

The paper introduces the Logit Refiner, a lightweight autoregressive module that restores intra‑scale dependencies in Visual Autoregressive Models (VAR) by sequentially sampling tokens conditioned on frozen backbone features. This refiner adds only about 10% more parameters and less than 5% of the base model’s training compute, and can be applied to any pretrained VAR checkpoint without retraining. Experiments on ImageNet 256×256 show that the refiner consistently improves generation quality across backbones ranging from 310 M to 2 B parameters, enabling a 1.1 B‑parameter model to outperform a model twice its size, and the method generalizes to text‑to‑image generation, demonstrating that the mean‑field bottleneck is effectively alleviated.

By Meimingwei Li, Stefan Andreas Baumann, Felix Krause, Bj\"orn Ommer