Channel-wise Vector Quantization
arXiv:2605. 26089v2 Announce Type: replace-cross Abstract: We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens.
VTBench is a new benchmark that evaluates visual tokenizers (VTs) used in autoregressive image generation. It assesses VTs on image reconstruction, detail preservation, and text preservation across diverse scenarios, revealing that continuous VAEs outperform discrete VTs in maintaining spatial structure and semantic detail. The study also explores GPT‑4o’s potential autoregressive behavior and releases the benchmark publicly to encourage development of robust, open‑source VTs.
arXiv:2605. 26089v2 Announce Type: replace-cross Abstract: We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens.
arXiv:2608. 08676v1 Announce Type: cross Abstract: Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation.
VTR-Bench is a new benchmark designed to evaluate how well video generation models render text within scenes. It includes 300 prompts across five real-world scenarios such as advertisements and scientific videos, and uses an automated pipeline with human alignment to assess text fidelity and scene/motion requirements. Experiments on 11 state‑of‑the‑art models show that even the best performer has a word error rate of 0.250, underscoring widespread challenges in visual text rendering.
arXiv:2609.13245v1 Announce Type: new Abstract: Speculative Jacobi Decoding (SJD) is an important approach for accelerating autoregressive image generation. Although SJD has shown superior performanc...
The paper investigates how attention sparsity behaves in autoregressive image generation, finding a distinct diagonal sparsity pattern due to spatial locality of visual tokens. It introduces a diagonal‑aware sparse attention mechanism that skips KV entries along the diagonal within a recent window, achieving up to 3.1× higher throughput and 1.19× lower latency with less than 2% quality loss compared to dense inference.
arXiv:2512. 08240v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs.
The paper introduces RAE-CoD, a diffusion-based compression method that operates in a representation autoencoder space to preserve recognizable content even at extremely low bitrates. It addresses the problem of semantic collapse observed in existing codecs when the bitrate approaches zero, showing that reconstruction losses conflict with semantic objectives and that VAE diffusion models lose efficiency in preserving semantics. Experiments on MSCOCO-30K demonstrate that RAE-CoD outperforms competitors, reducing VFM feature MSE and Fréchet Distance ratios by at least 25.7% and 69.1% at 0.001–0.008 bpp while maintaining stable recognizability and quality.
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through a unified multimodal interface.
arXiv:2608. 04515v1 Announce Type: cross Abstract: Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices.
Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual...
arXiv:2607.11233v2 Announce Type: replace Abstract: Virtual try-on (VTON) is a bi-conditional image generation problem that requires not only accurate person preservation but also faithful garment de...
arXiv:2609.36756v1 Announce Type: cross Abstract: One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressiv...