arXiv:2609.36756v1 Announce Type: cross
Abstract: One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressiv...
By Jiawei Zhang, Shuhao Liu, Rong Huang, Yuancheng Li, Zhihui Li, Xiaojun Chang, Changlin Li
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation qual...
arXiv:2609.13245v1 Announce Type: new
Abstract: Speculative Jacobi Decoding (SJD) is an important approach for accelerating autoregressive image generation. Although SJD has shown superior performanc...
By Baoquan Zhang, Bingqi Shan, Shihao Fang, Kenghong Lin, Xutao Li, Yunming Ye
Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained model, converting pixel values into token sequences that the LLM processes through its vocabulary head. This design shows that pretrained language models can provide probability estimates for image coding, but it also couples compression to tokenizer behavior, vocabulary-specific numeric tokens, and model-family-specific adaptation.
In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compression ratio is often accompanied by the loss of critical details.
arXiv:2509. 12159v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models have demonstrated exceptional performance in UI2Code tasks, significantly enhancing website development efficiency.
By Jingyu Xiao, Zhongyi Zhang, Yuxuan Wan, Yintong Huo, Yang Liu, Michael R. Lyu
arXiv:2505. 18227v4 Announce Type: replace-cross Abstract: In Transformer architectures, tokens\textemdash discrete units derived from raw data\textemdash are formed by segmenting inputs into fixed-length chunks.
By Zhenglun Kong, Yize Li, Fanhu Zeng, Lei Xin, Shvat Messica, Xue Lin, Pu Zhao, Manolis Kellis, Hao Tang, Marinka Zitnik
arXiv:2607. 00371v1 Announce Type: cross Abstract: Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation.
By Nuoyan Zhou, Zhijun Tu, Lei Yu, Kun Cheng, Jie Hu, Nannan Wang, Xinghao Chen
arXiv:2512. 08240v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs.
By Jusheng Zhang, Xiaoyang Guo, Tongyu Mo, Qinhan Lv, Wenhao Chai, Jian Wang, Keze Wang, Liang Lin
The paper introduces MTAR, a training framework for autoregressive image generation that enhances performance through multi-token prediction, token-level contrastive regularization, and semantic dropping. These components address sparse supervision, improve representation discriminability, and accelerate training without affecting inference. On ImageNet, MTAR outperforms LlamaGen with lower FID and faster training, achieving comparable results in only a third of the iterations.
By Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu
arXiv:2609.35232v2 Announce Type: replace-cross
Abstract: Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pr...
By Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo
arXiv:2605. 26089v2 Announce Type: replace-cross Abstract: We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens.
By Wei Song, Tianhang Wang, Yitong Chen, Tong Zhang, Zuxuan Wu, Min Li, Jiaqi Wang, Kaicheng Yu