arXiv AI By Chicago Y. Park, Jialin Mao, Xiaojian Xu, Taha Kass-Hout, Ulugbek S. Kamilov, Cao Xiao

Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling

Read the original on arXiv AI →

arXiv:2607. 09892v1 Announce Type: cross Abstract: We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 11

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.

By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
arXiv Computer Vision
Sep 23

MIAR: Medical Image Super-Resolution With Autoregressive Modeling

MIAR introduces a multi‑scale autoregressive framework for medical image super‑resolution, treating the task as a conditional, progressive next‑scale prediction. It incorporates a Scale‑Adaptive Structural Decoder to preserve structural fidelity and uses a hierarchical beam search during inference to reduce recursive error accumulation. Experiments show MIAR outperforms existing methods, achieving a 7.86% MUSIQ improvement and a 2.02× speedup over diffusion‑based approaches.

By Fang Li, Yinglong Li, Hongyu Wu, Yang Gao, Minwei Zhao, Aimin Hao
arXiv Computer Vision
Aug 27

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

The paper introduces MTAR, a training framework for autoregressive image generation that enhances performance through multi-token prediction, token-level contrastive regularization, and semantic dropping. These components address sparse supervision, improve representation discriminability, and accelerate training without affecting inference. On ImageNet, MTAR outperforms LlamaGen with lower FID and faster training, achieving comparable results in only a third of the iterations.

By Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu
arXiv Computer Vision
Aug 27

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

ARGenSeg introduces an autoregressive generation-based approach for image segmentation that integrates seamlessly with multimodal large language models (MLLMs). Unlike prior methods that use boundary points or dedicated segmentation heads, ARGenSeg generates dense masks directly through visual token output and detokenization via a universal VQ‑VAE, enabling fine‑grained pixel‑level perception. The framework employs a next‑scale‑prediction strategy to parallelize token generation, resulting in faster inference while outperforming state‑of‑the‑art segmentation models on multiple datasets.

By Xiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji, Dandan Zheng, Jingdong Chen, Jun Zhou