arXiv Computer Vision

SJD-SV: Speculative Jacobi Decoding with Semantics Verification for Autoregressive Image Generation

arXiv Computer Vision
Sep 25

MoVISA: Multi-Token Reasoning for Video Object Segmentation

MoVISA introduces Multi-Token Reasoning for Video Object Segmentation, using multiple segmentation tokens (e.g., SEG0, SEG1) instead of a single token to better localize multiple objects over time. This approach enhances fine-grained alignment between language prompts and spatio-temporal mask predictions, leading to improved performance and interpretability. On benchmarks such as MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS, MoVISA achieves significant gains, notably a 13.2% J and F improvement on MeViS and an 8.4% J and F improvement on ReVOS.

By Ruining Zhao, Ho Kei Cheng, Alexander G Schwing
arXiv Computer Vision
Aug 25

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

arXiv:2501.04001v4 Announce Type: replace Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...

By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
arXiv Computer Vision
Aug 27

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

The paper introduces MTAR, a training framework for autoregressive image generation that enhances performance through multi-token prediction, token-level contrastive regularization, and semantic dropping. These components address sparse supervision, improve representation discriminability, and accelerate training without affecting inference. On ImageNet, MTAR outperforms LlamaGen with lower FID and faster training, achieving comparable results in only a third of the iterations.

By Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu
arXiv AI
Sep 3

RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models

The paper introduces RVSD, a training‑free, plug‑and‑play decoding framework that combines token sparsification with Semantic‑Space Visual Retrieval (SSVR) to reduce visual hallucinations in large vision‑language models. RVSD employs a semantics‑directed token selection strategy to prune redundant tokens while preserving essential visual information, and uses SSVR to perform on‑demand cross‑modal retrieval within a shared semantic space. Experiments show that RVSD achieves state‑of‑the‑art performance in mitigating visual hallucinations while maintaining strong suppression in long‑context generation.

By Canjie Liu, Jiawen Kang, Jinbo Wen, Zishao Zhong
arXiv Computer Vision
Sep 17

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.

By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
arXiv Computer Vision
Aug 26

AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer

arXiv:2608.23864v1 Announce Type: new Abstract: Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be or...

By Junqiu Yu, Pandeng Li, Yikai Wang, Jiaxing Zhao, Yujie Wei, Kaixun Jiang, Quanhao Li, Hongtao Yu, Zhihang Liu, Zhaohe Liao, Junjie Zhou, Yun Zheng, Yu Liu, Yanwei Fu