arXiv AI By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Read the original on arXiv AI →

arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 24

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.

By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv AI
4d ago

VETO: Video Efficient Token Optimization for Vision Language Models

VETO (Video Efficient Token Optimization for Vision Language Models) is a plug‑in that reduces the quadratic cost of visual tokens in long‑video inference by applying dual‑axis compression: an intra‑frame compressor merges semantically similar tokens within each frame, and an inter‑frame compressor merges temporally redundant frames. By first compressing spatial dimensions, VETO lowers the cost of subsequent global temporal matching, surpassing single‑axis methods and achieving up to 45% faster inference on models such as LLaVA‑OneVision‑7B while maintaining or improving accuracy. The approach is universally applicable across LLaVA‑OneVision, InternVL‑2.5, and LongVA, preserving or enhancing zero‑shot accuracy even under extreme token budgets.

By Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen, Winston H. Hsu