arXiv AI

Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting

The paper proposes Metric-based Loss Weighting to enhance visual grounding in multimodal machine translation. By increasing loss for tokens that benefit from image context—identified via the Point-wise Cross-mutual Information (PCXMI) metric and its Congruency-based variant—the method improves translation accuracy on the CoMMuTE dataset by over 7 percentage points. Experiments fine-tune three pretrained multimodal LLMs across three language directions, showing superior performance compared to standard fine-tuning while preserving overall translation quality.

arXiv Computation and Language
Sep 18

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.

By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi
arXiv Computation and Language
Aug 28

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models introduces the VIG‑Sampler, a method that prioritizes tokens for decoding based on their attention to image tokens and penalizes redundancy in image‑attention distributions. The approach aims to improve the quality of multimodal generation by selecting more informative tokens during diffusion decoding. Experiments on seven captioning and VQA benchmarks with three open‑source dMLLMs show that VIG‑Sampler outperforms the Info‑Gain Sampler by an average of 19.3 CIDEr points and achieves better COCO Caption results using only half as many decoding steps.

By Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim
Hugging Face Trending Papers
Sep 8

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

The paper investigates image tokenizers as the visual language of unified multimodal models by creating a controlled autoregressive testbed that tracks task‑specific validation losses during multimodal continual pretraining across text, image, text‑to‑image, and image‑to‑text predictions. It shows that losses must be analyzed by task, that the loss–performance relationship varies with the token space, and that better reconstruction does not always lead to stronger downstream performance. The study also demonstrates how tokenizer design choices—such as discriminator use, semantic supervision, and vocabulary size—affect joint modeling and downstream results.

arXiv Computer Vision
Sep 3

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

The paper introduces TIC‑Bench, a new benchmark for evaluating multimodal large language models on deeply interleaved text‑image contexts. It covers logical, temporal, and spatial association tasks, totaling 2,280 questions across eight specific types. The authors benchmarked ten state‑of‑the‑art MLLMs, finding a significant performance gap versus human experts and highlighting persistent challenges in integrating evidence across interleaved visual and textual inputs.

By Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo
arXiv AI
Aug 28

VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation

The paper introduces Vision-Free Adaptation (VFA), a method that separates multilingual language enhancement from visual alignment in multimodal large language models. VFA fine‑tunes a base LLM on multilingual text to create a multilingual task vector, which is then merged with the vision‑aligned task vector of an existing MLLM. Experiments on five MLLMs and six multilingual benchmarks show consistent gains while preserving multimodal and text‑only performance, and using less than 2% of text data narrows the performance gap to fully multimodal‑trained models.

By Yixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang, Lingjie Jiang, Shaohan Huang, Yun Chen, Guanhua Chen, Furu Wei
arXiv Computation and Language
Aug 31

OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion

OmniFusion is an end‑to‑end multilingual multimodal translation system that fuses a pretrained multimodal foundation model (Omni 2.5‑7B) with a translation large language model (SeedX PPO‑7B). By connecting hidden states from multiple layers of the multimodal model to the translation LLM, OmniFusion can translate speech, speech‑and‑image, and text‑and‑image inputs while reducing simultaneous speech‑translation latency by about one second compared to cascaded pipelines. The approach improves overall translation quality by leveraging both audio and visual context.

By Sai Koneru, Matthias Huck, Jan Niehues
arXiv Computer Vision
Sep 24

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.

By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun