Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,568 stories · RSS feed

arXiv Computer Vision
2d ago

Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models

The paper investigates the gap between retrieval and reading in document vision‑language models, showing that even when the correct page is retrieved, the model often fails to use the text on that page. By comparing answers derived from page images alone versus images plus extracted OCR text, the authors demonstrate that adding OCR text can improve strict accuracy by 13 to 16 points on their FoveDoc‑Bench benchmark. The study also reveals that OCR benefits textual evidence but not charts or figures, and that its advantage diminishes as retrieval quality worsens.

By Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu
arXiv Computer Vision
2d ago

Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification

The paper evaluates how well generalist and dermatology-specific machine learning models perform on diverse skin lesion datasets, including dermoscopic images and smartphone photographs. It benchmarks a range of architectures—general-purpose vision-language models, foundation models, and task-specific dermatology classifiers—under conditions of distribution shift, modality change, and demographic variability. The study quantifies the performance gap between current state‑of‑the‑art models and the robustness needed for safe, equitable clinical deployment.

By Emanoel dos Santos, Kelvin Cunha, Rodrigo Mota, Fabio Papais, Thales Bezerra, Natalia Lopes, Erico Medeiros, Shirley Cruz, Jessica Araujo, Paulo Borba, Tsang Ing Ren
arXiv AI
2d ago

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

HyperBrowseComp is a multilingual and multimodal web‑browsing benchmark featuring 423 manually authored, human‑validated questions in 13 languages. The questions are intentionally difficult, requiring users to locate obscure evidence, follow multi‑step clue chains, or inspect heterogeneous sources such as videos, scanned documents, images, or maps. The benchmark filters out easier questions by testing models without internet access and evaluates performance using provider‑native search and a shared external retrieval harness, with a human evaluation on a sample to contextualize model effort.

By Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han, Ryandito Diandaru, Qinrong Cui, Jan Christian Blaise Cruz, Badrinath Chandana, Peerawat Chomphooyod, Ahmed Attia, Jonibek Mansurov, Emilio Villa-Cueva, Canh Duong Nguyen, Imran Turganov, Minghao Wu, Peerat Limkonchotiwat, Irina Nikishina
arXiv Machine Learning
2d ago

Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

The paper introduces SCOUT, an offline multi-agent reinforcement learning framework that combines a generative behavioral prior with a decomposed value function for test-time action refinement. SCOUT uses optimal unified transport and Stein variational gradient descent to steer behavioral samples toward high-value regions, with the number of transport steps providing adaptive scaling. The authors prove a single-term KL bound on the joint soft-value gap under the individual-global-max principle and demonstrate that SCOUT outperforms existing methods on both discrete and continuous offline MARL benchmarks, including offline-to-online settings.

By Dongsu Lee, Haoran Xu, Amy Zhang
arXiv Computation and Language
2d ago

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

The study investigates why multimodal large language models (VLMs) perform better at verifying scientific claims when evidence is presented as a table rather than a chart, despite both formats containing the same data. Using layer‑wise linear probing and attention analysis on three open‑weight VLMs, the authors find that chart information is indeed encoded in intermediate representations but never reaches the prediction layer, a gap absent for tables. Attention patterns reveal that this disconnect manifests differently across model families, suggesting the issue lies in how encoded visual data is utilized at prediction time rather than in the encoding process itself.

By Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, Florian Boudin, Akiko Aizawa
arXiv AI
2d ago

Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

The paper demonstrates that a targeted adversarial perturbation can reduce a vision‑language model’s training loss to near zero for a fixed target caption, yet the same model, when generating freely, still produces the correct description. This phenomenon, termed the train/inference gap, is traced to a single autoregressive step where the target token’s rank is fixed across all images, and further analysis shows that the language decoder, rather than the visual encoder, determines whether the corrupted signal is amplified or suppressed. The study uses a controlled two‑stage PGD attack on Qwen2.5‑VL‑7B‑Instruct and evaluates the effect on 200 held‑out COCO images, revealing that adversarial robustness in autoregressive VLMs largely depends on the language decoder’s prior. whyItMatters":"The findings suggest that defenses and faithfulness evaluations for deployed vision‑language models should focus on the language decoder rather than the visual encoder, as the former is the key determinant of robustness to adversarial perturbations."

By Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama
arXiv Computer Vision
2d ago

Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy

The paper introduces a unified framework for multimodal 3D human pose estimation that fuses RGB, LiDAR, and mmWave radar data while incorporating kinematics-based sensor fusion. It presents a black-box subject membership inference attack and a pointwise maximal leakage analysis to assess privacy risks, and proposes a user-level differential privacy method called Action Temporal Stratification to mitigate these risks. The framework is evaluated on the MM-Fi dataset under three experimental protocols, with source code to be released upon acceptance.

By Kaushik Bhargav Sivangi, Fani Deligianni
arXiv AI
2d ago

Multimodal Representation Learning Conditioned on Semantic Relations

Multimodal representation learning has largely relied on contrastive models like CLIP that produce a single embedding per sample, which limits their ability to capture relation-dependent relevance. The proposed Relation-Conditioned Multimodal Learning (RCML) framework explicitly conditions embeddings on natural‑language relation descriptions, enabling the same sample to be represented differently under various relational contexts. RCML builds relation‑aware training pairs, incorporates a relation‑conditioned module, and uses a unified contrastive objective to jointly model cross‑modal alignment and relation‑induced structure, achieving superior performance on retrieval and classification tasks across zero‑shot, fine‑tuned, and out‑of‑domain settings.

By Yang Qiao, Yuntong Hu, Bowen Zhu, Hasibul Haque, Liang Zhao
arXiv AI
2d ago

What do your logits know?

The paper investigates how much information can be extracted from different internal representations of vision‑language models, focusing on two bottlenecks: low‑dimensional projections of the residual stream (via tuned lenses) and the final top‑k logits. It systematically compares the amount of retained information at these levels and finds that even the easily accessible top‑logit bottleneck can leak task‑irrelevant details from image queries, sometimes matching the leakage seen from full residual projections.

By Masha Fedzechkina, Eleonora Gualdoni, Rita Ramos, Sinead Williamson
arXiv Computer Vision
2d ago

A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds

The paper introduces an automated pipeline that reconstructs editable 3D procedural models of field‑grown maize directly from raw 3D point clouds, eliminating the need for manual tuning or species‑specific training data. It uses a vision‑language model to annotate leaf midlines in rendered views, then applies deterministic geometric algorithms and differentiable NURBS fitting to generate accurate plant descriptors and refine leaf surfaces. The method achieves a median Chamfer distance of 5.4 mm on 100 diverse maize plants and recovers 99.4% of reference leaves with high overlap, outperforming previous semi‑automated approaches.

By Mozhgan Hadadi, Talukder Z. Jubery, Adarsh Krishnamurthy, Baskar Ganapathysubramanian
arXiv Computer Vision
2d ago

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT is a new training‑free metric for evaluating text‑to‑image alignment. It replaces fixed reference answers with an agreement rule that compares image‑based and caption‑only responses, weighting questions by how decisively the caption determines them. By merging this agreement score with a holistic score, DEPICT improves negation accuracy dramatically and outperforms existing training‑free metrics while matching or exceeding fine‑tuned evaluators on several benchmarks.

By Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins
arXiv Computation and Language
2d ago

How Robust Is Multimodal Claim Verification to LLM Rewriting?

The paper investigates how stylistic changes introduced by large language models (LLMs) affect multimodal claim verification, a task that determines whether a textual claim is supported by given evidence. Two rewriting strategies are used: natural rewriting, mimicking typical academic polishing, and controlled injection, adding a single LLM-associated word. Across 11 open‑weight models (2B–38B parameters) from five VLM families, the study finds that most models remain robust to these modifications, showing no significant accuracy drop, though consistent probability shifts—especially under hedging conditions—are observed.

By Yun-Ang Wu, Xanh Ho, Andre Greiner-Petter, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Akiko Aizawa
arXiv Computation and Language
2d ago

Evaluating VQA in Vision Language Models using Cooperative Principles

The paper evaluates Vision Language Models (VLMs) on Visual Question Answering tasks where questions violate Grice's maxims. By generating question modifiers that add non-essential, ambiguous, or false information, the authors show that VLMs such as ChatGPT, Claude, Gemini, and Llava exhibit reduced performance. They also compare human pragmatic reasoning to VLM reasoning, noting differences in how each handles human‑induced versus AI‑generated violations, and find that humans spend less time resolving VLM‑induced violations while VLMs are less accurate in those cases.

By Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal
arXiv AI
2d ago

RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction

RASPER is a reward‑aligned summarizer that tailors the extraction of information from unstructured discharge notes to improve downstream clinical predictions. It uses a tunable LLM summarizer trained with reinforcement learning, where the reward comes from the loss of a downstream predictor, and incorporates patient‑specific context via a longitudinal encoder that soft‑prompts the summarizer with structured codes. The approach consistently outperforms strong baselines on readmission prediction and medication recommendation tasks in the MIMIC‑III and MIMIC‑IV datasets.

By Arya Hadizadeh Moghaddam, Mohsen Nayebi Kerdabadi, Chen Chen, Dongjie Wang, Zijun Yao
arXiv AI
2d ago

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

This paper introduces a personalized automatic speech recognition system for a Czech speaker with a permanent tracheal stoma and severe dysarthria. The authors release a 33‑hour annotated dataset collected via an artificial conversation protocol and develop a multi‑stage training pipeline based on Whisper Base, fine‑tuning on Czech speech, simulated tracheostomic speech, and the speaker’s data. Evaluations in scripted, question‑answering, and spontaneous dialogue scenarios show a 50 % relative reduction in character error rate compared to the Whisper Base baseline and better accuracy than the speaker’s assistants on isolated utterances.

By David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus
arXiv Computation and Language
2d ago

An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

The paper introduces an automated pipeline that extracts conversational turns and backchannels from separate-channel recordings of spontaneous dyadic dialogue, combining voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post‑processing. Evaluated on 99 ten‑minute Danish conversations, the system achieved F1 scores around 0.62 for both turns and backchannels, with median onset/offset errors of roughly 0.15–0.18 s. Performance was consistent across normal and asymmetric listening conditions, and a case study showed the pipeline’s outputs were less variable than human annotations, supporting its use as a reliable first‑pass annotation tool in semi‑automated workflows.

By Hanlu He, Harald Vilhelm Skat-R{\o}rdam, Ingvi \"Orn\'olfsson, Ivana Konvalinka
arXiv Computation and Language
2d ago

Social bot detection in the age of ChatGPT: Challenges and opportunities

The article reviews the challenges and opportunities of detecting social bots amid the rise of advanced AI chatbots. It highlights gaps in current detection methods, especially regarding AI-generated conversations, and identifies emerging trends such as synthetic data generation, multimodal cross‑platform detection, low‑resource language support, and federated learning approaches. The authors propose these directions as promising avenues for future research.

By Emilio Ferrara
arXiv Computation and Language
2d ago

World Embedding Benchmark

The World Embedding Benchmark introduces 8,000 simulation-based video cases covering fluid, solid, dynamic, and optical physics, each paired with physical annotations. It supports three tasks—text‑video retrieval, physical‑property regression, and multiple‑choice classification—to assess how well video embeddings capture physical alignment versus quantitative information. Experiments show that while pre‑trained models perform poorly on retrieval and classification, lightweight probes can extract useful physical data, and physics‑specific contrastive training improves alignment but harms regression, highlighting a trade‑off. Retrieval‑augmented generation using these embeddings further enhances the physical fidelity of generated videos.

By Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li, Fengyu Cai, Hou Pong Chan, Jialin Yu, Hao Zhang, Chenghua Lin, Chenghao Xiao
arXiv Computation and Language
2d ago

BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions

BioMol-MQA is a new question‑answering dataset focused on polypharmacy that combines a multimodal knowledge graph—containing both text and molecular structure—with challenging questions designed to test large language models’ ability to retrieve and reason over this diverse information. The dataset highlights the limitations of current retrieval‑augmented generation systems, which typically handle only single‑modality text, by demonstrating that existing LLMs perform poorly unless provided with the necessary multimodal background data. This underscores the need for more robust RAG frameworks capable of integrating multiple data types for accurate responses.

By Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu, Fali Wang, Suhang Wang
arXiv Computation and Language
2d ago

The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?

The paper introduces Percept-V, a dataset of 6,000 program-generated images across 30 domains designed to test simple visual perception skills from the TVPS-4 framework. Experiments show that state‑of‑the‑art multimodal large language models perform poorly compared to humans, especially as image complexity increases, and that fine‑tuning yields only limited generalization to related datasets. The study highlights specific perception skills that remain challenging for current models.

By Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla