Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv Computer Vision
Sep 22

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

arXiv:2609.24576v1 Announce Type: cross Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...

By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv Computer Vision
Sep 22

M3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in Forests

M3GA-Wild is a new benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests, combining synchronized RGB imagery and LiDAR from ground traversals with high‑resolution aerial imagery and multi‑altitude LiDAR over 370 hectares. The dataset includes accurate geo‑referenced 6‑DoF poses and spans 36 km of forest traversals, enabling systematic evaluation of visual, LiDAR, cross‑modal, and multi‑modal methods. Baseline experiments show LiDAR outperforms vision‑only approaches under severe viewpoint changes, while current multi‑modal fusion offers limited gains due to poor cross‑modal alignment, highlighting challenges in cross‑platform localisation and domain gaps.

By Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani
arXiv Computer Vision
Sep 22

BrainIAC: Interactive 3D Brain Lesion Segmentation across Heterogeneous MRI Modalities with Online Adaptation

BrainIAC is a unified framework for 3D brain lesion segmentation that handles heterogeneous MRI modalities and adapts online during interactive segmentation. It combines a multi‑modal backbone trained with zero‑filling and random modality dropping, 3D interactive prompts that default to fully automatic predictions, and a two‑stage online adaptation guided by pseudo‑labels and a Click‑Centered Gaussian loss. Experiments on seven MRI datasets show that the components work synergistically, outperforming existing methods and generalizing to unseen modalities and pathologies.

By Wentian Xu, Anthony P Addison, Ziyun Liang, Harry Anthony, Guang Yang, Konstantinos Kamnitsas
arXiv Computer Vision
Sep 22

Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning

The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.

By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
arXiv Computer Vision
Sep 22

DeCo: Efficient Decouple-to-Couple Learning for Multi-Task Visual Grounding

DeCo introduces an efficient decouple-to-couple learning framework for multi-task visual grounding, addressing conflicts between localization and segmentation tasks. It first applies Task-aware Semantic Decoupling (TSD) to separate shared visual cues into task-specific features guided by salient words, then uses Hybrid Prior Coupling (HPC) to merge sentence-level semantic priors with mask-derived spatial priors for improved grounding. Experiments across multiple natural and remote sensing datasets show that DeCo achieves state‑of‑the‑art performance while requiring only lightweight trainable parameters on a frozen multimodal encoder.

By Xiaoqiang Lu, Licheng Jiao, Long Sun, Yuting Yang, Xu Liu, Lingling Li, Wenping Ma, Fang Liu
arXiv Computer Vision
Sep 22

SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models

SPHQuant introduces a rotation‑free spherical weight‑only quantization framework for Vision‑Language Models, decomposing 8‑dimensional weight vectors into sign, radius, and a positive unit direction. By isolating outlier magnitudes in the radius and allocating extra precision there, it mitigates accuracy loss at extreme low bit‑widths. The method also employs a compact positive‑direction codebook with angular fine‑tuning and a hardware‑friendly GEMV kernel, achieving state‑of‑the‑art performance while boosting decode throughput by 30.3% on RTX A6000 compared to QTIP.

By Kewei Zhang, Zheng Chen, Haotong Qin, Yulun Zhang
arXiv Computer Vision
Sep 22

CIG-MAE: Cross-Modal Information-Guided Masked Autoencoder for Self-Supervised WiFi Sensing

CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.

By Gang Liu, Yanling Hao, Yixuan Zou
arXiv Computer Vision
Sep 22

CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification

The CXR‑LT 2026 Challenge introduces a multi‑center, long‑tailed chest X‑ray classification benchmark with over 145,000 radiologist‑annotated images from PadChest and NIH datasets. It defines two core tasks: robust multi‑label classification on 30 known classes and open‑world generalization to 6 unseen rare disease classes. The paper outlines data collection, annotation, solution strategies, and evaluates performance across head‑vs‑tail, calibration, and cross‑center gaps, noting that vision‑language models improve in‑distribution and zero‑shot performance but rare‑finding detection under multi‑center shift remains difficult.

By Hexin Dong, Yi Lin, Pengyu Zhou, Fengnian Zhao, Alan Clint Legasto, Juno Cho, Dohui Kim, Justin Namuk Kim, Mingeon Kim, Sunwoo Kwak, Gabriel Moy\`a-Alcover, Ky Trung Nguyen, Thanh-Huy Nguyen, Ha-Hieu Pham, Huy-Hieu Pham, Huy Le Pham, Nikhileswara Rao Sulake, Aina Tur-Serrano, Ruichi Zhang, Ang Zu, Adam E. Flanders, Zhiyong Lu, Ronald M. Summers, Mingquan Lin, Hao Chen, Yuzhe Yang, George Shih, Yifan Peng
arXiv Computer Vision
Sep 22

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

MinCU is a new benchmark for grounded minimal‑change understanding that presents pairs of near‑identical images differing by a single atomic variation in object category, attribute, count, or spatial position. Models are evaluated on their ability to describe the change, localize the changed region, and identify the changed entity. The authors also introduce SG‑ISA, a structured autoregressive method that decomposes the task into a Think‑Locate‑Describe sequence, showing that fine‑tuning with SG‑ISA improves both grounding accuracy and description quality while reducing reasoning‑token overhead.

By Chaoqian Mu, Wenhao Wu, Zichen Liang, Jiaxu Li, Lijun Wang, Yifan Wang, Huchuan Lu
arXiv Computer Vision
Sep 22

Preoperative Prediction of Microvascular Invasion in Hepatocellular Carcinoma by Integrating Multimodal Ultrasound and Clinical Data: A Multicenter Study

The study developed a multimodal ultrasound and clinical data model to predict microvascular invasion (MVI) preoperatively in hepatocellular carcinoma (HCC). Using data from 489 patients across eight centers, the model combined B-mode ultrasound, color Doppler flow imaging, dynamic contrast-enhanced ultrasound, and clinical information, achieving an AUC of 0.8953 in external validation. Dynamic contrast-enhanced ultrasound contributed the most predictive power, while other modalities and clinical data added complementary value.

By Jun Cheng, Yuanyuan Kong, Qing Huang, Xiaotong Tan, Licong Dong, Yulong Han, Wufeng Xue, Ruobing Huang, Dong Ni, Qi Yang, Jie Yu, Ping Liang
arXiv Computer Vision
Sep 22

LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models

LIBERO-VPro is a benchmark designed to assess the closed‑loop visual robustness of robotic foundation models by systematically perturbing visual inputs during task execution. It spans four dimensions—Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task‑Relevant Scene Variation—across 12 challenge categories, 96 settings, and 3,296 task‑condition cases. Evaluations on six models over 196,000 simulated episodes and 200 real‑world rollouts show that high nominal performance can hide significant weaknesses in visual grounding, adaptation, and sensitivity to stale or missing observations.

By Huiqiong Li, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Yu-Gang Jiang, Bin Zhu
arXiv Machine Learning
Sep 22

Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing

The study reports that agitation—an early, subtle sign of distress—can be detected in autistic youth using multimodal wearable sensors. By fusing upper‑body movement, wrist‑worn physiological data, and vocalizations through pretrained foundation models, the authors achieved an AUC of 0.724 at the clinician‑annotated onset of agitation, with audio contributing the most signal. The approach works across 15 participants, showing that individualized agitation is detectable even before the onset and without needing a separate model per child.

By Nibraas Khan, Abigale Plunk, John Staubitz, Ingrid Shragge, Jordan Brooks, Suzanne Wright, Alec Brewer, James Dieffenderfer, Alper Bozkurt, Amy Weitlauf, Nilanjan Sarkar
arXiv Machine Learning
Sep 22

A Federated Artificial Intelligence Framework for Optimizing Pancreatic Cancer Treatment - Strategy Update

arXiv:2609.24718v1 Announce Type: new Abstract: While a centralized approach involving patient consent to collect and analyze data centrally would theoretically offer the best data quality and predic...

By Anne-Christin Hauschild, Amirreza Aleyasin, Nils H. Beyer, Lisa Fricke, Jonas H\"ugel, Maryam Moradpour, Anh-Tien Nguyen, Youngjun Park, Sophia Rheinl\"ander, Tim Beissbarth, Elisabeth Hessmann, Martin Middeke, Matthias Lauth, Maximilian Reichert, Ulrich Sax