Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv AI
Sep 23

Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information

The study introduces a multimodal framework that predicts postprandial glycemic response (PPGR) by combining image-derived macronutrient estimates with clinical variables and gut microbiome data. It jointly performs macronutrient estimation from meal images and glucose prediction, using an attention-based module to model interactions between dietary and host-specific information. Evaluated on a real-world dataset, the model outperforms existing PPGR baselines that use image-derived inputs and nearly matches methods relying on manually reported macronutrients.

By Varvara Kondratyeva, Kamilia Zaripova, Nassir Navab, Azade Farshad
arXiv AI
Sep 23

LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot

The paper introduces an LLM-based Conversational AI Knowledge Assistant for the Raspberry‑Pi‑powered 13‑Axis MyBuddy humanoid robot. It combines large language model-driven language understanding, real‑time speech recognition, internet‑based knowledge retrieval (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to support intelligent, multi‑turn conversations and emotional‑support interactions. This system aims to overcome the limitations of traditional rule‑based dialogue systems in humanoid robots.

By Hanxiao Chen
arXiv AI
Sep 23

Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanation: A Case Study in Decision Support for Rice Cultivation in Japan

arXiv:2512. 21066v4 Announce Type: replace Abstract: Explainable artificial intelligence (XAI) reveals how explanatory variables relate to a response variable, yet communicating XAI outputs to laypersons remains difficult, limiting trust in AI-based predictions.

By Tomoaki Yamaguchi, Yutong Zhou, Masahiro Ryo, Keisuke Katsura
arXiv AI
Sep 23

ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry

ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.

By Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su
arXiv Machine Learning
Sep 23

FREESIA: Covariance-Aware Posterior Transport for Expressive and Scalable Data Assimilation

The paper introduces FREESIA, a training‑free, covariance‑aware posterior transport method for data assimilation that embeds forecast cross‑covariance into a flow‑based transport to recover unobserved states while preserving non‑Gaussian posterior structure. It combines an observation‑adaptive proposal with posterior correction, providing an asymptotically exact approximation of nonlinear posteriors and a Wasserstein error bound. Experiments on Double‑Well, Lorenz‑96, and Kolmogorov flow demonstrate that FREESIA captures complex posterior structures and achieves up to a 56% reduction in RMSE compared to the best baseline in sparse, nonlinear, non‑injective observation scenarios.

By Shiwei Ni, Yangwen Zhang, Hang Qi, Xiaofei Guan, Lili Ju
arXiv Machine Learning
Sep 23

Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models

The paper investigates how different action tokenization methods affect closed‑loop control in autoregressive vision‑language‑action models. It compares analytical, linear, and nonlinear representations, showing that lower reconstruction error does not guarantee better policy performance. The study highlights the need to evaluate tokenization on multiple criteria, including sequence predictability and decoder stability, rather than relying solely on reconstruction fidelity.

By Yuxin Yang, Gaohan He, Changxue Guan, Hangming Liu
arXiv Machine Learning
Sep 23

Foundation model embeddings capture pre-diagnostic changes on screening mammograms

The study examined whether embeddings from four foundation models—Mammo-CLIP, HOPPR, MedImageInsight, and BiomedCLIP—could detect pre‑diagnostic changes in screening mammograms. Using 1,773 biopsied women and matched controls, the researchers measured the speed of movement along a data‑derived “cancer direction” in embedding space over successive screening intervals. They found that embeddings from clinically grounded models (Mammo‑CLIP, HOPPR, MedImageInsight) showed faster drift in malignant cases compared to controls, while the general biomedical model BiomedCLIP did not, indicating that foundation model embeddings can encode early tissue changes without task‑specific fine‑tuning.

By Kalina P. Slavkova, Eric Brattain, Aditya Gowd, Akash Pattnaik, Jean-Benoit Delbrouck, Matthew Morgan, Julie Bauml, Javid Abderezaei, Khan Siddiqui
arXiv Machine Learning
Sep 23

MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction

MMAP is a Multimodal Missing‑Aware Alignment Pretraining method designed to learn image‑tabular representations from incomplete data. It uses a sigmoid contrastive learning image encoder with generative reconstruction, a tabular encoder based on a foundation model, and a missing token generator to handle missing modalities. The approach is evaluated on longitudinal Alzheimer’s tasks—predicting disease stage conversion and amyloid status—and outperforms both multimodal and unimodal baselines.

By Fiona Kekwick, Matthew Baugh, Bernhard Kainz, Paul M. Matthews, Wenjia Bai
arXiv Machine Learning
Sep 23

Evidential Fusion Network for Multimodal Survival Prediction under Missing Modalities

The paper introduces the Evidential Missing Modality Survival Fusion (EMMS) model, which predicts survival outcomes using multimodal data even when some modalities are missing. EMMS applies Dempster‑Shafer theory and Gaussian Random Fuzzy Numbers to fuse information, accounting for both aleatoric and epistemic uncertainty and the reliability of each modality. Experiments on four cancer datasets show that EMMS achieves state‑of‑the‑art performance while providing calibrated, interpretable uncertainty estimates without extra computational cost.

By Yucheng Xing, Hailan Mo, Zi Wang, Ling Huang, Mengling Feng
arXiv Machine Learning
Sep 23

FMMD: A multimodal multidisciplinary dataset of open peer reviews from F1000Research

FMMD is a multimodal, multidisciplinary dataset of open peer reviews from F1000Research that pairs manuscript-level visual and structural data with version‑specific reviewer reports and editorial decisions. It addresses key gaps in existing datasets by preserving precise alignment between review comments and the exact manuscript version, and by including a wide range of scientific disciplines beyond computer science. The dataset supports tasks such as visual‑semantic consistency classification, figure‑related review comment generation, and editorial decision prediction, providing a comprehensive empirical resource for multimodal automated scholarly paper review research.

By Zhenzhen Zhuang, Yuqing Fu, Jing Zhu, Zhangping Zhou, Jialiang Lin
arXiv Machine Learning
Sep 23

Unified Multimodal Uncertain Inference

Unified Multimodal Uncertain Inference (UMUI) is a new task that requires models to generate calibrated probability estimates for hypotheses conditioned on premises across text, audio, and video modalities. The authors create a human‑annotated evaluation set with scalar probability judgments for audio, visual, and audiovisual settings, and benchmark their approach on existing text and audio datasets. Their CLUE framework, which blends self‑consistent teacher calibration with distribution‑based confidence probing, enables a 3B‑parameter model to match or surpass zero‑shot baselines up to 32B parameters across all modalities.

By Dengjia Zhang, Alexander Martin, William Jurayj, Kenton Murray, Benjamin Van Durme, Reno Kriz
arXiv Computation and Language
Sep 23

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real‑world multimodal productivity, offering enhanced multimodal understanding, reasoning, and long‑horizon agentic task performance. It builds on a sparse mixture‑of‑experts architecture, extends its context window to one million tokens, and supports long‑context multimodal reasoning and planning. The release includes Qwen-MM-Plugins for native audio and video support and Qwen-Live-Harness for building responsive, real‑time multimodal agents, with extensive evaluations confirming strong performance across multimodal tasks.

By Qwen Team
arXiv Computation and Language
Sep 23

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding

CausalEmbed is an auto‑regressive method for generating compact multi‑vector embeddings in visual document retrieval. By using iterative margin loss during contrastive training, it reduces the number of visual tokens needed by 30‑155× while keeping performance competitive across different backbones and benchmarks. The approach offers efficient training, scalable test‑time performance, and a flexible scaling strategy for multi‑vector representations.

By Jiahao Huo, Yu Huang, Yibo Yan, Ye Pan, Kening Zheng, Wei-Chieh Huang, Yi Cao, Mingdong Ou, Philip S. Yu, Xuming Hu
arXiv Computation and Language
Sep 23

Calibrated Confidence Expression for Radiology Report Generation

The paper introduces ConRad, a reinforcement learning framework that fine‑tunes large vision‑language models to generate calibrated verbalized confidence estimates for radiology reports. ConRad offers both a single report‑level confidence score and a sentence‑level variant, trained with the GRPO algorithm and logarithmic scoring rewards to encourage truthful self‑assessment. Experiments show significant calibration improvements over existing methods, and clinical evaluation indicates that report‑level scores align well with clinicians’ judgments, enabling targeted review of low‑confidence statements.

By David Bani-Harouni, Chantal Pellegrini, Julian L\"uers, Su Hwan Kim, Markus Baalmann, Benedikt Wiestler, Rickmer Braren, Nassir Navab, Matthias Keicher
arXiv Computer Vision
Sep 23

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

The paper introduces ImIR, a method that tunes a large pretrained image‑editing model for all‑in‑one image restoration by replacing text prompts with continuous image‑derived instructions. The approach uses a lightweight token mapper to shift the degraded image’s vision‑language embedding toward that of a clean image, enabling a single adapter to handle six restoration tasks in about three hours on one GPU. ImIR outperforms text conditioning in matched comparisons and supports task‑agnostic restoration without requiring a degradation label.

By S\"uleyman Aslan, G\"orkay Aydemir, M{\i}sra Yavuz, Yunus Bilge Kurt, Nasrin Rahimi, Ahmet Rasim Emirda\u{g}{\i}, Burak Can Biner, M. Ak{\i}n Y{\i}lmaz
arXiv Computer Vision
Sep 23

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

RGSQ introduces a Riemannian geometry‑aware post‑training quantization method for large vision‑language models, treating quantization as a reconstruction problem under a Fisher‑Riemannian metric. It identifies modality‑specific sensitive directions via manifold mappings and applies geometry‑aligned rotations and whitening to steer low‑bit perturbations toward loss‑insensitive axes. Experiments on diverse VLM benchmarks show RGSQ delivers the best accuracy and stability in extremely low‑bit settings, outperforming existing VLM‑aware baselines by up to 5.9% and single‑modality methods by up to 8.6%.

By Zhiping Wu, Dongdong Ren, Yangchengyu Zhou, Zhengjie Zhang, Wenbin Li, Hongbing Pan, Yang Gao
arXiv Computer Vision
Sep 23

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

The paper introduces STD, a hierarchical token pruning framework for Large Vision‑Language Models that aligns pruning strategies with the functional roles of different network stages. By using high‑frequency spectral analysis in shallow layers, Gaussian‑smoothed attention in intermediate layers, and a stability‑adaptive trigger in deep layers, STD preserves essential visual information while aggressively reducing token counts. Experiments demonstrate that STD outperforms existing pruning methods, achieving up to 94.4% token reduction and a 3.9× speed‑up on LLaVA‑NeXT‑7B.

By Shuo Zhang, Jintao Tong, Yixiong Zou, Yuhua Li, Ruixuan Li
arXiv Computer Vision
Sep 23

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

Fysiverse-3D-Vision is a unified vision‑language‑geometry framework that reconstructs executable 3D scenes from a single image. It separates spatial layout reasoning from asset synthesis, using a shared representation where spatial reasoning and geometric reconstruction reinforce each other. The model employs a Transformer that integrates textual supervision, semantic visual cues, and geometric representations, and includes an object‑conditioned layout module to predict object translation, rotation, and scale while maintaining physical consistency through collision‑aware optimization.

By Dingkang Yang, Yizhou Liu, Wendong Cheng, Zizhi Chen, Shunli Wang, Yang Liu, Hongsheng Li, Lihua Zhang
arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
arXiv Computer Vision
Sep 23

Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models

Video-HopChain introduces a new dataset of 22,550 multi‑hop video questions over 13,378 videos, each question consisting of three to six yes/no sub‑questions whose integer answers sum to a verifiable reward. Training a Qwen3‑VL‑8B model with GRPO on this dataset improves performance across eight video‑understanding benchmarks from 55.4 to 57.9, and the addition of Confidence‑Gated Exploration (CGE) raises the mean to 59.3. The authors release the dataset, checkpoint, and training code for further research.

By Trung Nguyen Quang, Yuhao Dong, Shuo Sun, Shuai Liu, Shulin Tian, Kim-Hui Yap, Ziwei Liu