Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Machine Learning
3d ago

Uncertainty Quantification for Flow-Based Generalist Robot Policies

The paper introduces a method for quantifying epistemic uncertainty in flow‑matching based generalist robot policies, such as vision‑language‑action models and world‑action models. By measuring velocity‑field disagreement across a small ensemble, the authors obtain better‑calibrated uncertainty estimates that can detect deployment failures and guide active fine‑tuning. Their SAVE approach reduces the need for expert demonstrations, improving real‑world task success from 39 % to 47 % while maintaining a fixed demonstration budget.

By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
arXiv AI
3d ago

THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS

The paper introduces THPL, a vision-to-language framework for precision feeding of rainbow trout in recirculating aquaculture systems. It uses Fishsort to extract feeding trajectories, a Hierarchical Behavior Encoder to model individual and collective dynamics, and integrates these with environmental data and expert rules to fine‑tune a large language model via LoRA and counterfactual multimodal Direct Preference Optimization. Experimental results show strong correlation between the Activity Coefficient and expert feeding intensity, and significant improvements in decision accuracy and language metrics when using dual‑evidence tokens and counterfactual optimization.

By Meng Liang, Guanbo Feng, Haozhuang Chi, Shilong Zhao, Zhixin Xiong, Yuhang He, Wenfeng Han, Tianhao Zhao, Zhihong Ma, Ying Liu
arXiv Computer Vision
3d ago

Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

The paper introduces IG‑VLA, a vision‑language‑action framework that learns to imagine task‑relevant future scene evolution in a latent spatiotemporal space, guiding action prediction without generating full pixel‑level videos. It further compresses this future reasoning into a compact Scene Gist Token via Scene Gist Memory, enabling efficient inference. Experiments on LIBERO, LIBERO‑Plus, and VLABench show that IG‑VLA improves success rates by nearly 6% and speeds up inference up to 6.38× compared to strong baselines.

By Shenglan Li, Zhendong Mi, Hengyi Zhu, Jingwu Luo, Chun Kit Chan, Geng Yuan, Yanzhi Wang, Pu Zhao, Shaoyi Huang
arXiv Computer Vision
3d ago

Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration

Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration proposes CDPM, a method that first aligns semantic representations across modalities and then refines correspondences with fine-grained CNN features. CDPM adapts DINOv3 using geometrically consistent cross-modal patch pairs, builds a DINO-Centric Feature Pyramid for stable cross-modal matching, and adds a lightweight CNN branch for precise local refinement. Experiments on three cross-modal datasets show that CDPM outperforms existing dense matchers, improving AUC metrics and reducing mACE while using fewer FLOPs.

By Zhiwei Wang, Defeng He, Yuxing Li, Meilu Zhu, Edmund Y. Lam
arXiv Computer Vision
3d ago

UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation

UniDynamics is a diffusion-based framework that generates future 4D dynamic scenes—comprising RGB, depth, and optical flow—from a single event-RGB pair, without needing long histories or control priors. It introduces an Event Latent Enhancement module to align event data into conditioning features and a Perceptual Dynamics Space within a multi-scale U‑Net to decouple and adaptively interact depth and flow, enforcing geometric and motion constraints for coherent predictions. Experiments on VKitti2 and DSEC show state‑of‑the‑art performance, producing high‑quality, temporally coherent, and 4D‑consistent predictions even under high‑speed motion blur.

By Daikun Liu, Xin Zhan, Teng Wang, Xiaoping Wang, Changyin Sun
arXiv AI
3d ago

DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

DriftTTS is a few‑step neural text‑to‑speech model that generates mel‑spectrograms without relying on a generative teacher, distillation, or adversarial discrimination. It employs a distribution‑matching drift objective in a mel‑domain feature space defined by raw mels and a frozen masked‑autoencoder encoder pretrained on LJSpeech. On the LJSpeech dataset, DriftTTS achieves competitive metrics (3.87 dB MCD, 3.7% WER) and a MOS of 4.18, rivaling the Matcha‑TTS baseline and approaching ground‑truth quality.

By Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam
arXiv AI
3d ago

Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition

The paper introduces ECHO-$k$, a self-supervised, task-agnostic method for selecting which modalities to acquire at test time in multimodal, high-dimensional learning. By using a deep model’s pretrained representations as proxy targets, ECHO-$k$ learns a reinforcement‑learning policy that sequentially chooses informative modalities, providing theoretical guarantees in a linear setting. Experiments show that ECHO-$k$ consistently improves budgeted downstream performance across various foundation‑model backends, offering a principled approach to cost‑aware test‑time deployment when measurements are expensive or time‑constrained.

By Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne
arXiv AI
3d ago

Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

The paper presents a mechanistic defense for Vision‑Language‑Action (VLA) models against adversarial patches. By using a sparse autoencoder, the authors identify a feature whose activation correlates strongly with the presence of an adversarial patch and suppress this feature only when a linear probe detects an attack. This conditional intervention improves robustness on the LIBERO‑10 benchmark while avoiding the performance degradation that occurs with continuous suppression.

By Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi
arXiv AI
3d ago

Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

The paper introduces AdsCVR, a benchmark for e‑commerce cross‑video reasoning with 2,483 videos and 6,110 QA pairs across six reasoning dimensions. It proposes AdSeek, an agentic framework that actively selects visual and audio tools during multi‑turn exploration and uses an offline trajectory rectification mechanism to improve reinforcement learning. AdSeek achieves 74.30% accuracy on AdsCVR, outperforming its backbone by 27.90 percentage points and generalizes to the open‑domain CrossVid benchmark.

By Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng
arXiv Machine Learning
3d ago

When Is Accuracy Evidence? A Unified Theory of Generalisation, Validation, and Information Fusion

The paper introduces a unified exponential framework for generalising K‑fold cross‑validation (CV) accuracy into conservative risk bounds, called gamma‑CUBV. It models dependence between folds via a joint sub‑Gaussian proxy matrix, yielding an effective number of folds and showing that more folds do not always increase evidence when data are strongly correlated. The framework extends to posterior predictor distributions and weighted multi‑source fusion, and is empirically evaluated on linear classifiers trained on heterogeneous multimodal Gaussian mixtures, demonstrating that corrected resubstitution can be tighter than K‑fold CV in low‑sample, high‑heterogeneity settings.

By JM Gorriz
arXiv AI
3d ago

TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods

TasteBench is a multimodal benchmark designed to accelerate sustainable protein discovery by providing computational proxies for sensory prediction. It includes a food-level ranking task based on over 21,000 human evaluations of 215 plant-based foods across 24 categories, and a molecular-level taste classification task covering 15,000 flavor molecules. The benchmark offers baseline models, characterizes inter-rater agreement and reliability limits, and demonstrates that the best model achieves pairwise accuracy comparable to individual human panelists.

By Anna T. Thomas, Sohum Patnaik, Caroline Cotto, Benjamin Sanchez-Lengeling
arXiv Computer Vision
3d ago

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

HakushoBench is a Japanese chart and table visual question answering benchmark created from 33 governmental white papers, comprising 2,053 images across more than ten types. The dataset includes manually annotated and independently verified QA pairs that evaluate holistic understanding of charts and tables rather than just local visual cues. Experiments show that HakushoBench is significantly harder than existing Japanese benchmarks, with sub‑10B open‑weight models achieving at most 58.6% accuracy and even large models like Qwen3.5-397B-A17B lagging behind Gemini‑3‑Pro by 8.1 points.

By Issa Sugiura, Shuhei Kurita, Yusuke Oda, Naoaki Okazaki
arXiv Machine Learning
3d ago

Balancing Multimodal Learning via Functional Progress

The paper introduces Function‑Space Guided Multimodal Optimization (FGMO) to address modality imbalance in multimodal learning. FGMO uses Functional Progress Estimation (FPE) to measure each modality’s function‑space response and Functional Response Control (FRC) to redistribute learning‑rate budgets, aiming to balance optimization progress across modalities. The authors provide theoretical analysis of FRC’s target‑contraction property and demonstrate FGMO’s effectiveness on several multimodal benchmarks.

By Zhongjing Gu, Fengqiang Wan, Yiming Cui, Yufa Feng, Yang Yang
arXiv AI
3d ago

Multi-Modal Environment-Aware Beam Management for Massive MIMO: A Geometry-Driven Virtual Base Station Framework

The paper presents a geometry-driven framework for beam management in high-frequency massive MIMO systems. It constructs an offline virtual base station database using 3D LiDAR point clouds and location data to model dominant reflection paths, enabling a coarse channel reconstruction. A VBS-assisted orthogonal-pilot scheme and a dual-agent dueling double deep Q-network are then employed to refine beam estimates and perform coordinated beam selection, yielding improved training efficiency and performance over existing baselines.

By Yijie Bian, Wei Guo, Jie Yang, Shenghui Song, Jun Zhang, Shi Jin, Khaled B. Letaief
arXiv Machine Learning
3d ago

Structural-Functional Brain Connectivity Generation via Multimodal Hypergraph-based Flow Matching

The paper introduces Multimodal Hypergraph Flow Matching (MHG‑FM), a framework that jointly generates structural connectivity (SC) and functional connectivity (FC) by constructing modality‑specific hypergraphs and learning higher‑order representations with Hypergraph Neural Network encoders. It employs Dual Cross‑Attention for bidirectional cross‑modal fusion, a variational autoencoder to map fused representations into a latent space, and conditional flow matching to synthesize connectivity and perform multimodal translation. Experiments on the Human Connectome Project Young Adult dataset demonstrate that MHG‑FM outperforms state‑of‑the‑art baselines in reconstruction quality, topology preservation, distributional similarity, and SC‑FC coupling, while achieving roughly eight times faster sampling than a comparable diffusion backbone.

By Chyong Yi Poh, Hwa Hui Tew, Junn Yong Loo, Rapha\"{e}l C. -W. Phan, Fuad Noman, Pew-Thian Yap, Chee-Ming Ting
arXiv Machine Learning
3d ago

Confidence-Gated Cloud-Edge Cascade Triage via Variational Risk Minimization for Medical Imaging

The paper introduces Variational Risk Minimization (VRM), a distillation framework that treats large vision‑language model (LVLM) report variants as Monte Carlo samples of latent clinical interpretations. VRM learns from a variationally marginalized teacher distribution, providing uncertainty‑aware supervision even when modalities are missing, and outperforms fine‑tuning baselines while improving calibration. In a compact edge‑student setup, a confidence‑gated cascade achieves an AUC of 0.941 at 103 ms latency with only 20.3 % cloud escalation, offering a clear reliability‑latency trade‑off for cloud‑edge medical imaging workflows.

By Xinye Yang, Zhusi Zhong, Scott Collins, Michael Bernstein, Grayson Baird, Terrence Healey, Michael Atalay, Mahesh Jayaraman, Xuyu Wang, Zhicheng Jiao
arXiv AI
3d ago

Zephon: Elastic Determinism for Online, Stateful Foundation Model Data Loading Pipelines

Zephon is a data loader designed for foundation model training that ensures deterministic ordering of training data batches even when GPU resources change, checkpoints are resumed, or different processing backends are used. It handles online, stateful pipelines—where tokenization, packing, and mixing of samples create complex n‑to‑m transformations—by partitioning the data stream into topology‑independent lanes, serializing ordering decisions, and checkpointing only bounded in‑flight state. Experiments on text and vision‑language tasks show Zephon delivers competitive throughput while offering guarantees that existing loaders lack for such pipelines.

By Maximilian B\"other, Josh Wills, Ties Robroek, Sonnet Xu, Paul Burstein, Daniel Zayas, Cody Blakeney, Siddharth Joshi, Haoli Yin, Rishabh Adiga, Haakon Mongstad, Luke Merrick, Pratyush Maini, Ari Morcos, Matthew Leavitt, Ana Klimovic, Bogdan Gaza
arXiv Computer Vision
3d ago

From Language Priors to Field Adaptation: Preference Learning for Traversability Estimation

The paper proposes a method for image-based traversability estimation that uses von Mises‑Fisher mixture prototypes in a frozen vision‑language feature space. By incorporating relative natural‑language rules as a commonsense prior and fine‑tuning with sparse image annotations, the approach achieves domain adaptation with fewer labels. Experiments on WayFAST show competitive accuracy to end‑to‑end models and improved dense predictions on semantic maps, while qualitative results illustrate zero‑shot language prior use and semantic interpretation of prototypes.

By Simon Schwaiger, David Seyser, Alessandro Scherl, Zlatan Ajanovi\'c, Wilfried W\"ober, Gerald Steinbauer-Wagner
arXiv Computer Vision
3d ago

FiberGeoText: A Vision-Language Model for Population- Level Organization of Superficial White Matter

FiberGeoText (FGT) is a vision‑language model that clusters short‑range superficial white matter streamlines from ultra‑high‑resolution diffusion MRI into population‑level groups. It jointly encodes each streamline’s 3‑D trajectory, cortical anatomical context (via text from multiple parcellation schemes), and shape, using a pretrained large language model to unify heterogeneous anatomical descriptions. Evaluations on 0.76 mm diffusion data show that FGT outperforms state‑of‑the‑art methods in cortical parcel coherence, shape consistency, cluster‑size consistency, and cross‑subject correspondence, and it generalizes well to unseen subjects, recovering 96.7 % of learned clusters.

By Yuqian Chen, R. Jarrett Rushmore, Guikun Chen, Fan Zhang, Edward Yeterian, Nikos Makris, Yogesh Rathi, Lauren J. O'Donnell
arXiv AI
3d ago

Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking

The paper introduces Causal-Invariant Masking (CIM) to better quantify epistemic uncertainty in Multimodal Large Language Models (MLLMs) by measuring semantic shift between original predictions and those conditioned on a causally-focused view. It proposes Semantic Divergence as a core metric that converges to the variance of the model’s sensitivity to non‑causal correlations, and introduces Expected Embedding Drift (EED) as a fast geometric proxy that estimates this shift directly in the embedding space. Experiments demonstrate state‑of‑the‑art uncertainty quantification performance and a nearly 50% speedup with EED.

By Haoyang Luo, Linwei Tao, Jie Gui, Xinghao Chen, Chang Xu, Jianyuan Guo, Minjing Dong