Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv AI
6d ago

FairSSL: Fair Multimodal Self-Supervised Learning

FairSSL is a multimodal self‑supervised learning framework that treats data heterogeneity as a fairness resource instead of a limitation. It replaces strict alignment with a subject‑aware Variance‑Invariance‑Covariance Regularization objective, enforcing alignment only across segments from the same subject. The method includes segment‑based pooling for variable‑length modalities and regularizes representations to promote within‑subject variability, cross‑modal and cross‑subject invariance, and decorrelation, theoretically bounding the score gap between protected groups and empirically outperforming baselines on heterogeneous multimodal datasets.

By Jiaee Cheong, Abtin Mogharabin, Paul Liang, Hatice Gunes, Sinan Kalkan
arXiv AI
6d ago

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

UrbanVLA is a Vision‑Language‑Action framework designed to enable delivery robots to navigate large‑scale urban environments using long‑horizon route instructions. The model aligns noisy route waypoints with visual observations and plans trajectories, trained through a two‑stage pipeline of supervised fine‑tuning on simulated data and reinforcement fine‑tuning on mixed simulation and real‑world data. Experiments show UrbanVLA outperforms strong baselines by over 55% on the SocialNav task and demonstrates reliable real‑world navigation in large urban settings.

By Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, Zhibo Chen, Zhizheng Zhang, He Wang
arXiv AI
6d ago

Do Vision Language Models Understand Human Engagement in Games?

The study investigates whether vision–language models (VLMs) can infer human engagement from gameplay videos using the GameVibe Few‑Shot dataset across nine first‑person shooter games. Three VLMs were tested under six prompting strategies—including zero‑shot, theory‑guided prompts based on Flow, GameFlow, Self‑Determination Theory, and MDA, and retrieval‑augmented prompting—evaluating both pointwise engagement prediction and pairwise prediction of engagement change. Results show that zero‑shot predictions are weak and often do not beat simple majority‑class baselines; retrieval‑augmented prompting improves pointwise prediction in some cases, while pairwise prediction remains difficult, and theory‑guided prompts do not reliably help and may reinforce superficial shortcuts. "whyItMatters":"The findings highlight a perception–understanding gap in current VLMs, indicating that while they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games."

By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
arXiv AI
6d ago

Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning

Rewind-IL is a training‑free online safeguard for generative action‑chunked imitation learning policies. It uses a zero‑shot failure detector based on Temporal Inter‑chunk Discrepancy Estimate (TIDE) and a state‑respawning mechanism that returns the robot to a verified safe intermediate state. The system builds a checkpoint library offline with a vision‑language model and monitors self‑consistency online, rewinding execution to the latest safe checkpoint when a failure is detected, thereby improving reliability in long‑horizon manipulation tasks.

By Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi
arXiv Machine Learning
6d ago

LEARN-TS: LLM-Enhanced Alignment and Reconstruction with Normality Guidance for Multivariate Time-Series Anomaly Detection

LEARN-TS is a method for multivariate time‑series anomaly detection that uses a frozen language model to generate two semantic representations: one that captures window‑specific observation context without exact numerical values, and another that provides a fixed, dataset‑agnostic normality prompt. These representations guide a patch‑masked reconstruction process, allowing the model to estimate normality discrepancy and produce timestamp‑level anomaly evidence. Experiments on four benchmarks show that LEARN‑TS achieves the best mean performance in 13 of 16 dataset‑metric comparisons, and ablation studies confirm the benefits of observation conditioning, joint normality alignment, and semantic references.

By Jahyeob Koo, Kio Yun, Byoungmo Koo, Jun-Geol Baek
arXiv Machine Learning
6d ago

Diffusion-2BC: Hybrid Diffusion and Regression Training for Offline Behavior Cloning in Autonomous Driving

Diffusion-2BC is a hybrid training method that combines a diffusion denoising objective with an auxiliary deterministic behavior‑cloning loss over a shared visual encoder. The auxiliary loss is used only during training, while inference remains diffusion‑based. Experiments on the Claw environment and CARLA navigation show that Diffusion‑2BC reduces mean mask‑distance error by about 10% compared to a diffusion baseline and by 85% compared to standard deterministic behavior cloning, and it enables the agent to travel farther and exhibit multimodal route choices.

By Bruno Maciel Machado, Eric Aislan Antonelo
arXiv Machine Learning
6d ago

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

The paper investigates a specific failure mode in vision‑language‑action models, termed instruction‑action binding, where models respond to language and vision separately but fail to combine them to select the correct action under counterfactual changes. Through behavioral analyses of fine‑tuned policies, the authors show that failed rollouts often preserve source behavior or switch to other demonstrated tasks, indicating that language is not ignored but mis‑bound. They propose Equivariant Counterfactual Training (ECT), which supplies counterfactual demonstrations and a paired loss to enforce correct action selection, achieving significant performance gains across simulated and real‑world benchmarks.

By Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
arXiv Machine Learning
6d ago

Data-Driven Priors for Uncertainty-Aware Risk Prediction of Clinical Deterioration using Multimodal Data

The paper introduces MedCertAIn, a framework that uses data‑driven priors to enhance uncertainty estimation in multimodal clinical models. By integrating cross‑modal similarity and modality‑specific corruptions into neural network priors, the authors improve predictive performance and selective prediction for in‑hospital mortality risk using MIMIC‑IV and MIMIC‑CXR data. The results demonstrate competitive accuracy and notable gains over deterministic and stochastic baselines, underscoring the potential of such priors for reliable, uncertainty‑aware clinical decision support.

By L. Juli\'an Lechuga L\'opez, Tim G. J. Rudner, Farah E. Shamout
arXiv Computer Vision
6d ago

GPEC: Efficient Pre-LLM Gaussian Process Embedding Correction for Cardiac Video Caption Generation

GPEC is a lightweight pre‑LLM Gaussian Process Embedding Correction that refines visual representations for cardiac ultrasound caption generation. By inserting a residual correction layer between the visual projection and the language model, it aligns projected embeddings with annotation‑guided targets derived from structured video annotations. Experiments show that GPEC improves caption similarity and content alignment while adding less than 0.05 s of inference overhead, all without fine‑tuning the pretrained multimodal backbone.

By Arefeh Rezaei
arXiv Computer Vision
6d ago

Harnessing Vision-Language Models for Perceptual Quality Assessment and Autonomous Content Adjustment in Augmented Reality

The paper introduces RateAR, a benchmark of AR images and videos that captures perceptual factors such as object placement, scale, and shadow consistency. Eleven commercial vision‑language models (VLMs) are evaluated on this benchmark, with VLM‑based quality predictions showing strong correlation with human judgments (Spearman’s ρ up to 0.8695). Using contextual prompting, the authors build an automated AR content adjustment system that, in a 21‑participant study, improved placement and size coherence of virtual content for over 90% of users.

By Elias Rotondo (Duke University), Lin Duan (Duke University), Yanming Xiu (Duke University), Sangjun Eom (Duke University), Conrad Li (Duke University), Maria Gorlatova (Duke University)
arXiv Computer Vision
6d ago

Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

The paper introduces Align Then Reason (ATR), a multilingual lip‑sync judge that aligns frame‑level lip representations with phonetic units of a candidate text line and then uses a language model to evaluate both content and timing. ATR achieves significant improvements over existing baselines on a seven‑language benchmark, with mean AUC gains of up to 59.4% for 2B reasoners and similar gains across other LLM families. The method also transfers well to unseen languages and outperforms lip‑reading baselines on real dubbing tasks such as dub‑line reranking and script‑to‑clip assignment.

By Rui Liu, Bhavin Jawade, Haoqi Li, Shivam Mehta, Karan Saxena, Yinghong Lan, Cameron R. Wolfe
arXiv Computer Vision
6d ago

Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning

Watch Your Speech (WYS) is a video‑to‑speech synthesis framework that uses textual conditioning to resolve the one‑to‑many mapping problem inherent in silent talking‑face videos. It fuses textual context with video sequences via an attention‑based embedding module and employs a conditional flow matching objective to produce high‑fidelity, phonemically accurate speech. Experiments on LRS2 and LRS3 show that WYS sets new state‑of‑the‑art results in audio‑visual synchronization while maintaining competitive word‑error rates, and subjective tests confirm near‑human naturalness.

By Gunwoo Lee, Yoori Oh, Yoseob Han
arXiv Computer Vision
6d ago

FutureWorlds: Learning Robotic World Models from Alternative Futures

FutureWorlds is a framework that learns robotic world models by generating alternative future predictions and using them as learning signals. It combines candidate construction, history maintenance, and learning from relative quality through a multimodal discrete autoregressive model and diverse beam search. The MemSPO algorithm further optimizes the world model by converting video trajectory rewards into group-relative advantages, leading to significant improvements in LPIPS scores across RT-1, BridgeV2, and RoboCasa datasets.

By Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang
arXiv Computer Vision
6d ago

Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models

Dyna3 is a training‑free framework that extends the depth foundation model DA3 to perform 4D dynamic scene reconstruction without fine‑tuning. By leveraging DA3’s cross‑view features and a best‑match search, it distinguishes static surfaces from moving objects, and uses vision‑language models to generate semantic prompts for SAM 3 to achieve precise instance‑level segmentation. Experiments on four datasets show Dyna3 outperforms correspondence‑trained methods, improving dynamic object segmentation by +5.5 pp, speeding pose estimation 13×, and reducing memory usage 4–8×.

By Xinhao Xiang, Weiyang Li, Zhijie Zheng, Abhijeet Rastogi, Jiawei Zhang
arXiv Computer Vision
6d ago

Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models

The paper introduces MOSAIC, a method for personalized federated vision‑language models that addresses domain heterogeneity by focusing on class‑specific cross‑domain residuals. It constructs a decision‑aware harmfulness score to identify residuals that hurt image‑text decision margins, then uses a low‑rank residual adapter with shared class factors and private domain factors, an image‑conditioned gate, and harmful‑pair‑aware reweighting to refine local updates. Experiments on Office31, OfficeHome, and DomainNet100 show consistent improvements in macro‑client top‑1 accuracy across various domain‑shift scenarios.

By Wentao Yue, Qingyu Mao, Tianyou Lai, Ahmed M. Abdelmoniem, Qilei Li
arXiv Computer Vision
6d ago

UnifiedAttack: Evaluating the Safety of Large Multimodal Models in Synergistic Harmful Image-Text Generation

UnifiedAttack introduces a benchmark for testing the safety of large multimodal models (LMMs) in tasks that combine text and image to produce harmful content. The benchmark focuses on the additional harm that arises from cross‑modal synergy and includes filtered multimodal samples and synthesized disinformation queries. A synergistic hijacking framework—comprising In‑Context Reskinning (ICR) and Cognitive Planning Injection (CPI)—is proposed to expose vulnerabilities, and extensive evaluations show that UnifiedAttack consistently bypasses current alignment defenses in state‑of‑the‑art architectures.

By Bingjun Luo, Jialin Guo, Tony Wang, Siqi Li
arXiv Computer Vision
6d ago

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

The paper introduces Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for multi‑label video safety detection. ATPO uses an Adaptive Tversky Reward (ATR) that dynamically adjusts false‑positive and false‑negative penalties, allowing controllable precision‑recall trade‑offs. Experiments on SafeWatch‑Bench and XD‑Violence demonstrate significant performance gains, raising the Jaccard Index from 40.66 to 75.44 on SafeWatch‑Bench‑Real.

By Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne
arXiv Computer Vision
6d ago

FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering

FocusGraph is a framework for selecting keyframes in egocentric long‑video question answering. It uses a lightweight Scene‑Graph LLM Selector to identify query‑relevant clips from graph‑based captions, then extracts keyframes with Patch‑wise Sparse‑Flow Retention (PSFR) before feeding them to a multimodal large language model for answer generation. The method achieves state‑of‑the‑art performance on FindingDory and HourVideo while reducing question‑time inference cost.

By Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin
arXiv Computer Vision
6d ago

DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation

DynamicVLA is a latency‑aware Vision‑Language‑Action model designed for dynamic object manipulation, featuring a compact 0.4B architecture and a convolutional vision encoder for efficient multimodal inference. It employs a continuous inference schedule that overlaps reasoning and execution, and a Latent‑aware Action Streaming mechanism that discards stale action prefixes to maintain action‑time alignment. The authors also introduce the Dynamic Object Manipulation (DOM) benchmark, comprising 200K synthetic episodes and 2K real‑world episodes, and demonstrate that DynamicVLA improves dynamic manipulation success in simulation and on real robots.

By Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu
arXiv AI
6d ago

VIDA: A Dataset for Visually Dependent Ambiguity in Multimodal Machine Translation

VIDA (Visually-Dependent Ambiguity) is a new dataset comprising 2,500 curated instances designed to test whether multimodal machine translation models can use visual input to resolve ambiguous source spans. The authors also introduce Disambiguation-Centric Metrics, which employ an LLM-as-a-judge classifier to verify correct span-level resolution. Experiments with recent large vision‑language models show that visual disambiguation remains difficult, and that chain‑of‑thought supervised fine‑tuning yields stronger out‑of‑distribution performance, especially for collective‑noun ambiguities.

By Jingheng Pan, Xintong Wang, Longyue Wang, Liang Ding, Weihua Luo, Chris Biemann