Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,568 stories · RSS feed

Hugging Face Trending Papers
2d ago

Mind the Accent Gap: British Accent Robustness in Speech-Driven Financial Voice Assistants

The paper highlights that AI voice assistants using ASR and LLMs struggle with regional British accents because most ASR models are trained on American English. It introduces CavaBench, a benchmark of spoken financial queries, to evaluate ASR models and their impact on downstream tool‑calling accuracy across British accents. The study finds that while WER predicts tool‑calling accuracy, it may not fully capture task‑level performance, revealing accent‑related failures that vary by model and acoustic conditions.

Hugging Face Trending Papers
2d ago

MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity

MTOR introduces a generalizable AI‑generated video detector that combines global visual features with caption‑derived textual semantics and a novel Temporal Over‑Regularity (TOR) component. The TOR module captures three levels of temporal consistency—coarse inter‑frame continuity, fine‑grained token correspondence, and frame‑to‑video stability—to exploit the stronger temporal persistence and lower variability found in synthetic videos. Extensive tests on five benchmarks with 46 generator variants show MTOR outperforms 16 baselines and remains robust against twelve real‑world video perturbations.

Hugging Face Trending Papers
2d ago

MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting

MoCAR (Motion-code Coordinate-aware AutoRegression) is a decoder‑only framework that treats continuous trajectory forecasting as next‑code prediction in a coordinate‑aware latent space. It learns a continuous motion‑code space from endpoint‑normalized trajectory segments, using historical codes as a teacher‑forced prefix and generating future codes autoregressively while updating local scene context. The method achieves top‑tier performance on Argoverse benchmarks, transfers well from AV2 to AV1 in zero‑shot evaluation, and improves on turn‑heavy scenarios without requiring trajectory‑space re‑tokenization or proposal‑and‑refinement pipelines.

Hugging Face Trending Papers
2d ago

Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation

The paper evaluates CLIP’s zero‑shot gender estimation on full‑face and periocular images from the Adience dataset. Using image‑text similarity with male/female prompts, CLIP achieves 95.54% accuracy on full faces without task‑specific training. Periocular predictions are initially biased toward males, but threshold alignment improves accuracy to 85.29%, and a linear SVM on CLIP features yields a modest 86.17% accuracy, slightly better than prior Adience results.

Hugging Face Trending Papers
2d ago

Scalable Minimal-Change Learning for Controllable Image Editing

The paper introduces ARRO, a reinforcement‑learning framework that enforces minimal‑change editing by auditing source images, instructions, and edited outputs for unimplemented and unintended changes. Using a vision‑language reward model and a group‑level rubric, ARRO achieves higher EditScore on multiple benchmarks and reduces off‑target pixel changes by 8.4% on FLUX.1 Kontext‑dev. The approach eliminates the need for per‑instruction human annotations and demonstrates strong performance across several evaluation sets.

arXiv Computer Vision
2d ago

Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild

The paper compares text‑based and feature‑based models for recognizing compound emotions in real‑world videos. It proposes textualizing non‑verbal cues from audio and visual modalities into text to leverage large language models, while feature‑based models directly combine extracted multimodal features. Experiments on the C‑EXPR‑DB dataset show that feature‑based models outperform textualization in the wild, though textual models can excel when rich transcripts are available.

By Nicolas Richet, Soufiane Belharbi, Haseeb Aslam, Meike Emilie Schadt, Manuela Gonz\'alez-Gonz\'alez, Gustave Cortal, Alessandro Lameiras Koerich, Marco Pedersoli, Alain Finkel, Simon Bacon, Eric Granger
arXiv Machine Learning
2d ago

Architecture-Dependent Fusion Pathways in MLLMs

The paper investigates how visual and textual information are fused in Multimodal Large Language Models (MLLMs). By analyzing concatenation and native multimodal architectures through alignment decoupling, attention routing, entropy, intrinsic dimensionality, and causal interventions, the authors uncover two distinct fusion pathways: concatenation models use a text‑first, vision‑later strategy, while native models integrate vision and text earlier and reorganize feature spaces. The study also employs visual CKA to test the Platonic Representation Hypothesis, offering a mechanistic view of multimodal fusion and informing architecture‑aware diagnostics.

By Hebao Zhu, Dongxia Wu
arXiv Machine Learning
2d ago

ZeroMAG: Zero-Shot Multimodal Adapter Generation for Plug-and-Play EEG Foundation Models

ZeroMAG is a zero‑shot multimodal adapter generation framework that extends frozen EEG foundation models to heterogeneous multimodal recordings using only unlabeled target data. It constructs a configuration‑invariant adapter and generates adapter weights in a latent space learned from source adapters, without target‑side optimization. Across six held‑out target datasets and three EFM backbones, ZeroMAG improves balanced accuracy by 7.22 percentage points over EEG‑only inference and 4.89 points over direct weight regression, approaching supervised multimodal adaptation.

By Yubo Wang, Jingying Ma, Xinliang Zhou, Yangxuan Zhou, Jiquan Wang, Sha Zhao, Yiyuan Yang, Yi Ding, Ziyu Jia, Chenyu Liu, Cuntai Guan
arXiv Computation and Language
2d ago

CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation

CLIMB is a training‑free inference‑time framework for multimodal retrieval‑augmented generation. It builds a compact complementary evidence pool using an MMR‑style objective that balances relevance and redundancy, then refines answers with a confidence‑controlled critic that scores relevance, specificity, and cross‑modal alignment. The method stops refinement when confidence no longer rises, improving performance on Encyclopedic‑VQA and InfoSeek without altering the retriever or language model.

By Hang Gao, Wujiang Xu, Zhixing Zhang, Kai Mei, Jingyi Yang, Dimitris N. Metaxas
arXiv AI
2d ago

VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models

VIDiff is a unified foundation model that uses diffusion techniques to perform a broad range of video tasks, including both understanding tasks like language‑guided video object segmentation and generative tasks such as video editing and enhancement. Unlike prior methods that focus on short clips and require time‑consuming tuning, VIDiff can edit and translate videos within seconds based on user instructions and employs an iterative auto‑regressive approach to maintain consistency in long‑form videos. The authors demonstrate convincing generative results across diverse input videos and written instructions, supported by qualitative and quantitative evidence.

By Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang
arXiv AI
2d ago

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

The paper introduces threat‑preserving representation sensitivity (TPRS) to assess how changes in agent‑visible wording affect attack success rates (ASR) while keeping the underlying task and evaluation criteria constant. Experiments on Agent Security Bench, MCPTox, and AgentDojo show that renaming threat‑related tools can significantly alter ASR—up to +13.21 percentage points on some models—yet may also degrade benign utility. The findings suggest that security scores based on a single representation may not generalize, urging robustness claims to be validated across multiple threat‑preserving representations.

By Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu
arXiv Computer Vision
2d ago

Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs

Gaze Attention is a new mechanism for multimodal large language models that selectively focuses on relevant visual regions during each generation step, rather than attending to all visual tokens. By grouping tokens into spatial regions and using learnable context tokens to retain global information, it reduces attention computation and visual key‑value entries by up to 90%. Experiments on 13 image and 6 video benchmarks show that Gaze Attention matches or outperforms dense‑attention baselines while using fewer visual resources.

By Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun
arXiv Machine Learning
2d ago

Uncertainty Quantification for Flow-Based Generalist Robot Policies

The paper introduces a method for quantifying epistemic uncertainty in flow‑matching based generalist robot policies, such as vision‑language‑action models and world‑action models. By measuring velocity‑field disagreement across a small ensemble, the authors obtain better‑calibrated uncertainty estimates that can detect deployment failures and guide active fine‑tuning. Their SAVE approach reduces the need for expert demonstrations, improving real‑world task success from 39 % to 47 % while maintaining a fixed demonstration budget.

By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
arXiv AI
2d ago

THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS

The paper introduces THPL, a vision-to-language framework for precision feeding of rainbow trout in recirculating aquaculture systems. It uses Fishsort to extract feeding trajectories, a Hierarchical Behavior Encoder to model individual and collective dynamics, and integrates these with environmental data and expert rules to fine‑tune a large language model via LoRA and counterfactual multimodal Direct Preference Optimization. Experimental results show strong correlation between the Activity Coefficient and expert feeding intensity, and significant improvements in decision accuracy and language metrics when using dual‑evidence tokens and counterfactual optimization.

By Meng Liang, Guanbo Feng, Haozhuang Chi, Shilong Zhao, Zhixin Xiong, Yuhang He, Wenfeng Han, Tianhao Zhao, Zhihong Ma, Ying Liu
arXiv Computer Vision
2d ago

Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

The paper introduces IG‑VLA, a vision‑language‑action framework that learns to imagine task‑relevant future scene evolution in a latent spatiotemporal space, guiding action prediction without generating full pixel‑level videos. It further compresses this future reasoning into a compact Scene Gist Token via Scene Gist Memory, enabling efficient inference. Experiments on LIBERO, LIBERO‑Plus, and VLABench show that IG‑VLA improves success rates by nearly 6% and speeds up inference up to 6.38× compared to strong baselines.

By Shenglan Li, Zhendong Mi, Hengyi Zhu, Jingwu Luo, Chun Kit Chan, Geng Yuan, Yanzhi Wang, Pu Zhao, Shaoyi Huang
arXiv Computer Vision
2d ago

Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration

Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration proposes CDPM, a method that first aligns semantic representations across modalities and then refines correspondences with fine-grained CNN features. CDPM adapts DINOv3 using geometrically consistent cross-modal patch pairs, builds a DINO-Centric Feature Pyramid for stable cross-modal matching, and adds a lightweight CNN branch for precise local refinement. Experiments on three cross-modal datasets show that CDPM outperforms existing dense matchers, improving AUC metrics and reducing mACE while using fewer FLOPs.

By Zhiwei Wang, Defeng He, Yuxing Li, Meilu Zhu, Edmund Y. Lam
arXiv Computer Vision
2d ago

UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation

UniDynamics is a diffusion-based framework that generates future 4D dynamic scenes—comprising RGB, depth, and optical flow—from a single event-RGB pair, without needing long histories or control priors. It introduces an Event Latent Enhancement module to align event data into conditioning features and a Perceptual Dynamics Space within a multi-scale U‑Net to decouple and adaptively interact depth and flow, enforcing geometric and motion constraints for coherent predictions. Experiments on VKitti2 and DSEC show state‑of‑the‑art performance, producing high‑quality, temporally coherent, and 4D‑consistent predictions even under high‑speed motion blur.

By Daikun Liu, Xin Zhan, Teng Wang, Xiaoping Wang, Changyin Sun
arXiv AI
2d ago

DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

DriftTTS is a few‑step neural text‑to‑speech model that generates mel‑spectrograms without relying on a generative teacher, distillation, or adversarial discrimination. It employs a distribution‑matching drift objective in a mel‑domain feature space defined by raw mels and a frozen masked‑autoencoder encoder pretrained on LJSpeech. On the LJSpeech dataset, DriftTTS achieves competitive metrics (3.87 dB MCD, 3.7% WER) and a MOS of 4.18, rivaling the Matcha‑TTS baseline and approaching ground‑truth quality.

By Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam
arXiv AI
2d ago

Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition

The paper introduces ECHO-$k$, a self-supervised, task-agnostic method for selecting which modalities to acquire at test time in multimodal, high-dimensional learning. By using a deep model’s pretrained representations as proxy targets, ECHO-$k$ learns a reinforcement‑learning policy that sequentially chooses informative modalities, providing theoretical guarantees in a linear setting. Experiments show that ECHO-$k$ consistently improves budgeted downstream performance across various foundation‑model backends, offering a principled approach to cost‑aware test‑time deployment when measurements are expensive or time‑constrained.

By Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne