Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Machine Learning
1d ago

Constant-Curvature Sliced Gromov-Wasserstein for Heterogeneous Cross-Curvature Alignment

The paper introduces Constant‑Curvature Sliced Gromov‑Wasserstein (CCSGW), a new divergence for aligning probability distributions on heterogeneous constant‑curvature spaces such as hyperbolic and spherical manifolds. It extends sliced Gromov‑Wasserstein by adding geodesic‑based one‑dimensional projections for spherical spaces, enabling efficient and principled comparison across manifolds with different curvatures while preserving intrinsic geometric relationships. The authors provide theoretical analysis showing that CCSGW controls intrinsic geometric discrepancy and demonstrate consistent performance gains when integrated into mixed‑curvature learning tasks like graph anomaly detection, node classification, and multimodal learning.

By Shanglin Li, Wenjing Lu, Muyang Li, Nicu Sebe, Ziheng Chen
arXiv Machine Learning
1d ago

Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors

Mu-DisCoCat is a multimodal variational quantum learning framework that maps Compositional Distributional Semantics (DisCoCat) onto Variational Quantum Circuits (VQCs) to achieve compositional concept generalization (CoCoGen). The pipeline first learns stable object representations from single-object image-text pairs, then fixes these to learn relations in multi-object scenarios. In classical simulations it outperformed a CLIP baseline on out‑of‑distribution relational accuracy, and on noisy quantum emulators and real IBM and IQM hardware it maintained strong fidelity correlations, reliably distinguishing unseen similar and dissimilar pairs.

By Mina Abbaszadeh, Matilda Karabina Moore, Raem Haq, Martha Lewis, Mehrnoosh Sadrzadeh
Hugging Face Trending Papers
2d ago

Mind the Accent Gap: British Accent Robustness in Speech-Driven Financial Voice Assistants

The paper highlights that AI voice assistants using ASR and LLMs struggle with regional British accents because most ASR models are trained on American English. It introduces CavaBench, a benchmark of spoken financial queries, to evaluate ASR models and their impact on downstream tool‑calling accuracy across British accents. The study finds that while WER predicts tool‑calling accuracy, it may not fully capture task‑level performance, revealing accent‑related failures that vary by model and acoustic conditions.

Hugging Face Trending Papers
2d ago

MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity

MTOR introduces a generalizable AI‑generated video detector that combines global visual features with caption‑derived textual semantics and a novel Temporal Over‑Regularity (TOR) component. The TOR module captures three levels of temporal consistency—coarse inter‑frame continuity, fine‑grained token correspondence, and frame‑to‑video stability—to exploit the stronger temporal persistence and lower variability found in synthetic videos. Extensive tests on five benchmarks with 46 generator variants show MTOR outperforms 16 baselines and remains robust against twelve real‑world video perturbations.

Hugging Face Trending Papers
2d ago

MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting

MoCAR (Motion-code Coordinate-aware AutoRegression) is a decoder‑only framework that treats continuous trajectory forecasting as next‑code prediction in a coordinate‑aware latent space. It learns a continuous motion‑code space from endpoint‑normalized trajectory segments, using historical codes as a teacher‑forced prefix and generating future codes autoregressively while updating local scene context. The method achieves top‑tier performance on Argoverse benchmarks, transfers well from AV2 to AV1 in zero‑shot evaluation, and improves on turn‑heavy scenarios without requiring trajectory‑space re‑tokenization or proposal‑and‑refinement pipelines.

Hugging Face Trending Papers
2d ago

Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation

The paper evaluates CLIP’s zero‑shot gender estimation on full‑face and periocular images from the Adience dataset. Using image‑text similarity with male/female prompts, CLIP achieves 95.54% accuracy on full faces without task‑specific training. Periocular predictions are initially biased toward males, but threshold alignment improves accuracy to 85.29%, and a linear SVM on CLIP features yields a modest 86.17% accuracy, slightly better than prior Adience results.

Hugging Face Trending Papers
2d ago

Scalable Minimal-Change Learning for Controllable Image Editing

The paper introduces ARRO, a reinforcement‑learning framework that enforces minimal‑change editing by auditing source images, instructions, and edited outputs for unimplemented and unintended changes. Using a vision‑language reward model and a group‑level rubric, ARRO achieves higher EditScore on multiple benchmarks and reduces off‑target pixel changes by 8.4% on FLUX.1 Kontext‑dev. The approach eliminates the need for per‑instruction human annotations and demonstrates strong performance across several evaluation sets.

arXiv Computer Vision
3d ago

Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild

The paper compares text‑based and feature‑based models for recognizing compound emotions in real‑world videos. It proposes textualizing non‑verbal cues from audio and visual modalities into text to leverage large language models, while feature‑based models directly combine extracted multimodal features. Experiments on the C‑EXPR‑DB dataset show that feature‑based models outperform textualization in the wild, though textual models can excel when rich transcripts are available.

By Nicolas Richet, Soufiane Belharbi, Haseeb Aslam, Meike Emilie Schadt, Manuela Gonz\'alez-Gonz\'alez, Gustave Cortal, Alessandro Lameiras Koerich, Marco Pedersoli, Alain Finkel, Simon Bacon, Eric Granger
arXiv Machine Learning
3d ago

Architecture-Dependent Fusion Pathways in MLLMs

The paper investigates how visual and textual information are fused in Multimodal Large Language Models (MLLMs). By analyzing concatenation and native multimodal architectures through alignment decoupling, attention routing, entropy, intrinsic dimensionality, and causal interventions, the authors uncover two distinct fusion pathways: concatenation models use a text‑first, vision‑later strategy, while native models integrate vision and text earlier and reorganize feature spaces. The study also employs visual CKA to test the Platonic Representation Hypothesis, offering a mechanistic view of multimodal fusion and informing architecture‑aware diagnostics.

By Hebao Zhu, Dongxia Wu
arXiv Machine Learning
3d ago

ZeroMAG: Zero-Shot Multimodal Adapter Generation for Plug-and-Play EEG Foundation Models

ZeroMAG is a zero‑shot multimodal adapter generation framework that extends frozen EEG foundation models to heterogeneous multimodal recordings using only unlabeled target data. It constructs a configuration‑invariant adapter and generates adapter weights in a latent space learned from source adapters, without target‑side optimization. Across six held‑out target datasets and three EFM backbones, ZeroMAG improves balanced accuracy by 7.22 percentage points over EEG‑only inference and 4.89 points over direct weight regression, approaching supervised multimodal adaptation.

By Yubo Wang, Jingying Ma, Xinliang Zhou, Yangxuan Zhou, Jiquan Wang, Sha Zhao, Yiyuan Yang, Yi Ding, Ziyu Jia, Chenyu Liu, Cuntai Guan
arXiv Computation and Language
3d ago

CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation

CLIMB is a training‑free inference‑time framework for multimodal retrieval‑augmented generation. It builds a compact complementary evidence pool using an MMR‑style objective that balances relevance and redundancy, then refines answers with a confidence‑controlled critic that scores relevance, specificity, and cross‑modal alignment. The method stops refinement when confidence no longer rises, improving performance on Encyclopedic‑VQA and InfoSeek without altering the retriever or language model.

By Hang Gao, Wujiang Xu, Zhixing Zhang, Kai Mei, Jingyi Yang, Dimitris N. Metaxas
arXiv AI
3d ago

VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models

VIDiff is a unified foundation model that uses diffusion techniques to perform a broad range of video tasks, including both understanding tasks like language‑guided video object segmentation and generative tasks such as video editing and enhancement. Unlike prior methods that focus on short clips and require time‑consuming tuning, VIDiff can edit and translate videos within seconds based on user instructions and employs an iterative auto‑regressive approach to maintain consistency in long‑form videos. The authors demonstrate convincing generative results across diverse input videos and written instructions, supported by qualitative and quantitative evidence.

By Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang
arXiv AI
3d ago

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

The paper introduces threat‑preserving representation sensitivity (TPRS) to assess how changes in agent‑visible wording affect attack success rates (ASR) while keeping the underlying task and evaluation criteria constant. Experiments on Agent Security Bench, MCPTox, and AgentDojo show that renaming threat‑related tools can significantly alter ASR—up to +13.21 percentage points on some models—yet may also degrade benign utility. The findings suggest that security scores based on a single representation may not generalize, urging robustness claims to be validated across multiple threat‑preserving representations.

By Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu
arXiv Computer Vision
3d ago

Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs

Gaze Attention is a new mechanism for multimodal large language models that selectively focuses on relevant visual regions during each generation step, rather than attending to all visual tokens. By grouping tokens into spatial regions and using learnable context tokens to retain global information, it reduces attention computation and visual key‑value entries by up to 90%. Experiments on 13 image and 6 video benchmarks show that Gaze Attention matches or outperforms dense‑attention baselines while using fewer visual resources.

By Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun