Multimodal models
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
Mind the Accent Gap: British Accent Robustness in Speech-Driven Financial Voice Assistants
The paper highlights that AI voice assistants using ASR and LLMs struggle with regional British accents because most ASR models are trained on American English. It introduces CavaBench, a benchmark of spoken financial queries, to evaluate ASR models and their impact on downstream tool‑calling accuracy across British accents. The study finds that while WER predicts tool‑calling accuracy, it may not fully capture task‑level performance, revealing accent‑related failures that vary by model and acoustic conditions.
MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity
MTOR introduces a generalizable AI‑generated video detector that combines global visual features with caption‑derived textual semantics and a novel Temporal Over‑Regularity (TOR) component. The TOR module captures three levels of temporal consistency—coarse inter‑frame continuity, fine‑grained token correspondence, and frame‑to‑video stability—to exploit the stronger temporal persistence and lower variability found in synthetic videos. Extensive tests on five benchmarks with 46 generator variants show MTOR outperforms 16 baselines and remains robust against twelve real‑world video perturbations.
MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting
MoCAR (Motion-code Coordinate-aware AutoRegression) is a decoder‑only framework that treats continuous trajectory forecasting as next‑code prediction in a coordinate‑aware latent space. It learns a continuous motion‑code space from endpoint‑normalized trajectory segments, using historical codes as a teacher‑forced prefix and generating future codes autoregressively while updating local scene context. The method achieves top‑tier performance on Argoverse benchmarks, transfers well from AV2 to AV1 in zero‑shot evaluation, and improves on turn‑heavy scenarios without requiring trajectory‑space re‑tokenization or proposal‑and‑refinement pipelines.
Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation
The paper evaluates CLIP’s zero‑shot gender estimation on full‑face and periocular images from the Adience dataset. Using image‑text similarity with male/female prompts, CLIP achieves 95.54% accuracy on full faces without task‑specific training. Periocular predictions are initially biased toward males, but threshold alignment improves accuracy to 85.29%, and a linear SVM on CLIP features yields a modest 86.17% accuracy, slightly better than prior Adience results.
Scalable Minimal-Change Learning for Controllable Image Editing
The paper introduces ARRO, a reinforcement‑learning framework that enforces minimal‑change editing by auditing source images, instructions, and edited outputs for unimplemented and unintended changes. Using a vision‑language reward model and a group‑level rubric, ARRO achieves higher EditScore on multiple benchmarks and reduces off‑target pixel changes by 8.4% on FLUX.1 Kontext‑dev. The approach eliminates the need for per‑instruction human annotations and demonstrates strong performance across several evaluation sets.
Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild
The paper compares text‑based and feature‑based models for recognizing compound emotions in real‑world videos. It proposes textualizing non‑verbal cues from audio and visual modalities into text to leverage large language models, while feature‑based models directly combine extracted multimodal features. Experiments on the C‑EXPR‑DB dataset show that feature‑based models outperform textualization in the wild, though textual models can excel when rich transcripts are available.
Architecture-Dependent Fusion Pathways in MLLMs
The paper investigates how visual and textual information are fused in Multimodal Large Language Models (MLLMs). By analyzing concatenation and native multimodal architectures through alignment decoupling, attention routing, entropy, intrinsic dimensionality, and causal interventions, the authors uncover two distinct fusion pathways: concatenation models use a text‑first, vision‑later strategy, while native models integrate vision and text earlier and reorganize feature spaces. The study also employs visual CKA to test the Platonic Representation Hypothesis, offering a mechanistic view of multimodal fusion and informing architecture‑aware diagnostics.
ZeroMAG: Zero-Shot Multimodal Adapter Generation for Plug-and-Play EEG Foundation Models
ZeroMAG is a zero‑shot multimodal adapter generation framework that extends frozen EEG foundation models to heterogeneous multimodal recordings using only unlabeled target data. It constructs a configuration‑invariant adapter and generates adapter weights in a latent space learned from source adapters, without target‑side optimization. Across six held‑out target datasets and three EFM backbones, ZeroMAG improves balanced accuracy by 7.22 percentage points over EEG‑only inference and 4.89 points over direct weight regression, approaching supervised multimodal adaptation.
CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation
CLIMB is a training‑free inference‑time framework for multimodal retrieval‑augmented generation. It builds a compact complementary evidence pool using an MMR‑style objective that balances relevance and redundancy, then refines answers with a confidence‑controlled critic that scores relevance, specificity, and cross‑modal alignment. The method stops refinement when confidence no longer rises, improving performance on Encyclopedic‑VQA and InfoSeek without altering the retriever or language model.
VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models
VIDiff is a unified foundation model that uses diffusion techniques to perform a broad range of video tasks, including both understanding tasks like language‑guided video object segmentation and generative tasks such as video editing and enhancement. Unlike prior methods that focus on short clips and require time‑consuming tuning, VIDiff can edit and translate videos within seconds based on user instructions and employs an iterative auto‑regressive approach to maintain consistency in long‑form videos. The authors demonstrate convincing generative results across diverse input videos and written instructions, supported by qualitative and quantitative evidence.
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
The paper introduces threat‑preserving representation sensitivity (TPRS) to assess how changes in agent‑visible wording affect attack success rates (ASR) while keeping the underlying task and evaluation criteria constant. Experiments on Agent Security Bench, MCPTox, and AgentDojo show that renaming threat‑related tools can significantly alter ASR—up to +13.21 percentage points on some models—yet may also degrade benign utility. The findings suggest that security scores based on a single representation may not generalize, urging robustness claims to be validated across multiple threat‑preserving representations.
Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs
Gaze Attention is a new mechanism for multimodal large language models that selectively focuses on relevant visual regions during each generation step, rather than attending to all visual tokens. By grouping tokens into spatial regions and using learnable context tokens to retain global information, it reduces attention computation and visual key‑value entries by up to 90%. Experiments on 13 image and 6 video benchmarks show that Gaze Attention matches or outperforms dense‑attention baselines while using fewer visual resources.
Uncertainty Quantification for Flow-Based Generalist Robot Policies
The paper introduces a method for quantifying epistemic uncertainty in flow‑matching based generalist robot policies, such as vision‑language‑action models and world‑action models. By measuring velocity‑field disagreement across a small ensemble, the authors obtain better‑calibrated uncertainty estimates that can detect deployment failures and guide active fine‑tuning. Their SAVE approach reduces the need for expert demonstrations, improving real‑world task success from 39 % to 47 % while maintaining a fixed demonstration budget.
THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS
The paper introduces THPL, a vision-to-language framework for precision feeding of rainbow trout in recirculating aquaculture systems. It uses Fishsort to extract feeding trajectories, a Hierarchical Behavior Encoder to model individual and collective dynamics, and integrates these with environmental data and expert rules to fine‑tune a large language model via LoRA and counterfactual multimodal Direct Preference Optimization. Experimental results show strong correlation between the Activity Coefficient and expert feeding intensity, and significant improvements in decision accuracy and language metrics when using dual‑evidence tokens and counterfactual optimization.
Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination
The paper introduces IG‑VLA, a vision‑language‑action framework that learns to imagine task‑relevant future scene evolution in a latent spatiotemporal space, guiding action prediction without generating full pixel‑level videos. It further compresses this future reasoning into a compact Scene Gist Token via Scene Gist Memory, enabling efficient inference. Experiments on LIBERO, LIBERO‑Plus, and VLABench show that IG‑VLA improves success rates by nearly 6% and speeds up inference up to 6.38× compared to strong baselines.
Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration
Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration proposes CDPM, a method that first aligns semantic representations across modalities and then refines correspondences with fine-grained CNN features. CDPM adapts DINOv3 using geometrically consistent cross-modal patch pairs, builds a DINO-Centric Feature Pyramid for stable cross-modal matching, and adds a lightweight CNN branch for precise local refinement. Experiments on three cross-modal datasets show that CDPM outperforms existing dense matchers, improving AUC metrics and reducing mACE while using fewer FLOPs.
UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation
UniDynamics is a diffusion-based framework that generates future 4D dynamic scenes—comprising RGB, depth, and optical flow—from a single event-RGB pair, without needing long histories or control priors. It introduces an Event Latent Enhancement module to align event data into conditioning features and a Perceptual Dynamics Space within a multi-scale U‑Net to decouple and adaptively interact depth and flow, enforcing geometric and motion constraints for coherent predictions. Experiments on VKitti2 and DSEC show state‑of‑the‑art performance, producing high‑quality, temporally coherent, and 4D‑consistent predictions even under high‑speed motion blur.
DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
DriftTTS is a few‑step neural text‑to‑speech model that generates mel‑spectrograms without relying on a generative teacher, distillation, or adversarial discrimination. It employs a distribution‑matching drift objective in a mel‑domain feature space defined by raw mels and a frozen masked‑autoencoder encoder pretrained on LJSpeech. On the LJSpeech dataset, DriftTTS achieves competitive metrics (3.87 dB MCD, 3.7% WER) and a MOS of 4.18, rivaling the Matcha‑TTS baseline and approaching ground‑truth quality.
Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition
The paper introduces ECHO-$k$, a self-supervised, task-agnostic method for selecting which modalities to acquire at test time in multimodal, high-dimensional learning. By using a deep model’s pretrained representations as proxy targets, ECHO-$k$ learns a reinforcement‑learning policy that sequentially chooses informative modalities, providing theoretical guarantees in a linear setting. Experiments show that ECHO-$k$ consistently improves budgeted downstream performance across various foundation‑model backends, offering a principled approach to cost‑aware test‑time deployment when measurements are expensive or time‑constrained.