Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv Computation and Language
Sep 24

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

The paper introduces a trilingual spoken hallucination detection benchmark covering English, Russian, and Kazakh news, with 12,013 samples that include synthetic alterations and severity levels, as well as 290 fact‑checked misinformation items. Detectors are evaluated in a reference‑free setting on text, ASR transcripts, and audio, revealing that most models underperform baseline classifiers, except Gemma‑3n on transcripts. Synthetic‑trained detectors achieve high macro‑F1 scores on real‑world misinformation, but Russian provenance analysis highlights model‑dependent signals that confound synthetic benchmarks.

By Meruyert Aristombayeva, Jason S. Lucas, Chaewan Chun, Dongwon Lee
arXiv Computer Vision
Sep 24

Gender Bias in Vision-Language In-Context Learning

The paper investigates how in‑context learning (ICL) in large vision‑language models (LVLMs) can amplify gender bias. Using the VL‑BICLE framework, the authors show that gendered ICL demonstrations shift model bias toward the demonstrated gender, especially in tasks involving gendered language such as image captioning and pronoun prediction. They find that similarity‑based retrieval does not mitigate this bias and that replacing real images with synthetic ones from stable diffusion reduces bias without hurting caption quality.

By Tong Xiang, Noa Garcia, Yuta Nakashima
arXiv AI
Sep 24

BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport

BiCFlow-MER introduces a conditional-flow framework for audio-text multimodal emotion recognition, treating the task as generative evidence transport within a structured emotion space. It disentangles emotion-oriented evidence from speaker style and lexical content, creating a conflict-aware affective condition that guides bidirectional rectified flow to an explicit emotion-space endpoint. The model verifies candidate emotions via adaptive prototype-cloud scoring and backward class-to-condition consistency, achieving superior performance on IEMOCAP, MELD, and the zero-shot CASE benchmark.

By Yanbing Wang, Shenyue Wang, Chunyang Yu
arXiv AI
Sep 24

SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference

SlackDrive is a pre‑inference compute allocator that dynamically selects the compute budget for each driving control step by reusing the realized latency from previous inferences. By profiling a small set of discrete budgets once, it estimates the current compute state online and chooses the highest‑utility budget that stays within the admissible latency envelope. On the NAVSIM v2 benchmark with DriveDreamer‑Policy, SlackDrive boosts latency‑constrained EPDMS performance by 21.7% compared to the best baseline, while full‑budget and token‑pruning approaches exceed the latency limits under runtime contention.

By Xiaohuan Pei, Hengguang Zhou, Yuanhao Ban, Justin Cui, Jiaqi Feng, Haoyu Xie, Tao Huang, Pichao Wang, Yanchao Yang, Cho-Jui Hsieh
arXiv AI
Sep 24

Cross-Modal Contrastive Learning from Histopathology and CT for Automated Renal Cell Carcinoma Grading

The study introduces RCC-Align, a cross‑modal contrastive learning framework that aligns paired histopathology whole‑slide images and CT scans to enhance noninvasive grading of clear cell renal cell carcinoma (ccRCC). Using patient‑level five‑fold cross‑validation on TCGA and CPTAC cohorts, RCC‑Align achieved an AUC of 0.601 and AUPRC of 0.599 for low‑ versus high‑grade ccRCC classification, outperforming CT‑only baselines and showing stronger WSI‑CT embedding alignment. The approach relies solely on CT at inference, potentially aiding grading when biopsy is unsafe or limited by tumor heterogeneity.

By Amit Das, Tanmay Shukla, Naofumi Tomita, Faraz Farhadi, Jessica Sin, Ari Hakimi, Chad Vanderbilt, Jie-Fu Chen, Ritesh Kotecha, Weijie Ma, Bing Ren, Saeed Hassanpour
arXiv AI
Sep 24

Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR

The study investigates how verbal attunement and real‑time behavioral mimicry affect users’ perceptions of an embodied AI counselor in virtual reality. Participants interacted with a system that varied in verbal attunement (attuned vs. neutral) and behavioral mimicry (present vs. absent). Results indicated that verbal attunement most reliably increased perceived empathy, while mimicry had a marginal effect on perceived humanness and showed exploratory positive associations with empathy, positivity, and humanness, especially among female participants.

By Nathalia Gomez, Haig Shamlian, Omar Khan, Tiffany D. Do
arXiv AI
Sep 24

Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition

Ruby-ASR introduces an evidence-preserving supervision method for Japanese automatic speech recognition that refines the conventional orthographic target into a span-bound orthographic–lexical-reading sequence. This ruby representation locally binds each written span to its spoken reading, enabling deterministic recovery of both orthographic and lexical views. Experiments on five Japanese benchmarks show that this refined target improves lexical-reading recovery while maintaining readable orthographic transcription.

By Hao Shi, Yun Liu, Xuehao Yang, Jun Liu, Chuanbo Hua, Xuanjun Chen, Lianbo Liu, Shiao Zhu, Zixiong Su
arXiv AI
Sep 24

What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit

The paper argues that apparent capability limits in vision‑language benchmarks often stem from the way answers are presented rather than from the models themselves. By comparing performance on COCO images with answer choices given as English names versus pixel coordinates, the authors show that models like Qwen3‑VL‑4B perform far better when answers are in natural language, and that the choice of answer format can swing model rankings by dozens of points. The study also demonstrates that different conventions (e.g., hue angles vs. pixel coordinates) reveal which formats a model can actually interpret, highlighting that a fixed answer vocabulary is not neutral across models.

By Alfredo F. Frontera Del Valle
arXiv AI
Sep 24

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

InfiNoVA is a data‑augmentation framework that transforms synchronized multi‑camera demonstrations into a dense, geometrically consistent set of training views by reconstructing each manipulation trajectory as a time‑varying 3D Gaussian. The method renders novel observations from sampled camera poses while preserving the original state‑action pairs, improving frame‑level fidelity and temporal consistency compared to generative synthesis. Across four real‑world manipulation tasks, policies trained with InfiNoVA achieve 5.4× higher average success under unseen randomized viewpoints than VISTA‑based augmentation and 1.7× higher success than training on all five physical camera views.

By Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave
arXiv AI
Sep 24

RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction

RelCheck is a training‑free post‑hoc correction pipeline that addresses relational hallucinations in multimodal large language models. It augments object‑level visual grounding with two forms of relational evidence—learned scene‑graph triples from RelTR and deterministic spatial predicates derived from bounding‑box geometry—forming a three‑layer visual knowledge base. When applied to LLaVA v1 13B, RelCheck improves the overall MME hallucination score from 585.0 to 630.0, with the most significant gain on spatial position accuracy.

By Siddhi Patil, Navrati Saxena, William B. Andreopoulos
arXiv AI
Sep 24

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

MemBodied introduces a fixed‑size episodic memory for Vision‑Language‑Action models, comprising an associative state that tracks interactions across policy calls and an episode anchor that stores a compact representation of the initial scene. By conditioning action generation on these memory components instead of raw past observations, MemBodied reduces context bloat and inference latency. In five memory‑dependent RMBench tasks, it outperforms stateless and vanilla recurrent policies by significant margins, and achieves a 90.6% success rate on the LIBERO‑Long suite, improving over the baseline by 5.4%.

By Tej Deep Pala, Navonil Majumder, Bryce Goh, Raphael Yee, Jianfei Yang, Liming Chen, Soujanya Poria
arXiv AI
Sep 24

Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction

The paper introduces a language‑guided approach for robots to join human groups by predicting socially compliant joining poses. It uses recursive spectral partitioning to generate candidate group subsets, ranks them with a language‑conditioned image–geometry model, and then applies a goal predictor that incorporates human‑formation priors to produce a multimodal energy–orientation map of feasible robot poses. Experiments on various group scenarios show competitive grounding accuracy, superior joining‑pose prediction, and successful real‑robot demonstrations in both static and dynamic settings.

By Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu
arXiv Computer Vision
Sep 24

Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection

The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.

By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
arXiv AI
Sep 24

Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach

The paper introduces UGTPhon, a grapheme-to-phoneme benchmark for user‑generated text in English, Vietnamese, and Korean, and presents a taxonomy for diagnosing pronunciation errors. It shows that existing G2P models and large language models struggle with canonical‑to‑non‑canonical text, with errors up to 66.8 PER points. A compositional G2P approach that uses exact‑match lookup and staged decoding reduces these errors and performs competitively with larger few‑shot LLMs.

By MinJu Jeon, Younghan Park, Han Sung Park, Jong-Hwan Kim, Dong-Jin Kim, Hoyeon Lee
arXiv AI
Sep 24

Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech

The paper introduces GUARD, a lightweight speaker identity unlearning framework designed to prevent re-identification in zero-shot text-to-speech systems. GUARD employs a learned speaker gate and speaker-agnostic activation steering on a frozen TTS backbone, optimizing steering vectors through group-relative reward optimization to reduce similarity to forgotten speakers while maintaining intelligibility and naturalness. Experiments on CosyVoice2 show that GUARD significantly lowers forget-speaker similarity and re-identification accuracy while preserving the ability to reproduce retained speakers.

By Hyoeun Kim, Yujun Lee, Kyuhong Shim
arXiv AI
Sep 24

Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning

The paper introduces FedMAST, a Federated Multi‑Axis Structural Tracing defense designed to detect and contain backdoor attacks in federated learning. FedMAST evaluates client updates through complementary structural, spectral, and historical evidence, applying tiered filtering and round‑level containment. In experiments across six backdoor attacks, FedMAST consistently achieves lower attack success rates while preserving high main‑task accuracy.

By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
arXiv AI
Sep 24

AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

arXiv:2609. 28366v1 Announce Type: cross Abstract: Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning.

By Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que, Chence Yang, Geng Yuan, Qianwen Li
arXiv Machine Learning
Sep 24

Confidence Falls Short: Asymmetric Certainty Gains from Optimization Hinder Multimodal Classification

The paper identifies that in multimodal learning, optimization often produces asymmetric certainty gains, with the stronger modality becoming more confident than the weaker one, which leads to imbalanced contributions and suboptimal performance. The authors attribute this issue to unimodal characteristics and propose a Max Confidence Regularization (MaxCR) method that tracks each modality’s semantic confidence via a nonlinear sparsity measure and applies max suppression and excitation to balance confidence levels. Experiments on standard datasets demonstrate that MaxCR improves overall performance compared to state‑of‑the‑art multimodal baselines.

By Longfei Huang, Xiangyu Wu, Yang Yang
arXiv Computer Vision
Sep 24

Task-Prototype Guided Flow Matching for Few-Shot Generalization in Vision-Language Robot Manipulation

Task-Prototype Guided Flow Matching (TP-Flow) is a few‑shot manipulation framework that transforms support demonstrations into structured task‑prototype tokens to guide both the initial flow prior and the velocity field. It uses symmetric cross‑attention with learnable queries to extract phase‑level prototypes, parameterizes a task‑adaptive initial distribution, and injects prototype information through gated adaptive normalization. TP‑Flow is trained with an episodic support‑query objective and prototype contrastive regularization, achieving high success rates on the LEROBOT‑ARM‑SO101 platform while maintaining real‑time execution and low latency.

By Yizhao Wang, Guantao Zhang, Jingbo Wang
arXiv AI
Sep 24

PhyMo: A Physical-Field Modality for Multimodal AI4Physics

The paper introduces PhyMo, a physics‑grounded multimodal framework that uses a physical‑field modality to represent heterogeneous measurements via PDE‑associated operators. It follows a three‑stage learning process: pretraining a physical‑field encoder with PDE residual supervision, aligning its representations with visual embeddings in a shared latent space, and applying downstream prediction heads to the fused multimodal representations. Experiments on five diverse physical datasets show that PhyMo outperforms the strongest baseline on each dataset, establishing its effectiveness for multimodal representation learning in AI for Physics.

By Henan Sun, Haitao Hu, Jin Liu, Jianfeng Zhang, Lujia Pan, Nuo Chen, Jia Li