Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv Computation and Language
Sep 24

Causal Tracing of Audio-Text Fusion in Large Audio Language Models

The study applies causal tracing to large audio language models (LALMs) to uncover how they fuse acoustic and textual information. Layer‑wise analysis reveals distinct fusion strategies—progressive integration in DeSTA versus abrupt late‑stage fusion in Qwen—while token‑wise analysis identifies the final sequence token as an informational bottleneck that decisively retrieves audio content. Additionally, an attention‑like query mechanism at intermediate tokens is observed, prompting the model to pull task‑relevant audio context.

By Wei-Chih Chen, Chien-yu Huang, Hung-yi Lee
arXiv Computer Vision
Sep 24

AgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

AgroBench is a reproducible benchmark that converts U.S. county-level crop yield statistics into weakly supervised pixel‑level crop time series. The data generation pipeline fuses USDA yield data with land cover masks, Sentinel‑2 and Sentinel‑1 imagery, climatic variables, and terrain information to produce multimodal sequences for individual crop pixels across the growing season. The benchmark includes over 13 million observations from 788,654 crop pixels, covering 5,107 county‑year combinations for five major U.S. crops from 2017 to 2024, and establishes a Leave‑One‑Year‑Out evaluation protocol with baseline machine learning results.

By Udaiveer Singh, Rajiv Ranjan, Shashank Tamaskar, Dharmendra Saraswat
arXiv Computer Vision
Sep 24

Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents

The paper introduces Motion Vision CAPTCHA (MVCAP), a new CAPTCHA framework that relies on motion-defined foreground structures to create challenges that are only solvable through temporal analysis of a dynamic background. MVCAP is implemented in three progressive levels—coherent motion, structural motion, and biological motion—and evaluated using the MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances. Human participants achieve 99.6% accuracy, whereas the best GUI agent scores only 16.8%, highlighting a significant human–agent perception gap and demonstrating that dynamic background camouflage is the key difficulty.

By Zeyu Zhang, Dingyi Rong, Zijian Chen, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai
arXiv Computer Vision
Sep 24

CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution

CereVLA is a cerebellum-inspired framework that enhances vision‑language‑action (VLA) policies by adding lightweight residual refinement and predictive consequence evaluation to frozen VLA execution. It generates corrective actions via flow‑based residual refinement, then assesses their short‑ and interval‑horizon impacts using a recurrent state‑space model and a history‑aware classifier, suppressing unfavorable corrections with a lightweight governor. Experiments on LIBERO‑10, LIBERO‑GOAL, and SO‑101 show that CereVLA improves task success rates and reduces control steps compared to state‑of‑the‑art baselines.

By Shuai Zeng, Yuxuan Liang, Hangmiao Hu, Fobao Zhou, Zixiang Wang, Wenxi Hong, Hang Zhao
arXiv Computer Vision
Sep 24

AWM-VLA: AlignedWorld Modeling for Efficient and Explainable Vision-Language-Action Policies

AWM‑VLA introduces a unified framework that embeds aligned world modeling directly into a diffusion‑transformer vision‑language‑action policy. By adding learnable future tokens aligned with vision‑language embeddings of future observations, the policy can anticipate long‑term consequences while generating actions. The method extends this with an object‑centric alignment objective and a principled weighting scheme, achieving up to 21% higher success rates on RoboCasa and humanoid tabletop benchmarks and producing object‑centric rationales preferred by human raters in 83% of cases.

By An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian
arXiv Computer Vision
Sep 24

DMM-Align: Closed-Loop Optimization for 2D-3D Registration with Dual-Role Diffusion

DMM-Align introduces a closed‑loop framework for 2D‑3D registration that jointly refines correspondences, estimates pose, and learns representations using a shared differentiable geometric state. The method employs two diffusion processes: a geometry‑aware diffusion that improves the soft matching matrix for robust correspondence estimation, and a geometry‑conditioned diffusion teacher that feeds pose‑induced supervision back into feature learning. Experiments on 7‑Scenes and RGB‑D Scenes V2 show that DMM‑Align outperforms strong baselines, particularly in low‑overlap and heavily occluded scenarios, demonstrating the value of closed‑loop geometric feedback.

By Chongjian Wang, Junjie Gao
arXiv Computer Vision
Sep 24

Groundbench: Multi-Resolution Polygon Grounding Exposes the Geometry Gap in Vision-Language Models

Groundbench is a new benchmark that evaluates vision‑language models on multi‑resolution polygon grounding, using the same 1,500 image‑expression‑referent triples but targeting exact‑N polygons with five different vertex budgets. It audits both filled‑region intersection‑over‑union (IoU) and legal‑polygon completion, revealing that the best models achieve 88.2 box IoU and 97.1 accuracy at IoU ≥ 0.5, while direct polygon predictions lag at 57.7 and 69.2. The study shows performance is non‑monotonic across budgets, collapses at the densest budget due to legality failures, and highlights that false spatial cues hurt more than false colour cues, underscoring an operational geometry gap beyond latent boundary perception.

By Zhonghan Bian, Zhenran Wang, Jinsong Li, Zhangyang Qi
arXiv Computer Vision
Sep 24

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.

By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv Computer Vision
Sep 24

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.

By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv Computer Vision
Sep 24

AstraLOD3: Zero-shot multimodal agentic reconstruction of LOD3 building models

AstraLOD3 is a zero‑shot multimodal agentic system that reconstructs LOD3 building models using multi‑view images, calibrated cameras, a filtered sparse SfM point cloud, and a natural‑language specification. The Astra agent selects and executes computational steps in Python and Blender, achieving a mean FRDS of 0.9647 across 35 runs, including 24 benchmark buildings, with geometric agreement comparable to purpose‑built methods. Ablation studies show the impact of reconstruction guidance, evidence modalities, model configuration, and run‑to‑run variability, demonstrating that LOD3 reconstruction can be framed as a constrained agentic process rather than a fixed pipeline.

By Bryan G. Pantoja-Rosero
arXiv Computer Vision
Sep 24

A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing

The paper introduces O$^2$-VG, a unified framework for oriented object visual grounding in remote sensing images, comprising three complementary models: O$^2$-VG-Trans, a cross‑modality transformer; O$^2$-VG-Uni, which predicts universal oriented proposals; and O$^2$-VG-VLM, an autoregressive vision‑language model that generates oriented bounding boxes. It also presents DIOR‑R‑SVG, a new dataset containing image, expression, and oriented box triplets for training and evaluation. The framework demonstrates superior performance across multiple benchmarks and is supported by publicly available code.

By Zeyu Ding, Yong Zhou, Jiaqi Zhao, Wen-Liang Du, Xixi Li, Hancheng Zhu, Rui Yao, Abdulmotaleb El Saddik
arXiv Computer Vision
Sep 24

Diff-RF: Mutually Reinforced Image Registration and Fusion via Degradation-Aware Learning

Diff‑RF is a diffusion‑based framework that jointly performs image registration and fusion while accounting for degradation in multi‑modal images. It first restores modality‑specific degradations within each image, then uses a cross‑modal diffusion module that couples registration and fusion, refining alignment and enhancing complementary information. Experiments on extended datasets show that this coupled approach yields higher registration accuracy and fusion quality under diverse degraded conditions.

By Xunpeng Yi, Zaixi Du, Qinglong Yan, Yibing Zhang, Han Xu, Jiayi Ma
arXiv Computer Vision
Sep 24

MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

MultiVENT‑Raw is a new multilingual benchmark comprising nearly 120,000 raw videos—continuous footage from cell phones, hand‑held cameras, or CCTV—totaling over 5,300 hours. The dataset includes 130 events and 222 event‑centric queries, along with human‑annotated relevance judgments and extracted key facts for relevant videos. It supports two tasks: retrieving videos relevant to a query event and generating a coherent report summarizing event‑related videos for a target user, with baseline models showing these tasks remain challenging.

By Reno Kriz, David Etter, Alexander Martin, Cameron Carpenter, Debashish Chakraborty, Hannah Recknor, Reihaneh Iranmanesh, Matthew Maciejewski, Kenton Murray, Eugene Yang, Benjamin Van Durme, Aaron Steven White, Andrew Yates, William Walden
arXiv Computer Vision
Sep 24

Semantic-Guided Fusion Network for Multi-Source Remote Sensing Image Classification

The paper introduces SGFNet, a Semantic‑Guided Fusion Network for classifying multi‑source remote sensing images. It features a Semantic Mixing Convolution Block that generates semantic‑aware kernels based on contextual relationships, and a Frequency Modulated Fusion Block that fuses cross‑modal information in the frequency domain to mitigate spatial misalignment. Experiments on the Augsburg and Houston 2018 datasets show SGFNet consistently outperforms state‑of‑the‑art methods.

By Yuwei Zhao, Chuanzheng Gong, Baogui Huan, Feng Gao, Junyu Dong, Qian Du
arXiv Computer Vision
Sep 24

Copy-Move Forgery Detection and Question Answering for Remote Sensing Image

The paper introduces the Remote Sensing Copy-Move Question Answering (RSCMQA) task, aimed at interpreting complex tampering scenarios in land resource monitoring and national defense. It presents five datasets—RS-CMQA, RS-CMQA-B, Real-RSCM, RS-TQA, and RS-TQA-B—spanning 29 regions in 14 countries, and a region-discrimination-guided multimodal copy-move forgery perception framework (CMFPF) that improves question answering accuracy on tampered images. Experiments show CMFPF outperforms general VQA and RSVQA models, establishing a new benchmark for RSCMQA.

By Ze Zhang, Enyuan Zhao, Di Niu, Jie Nie, Xinyue Liang, Lei Huang
arXiv Computer Vision
Sep 24

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.

By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv AI
Sep 24

EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory

EnSIMem is an entity‑structured long‑term memory architecture designed for agents that interact with users over extended periods. It organizes interactions into theme‑coherent episodes and creates dialogue‑grounded index entries of the form [entity][entity type][property:value], preserving source turns, temporal data, and multimodal fields. During online interaction, the agent decomposes requests into evidence requirements, performs entity‑property lookup, and retrieves the necessary evidence to generate responses directly from preserved source material rather than lossy summaries.

By Xuanyu Meng, Xing Fan, Xinyi Fan, Chenlei Guo, Yixuan Xie, Jiawei Han
arXiv AI
Sep 24

A Leakage-Aware Multimodal Evaluation Framework for Early Intraoperative Acute Kidney Injury Prediction

The study introduces SynerT, a waveform-only hybrid temporal model that uses a causal dilated TCN and dilated recurrent layers to predict early intraoperative acute kidney injury (AKI). Two extensions, SynerT-MM and SynerT-Stack, incorporate hemodynamic summaries, preoperative covariates, and a leakage-safe stacked ensemble to improve discrimination and calibration. Evaluated on the VitalDB database with a strict 60‑minute prediction window, SynerT-Stack achieved the best performance across AUROC, AUPRC, and F1‑max, and demonstrated the greatest net clinical benefit after recalibration.

By Quang Minh Nguyen, Duc Minh Le, Ho Nhat Minh Nguyen, Thuy Quynh Nguyen, Trong Nghia Nguyen
arXiv Computation and Language
Sep 24

NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task

NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification shared task series and the second focused on multidialectal Arabic speech processing. It includes five main tasks—Automatic Speech Recognition, Spoken Dialect Identification, Text-to-Speech, Spoken Language Translation, and Spoken Language Understanding—along with eight subtasks that test realistic scenarios such as low‑bandwidth, mixed dialects, code‑switching, out‑of‑domain, and zero‑shot settings. The event attracted 21 teams from at least 13 countries, with 48 test‑phase submissions and 14 system‑description papers, and the results highlight out‑of‑domain generalization as a major bottleneck while showcasing the strengths of Arabic‑specialized speech models, multimodal dialect identification, and ensemble methods.

By Peter Sullivan, Bashar Talafha, Ahmed Ashraf, Fethi Bougares, Haroun Elleuch, Chiyu Zhang, AbdelRahim Elmadany, Youssef Mohamed, Salima Mdhaffar, Yannick Est\`eve, Mohamed Elhoseiny, Hamzah Luqman, Nizar Habash, Muhammad Abdul-Mageed
arXiv Computation and Language
Sep 24

When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR

SpeechLLMs used in professional settings often undergo domain customisation through prompts or fine‑tuning, which can inadvertently cause the model to transcribe phonetically similar words from its context or training data, leaking private information. The authors systematically investigate this overlooked privacy risk, creating benchmarks to measure leakage rates for both prompting and fine‑tuning, and find that both mechanisms cause measurable leakage that compounds when combined. They evaluate a prompt‑level mitigation strategy and analyse the accuracy‑leakage trade‑off, concluding that fine‑tuning without context prompts offers the best balance between performance and privacy.

By Maike Z\"ufle, Jan Niehues