Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Statistics ML
6d ago

Pragmatic DML with AI-Learned Representations

The paper investigates the validity of using AI‑learned representations—such as compressed text, images, and other covariates—as controls in causal analysis. It shows that for a wide class of estimands, the bias introduced by imperfect representations is the product of two errors: one in the outcome regression and one in the balancing weight (or Riesz representer). The authors provide three key results: (1) cross‑fitted double machine learning yields valid Wald inference for the representation‑dependent target and can achieve semiparametric efficiency when representation errors are small; (2) fold‑wise representation learning (fine‑tuning) can be combined with DML inference using convex‑ and star‑aggregation pipelines; and (3) when representation errors are large, the framework offers interpretable sensitivity regions and root‑n inference for their endpoints. In a multi‑modal demand study, the approach produced consistent negative near‑unit elasticity estimates across seven representation‑specific models and their star aggregate, remaining robust across the sensitivity grid.

By Andres Aradillas Fernandez, Victor Chernozhukov, Carlos Cinelli, Sven Klaassen, Whitney Newey, Martin Spindler, Jan Teichert-Kluge, Suhas Vijaykumar
arXiv AI
6d ago

SpikeMoE: Brain-Inspired Competitive Routing for Flexible Spiking Mixture-of-Experts

SpikeMoE introduces a spike-based k‑WTA router that uses lateral inhibition and refractory periods to select the top‑K experts based on discrete spike counts, inspired by hippocampal CA1 competition. The framework combines spiking neural network dynamics with mixture‑of‑experts conditional computation and adds a two‑stage missing‑modality module for robust multimodal processing. Experiments on vision, language, and multimodal tasks show that SpikeMoE matches or surpasses ANN baselines while offering energy‑efficient performance.

By Xiaoli Liu, Yujie Liang, Jialin Li, Malu Zhang
arXiv AI
6d ago

CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

CoEvolve is a construct-to-edit framework for visual grounding that separates the task into explicit state construction and state editing. It uses Region‑Evolution Reinforcement to progressively refine candidate regions and Bidirectional Denoising Refiner to adjust coordinate fields based on fixed semantic context. The approach achieves high grounding accuracy with a 9B backbone, rivaling much larger models, and can recover over 27 percentage points in mean box overlap after a single refinement pass.

By Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao
arXiv AI
6d ago

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.

By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
arXiv Computer Vision
6d ago

Discrete Annotation, Continuous Preference: Rethinking Supervision for Accurate and Generalizable Aesthetic Image Cropping

The paper introduces a new approach to aesthetic image cropping by modeling human preference as a continuous, multi-peaked field rather than relying on discrete, grid‑based annotations. It presents the Continuous Preference Field (CPF) that reconstructs a dense preference landscape from sparse labels, and uses this to train a VLM‑based cropping model (CPIC) that achieves state‑of‑the‑art accuracy and strong out‑of‑domain generalization. Additionally, the authors propose CPICD, a recalibrated benchmark that corrects grid‑bound artifacts in existing datasets, providing a more reliable evaluation framework.

By Ziqing Zhang, Xiao Liu, Kai Liu, Jianze Li, Weihang Zhang, Linghe Kong, Yulun Zhang
arXiv AI
6d ago

Backdoor Containment via Expert Quarantine and Shutdown in LLMs

The paper introduces Quarantined Expert Shutdown (QES), a new backdoor containment strategy for large language models. QES allows backdoor learning to occur during training but routes it into a designated, quarantined expert that can be disabled at deployment. The method achieves significant reductions in attack success rates while largely preserving model utility.

By Jianwei Li, Min-Seon Kim, Jung-Eun Kim
arXiv AI
6d ago

DeFA: Dependency-Guided Failure Attribution for LLM Agents

DeFA is a dependency-guided framework that attributes failures in large language model agents by constructing an event dependency graph and a failure propagation graph from protocol relations and semantic dependencies. It identifies violating events, traces their sources and effects, and determines the decisive error, responsible agent, and error category. The method supports long trajectories through segmentation and has shown superior accuracy on text, image, and video tasks, while its diagnostic feedback can improve agent performance on subsequent tasks.

By Bo Deng, Xinlei Zheng, Yi Wei, Kang Zhou, Chongyang Tao, Renzhao Liang, Xuanren Chen, Lifan Guo, Chi Zhang
arXiv AI
6d ago

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

The paper introduces TRACE, a token‑level objective designed to reduce multi‑turn safety risks in large language models. TRACE assigns each token a weight based on the discounted return of a refusal‑attributable advantage, comparing a frozen reference model with a refusal‑ablated copy to credit early tokens for later refusal evidence. Evaluated across five open‑weight models and seven multi‑turn attacks, TRACE achieves the lowest attack success rate in all 35 model‑attack pairs while maintaining model utility within 1.23 points on MMLU and HellaSwag.

By Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang
arXiv AI
6d ago

PRISM: A Category-Theoretic Framework for Measuring and Refining Multimodal Analogies

PRISM is a modality‑agnostic, category‑theoretic framework that measures and refines multimodal analogies by representing them as explicit relational mappings. It introduces a pullback score to quantify relational alignment and an iterative refinement loop that uses this score as feedback to improve generated images. On the AnaloBench benchmark, PRISM’s pullback score alone achieves 82.5% accuracy, and human evaluations show a 57.65% preference for refined outputs, though refinement may sometimes favor visually crowded compositions.

By Mirella Zeisler, Ojas Shirekar, Mircea Lic\v{a}, Chirag Raman
arXiv AI
6d ago

Towards Reliable Vision-Language Models for Autonomous Driving

The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.

By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
arXiv AI
6d ago

LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification

LineupRL introduces a reinforcement learning framework with verifiable rewards for time series captioning, using a frozen large language model to identify the correct time series from a set of distractors based on a generated caption. This approach bypasses the limitations of supervised fine‑tuning and traditional RL rewards that poorly transfer to open‑ended time series generation. Experiments on two captioning benchmarks, as well as forecasting and reconstruction tasks, show that LineupRL outperforms both SFT and RL baselines across all metrics, and its trained 3B vision‑language model surpasses a 72B model distilled from SFT captions. The method also demonstrates resistance to reward hacking and produces captions that accurately trace trends and name key values.

By Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen
arXiv AI
6d ago

AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

AVSD-Scenes is a new dataset of 12,291 audio‑visual scene descriptions for urban environments, built from the TAU Urban Audio‑Visual Scenes dataset. The descriptions are generated by first creating modality‑specific text with Qwen2‑Audio‑7B and Qwen2.5‑VL‑7B, then merging them with large language models (Qwen3‑14B, Mistral‑Small‑3.2‑24B‑Instruct‑2506, Gemma‑3‑27B‑it) to produce multimodal narratives that combine auditory and visual cues. Benchmarks show that these multimodal descriptions improve semantic alignment, cross‑modal retrieval, and scene classification accuracy (up to 95.4%) while remaining discriminative even without explicit scene labels.

By Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley
arXiv AI
6d ago

SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL

SPHERE is an adaptive VR indoor scene generation framework that turns isolated 3D synthesis into continuous human‑AI co‑creation. It learns persistent spatial preferences from multimodal user interactions, abstracts these into hierarchical constraints for geometric resilience, and employs a human‑in‑the‑loop reinforcement learning loop to refine retrieval policies. A mixed‑design study with 42 participants and offline ablation show that SPHERE reduces corrective edits and physical effort while avoiding bias toward shallow object‑level traits, producing geometrically resilient, profile‑aligned layouts.

By Hyeonmin Lee, Zheng Wei, Kyungmin Kwon, Jumin Seo, Jiwon Park, Hayoung Oh
arXiv AI
6d ago

VISTA: A Visual Harness for Reasoning in an Interactive World

The paper introduces VISTA, a visual harness that equips a general-purpose multimodal model with long‑horizon vision and a lossless visual memory. VISTA enables the model to directly perceive and actively retrieve past observations, allowing it to reorganize visual input during reasoning. On the ARC‑AGI‑3 benchmark, VISTA boosts Claude Opus 5.0’s Relative Human Action Efficiency from 40.68 to a perfect 100.00, completing all 25 public games with 57.4% fewer actions than first‑time human participants, and it also outperforms baselines on three additional visual game and puzzle benchmarks.

By Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He
arXiv AI
6d ago

Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing

The paper presents a billing‑aware neural text‑to‑speech system for serverless CPUs, focusing on minimizing CPU‑seconds and GB‑seconds rather than just throughput or latency. By using request‑sized concurrent inference and a reclaimable instance lifecycle, the system limits per‑request CPU parallelism and releases idle memory while keeping the server process alive. On the Kokoro‑82M benchmark, it achieves 2.71 audio‑seconds per CPU‑second versus 0.89 with ONNX Runtime, cuts cost per audio‑hour from $0.0631 to $0.0153, and reduces idle billed memory from 8.7 GB to 1.33 GB, with faster restoration times.

By Pakorn Nathong, Kunat Pipatanakul
arXiv AI
6d ago

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.

By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv AI
6d ago

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.

By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
arXiv AI
6d ago

VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision

VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.

By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
arXiv Computer Vision
6d ago

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

PhysVista is a new benchmark that evaluates physical intelligence in Vision‑Language Models (VLMs) by integrating perception, reasoning, and plausibility assessment into a closed cognitive loop. It distinguishes between event‑level and scale‑level reasoning and tests models on both real‑world and AI‑generated videos to provide a holistic, fine‑grained analysis of physical understanding. Experiments show significant gaps in VLMs’ physical reasoning and plausibility assessment, underscoring the need for more principled, physically grounded multimodal designs.

By Xinge Peng, Yiting Lu, Tianwu Zhi, Wen Wen, Jianzhao Liu, Xin Li, Zhibo Chen
arXiv Computer Vision
6d ago

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

Omni-Embed-Mini is a 0.9B‑parameter model that embeds text, speech, audio, images, video, and visually‑rich documents into a single shared cosine space without updating any text‑side parameters. It uses a dense cascaded caption as a teacher signal, allowing the teacher and student to share identical backbone weights and requiring only lightweight projectors and phased LoRA adapters for alignment. The model achieves strong text retrieval performance (49.57 nDCG@10 on MTEB‑v2 BEIR‑8) while extending to five additional modalities and is significantly smaller than other open omni‑modal embedders.

By Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal