Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv AI
3d ago

RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction

RASPER is a reward‑aligned summarizer that tailors the extraction of information from unstructured discharge notes to improve downstream clinical predictions. It uses a tunable LLM summarizer trained with reinforcement learning, where the reward comes from the loss of a downstream predictor, and incorporates patient‑specific context via a longitudinal encoder that soft‑prompts the summarizer with structured codes. The approach consistently outperforms strong baselines on readmission prediction and medication recommendation tasks in the MIMIC‑III and MIMIC‑IV datasets.

By Arya Hadizadeh Moghaddam, Mohsen Nayebi Kerdabadi, Chen Chen, Dongjie Wang, Zijun Yao
arXiv AI
3d ago

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

This paper introduces a personalized automatic speech recognition system for a Czech speaker with a permanent tracheal stoma and severe dysarthria. The authors release a 33‑hour annotated dataset collected via an artificial conversation protocol and develop a multi‑stage training pipeline based on Whisper Base, fine‑tuning on Czech speech, simulated tracheostomic speech, and the speaker’s data. Evaluations in scripted, question‑answering, and spontaneous dialogue scenarios show a 50 % relative reduction in character error rate compared to the Whisper Base baseline and better accuracy than the speaker’s assistants on isolated utterances.

By David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus
arXiv Computation and Language
3d ago

An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

The paper introduces an automated pipeline that extracts conversational turns and backchannels from separate-channel recordings of spontaneous dyadic dialogue, combining voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post‑processing. Evaluated on 99 ten‑minute Danish conversations, the system achieved F1 scores around 0.62 for both turns and backchannels, with median onset/offset errors of roughly 0.15–0.18 s. Performance was consistent across normal and asymmetric listening conditions, and a case study showed the pipeline’s outputs were less variable than human annotations, supporting its use as a reliable first‑pass annotation tool in semi‑automated workflows.

By Hanlu He, Harald Vilhelm Skat-R{\o}rdam, Ingvi \"Orn\'olfsson, Ivana Konvalinka
arXiv Computation and Language
3d ago

Social bot detection in the age of ChatGPT: Challenges and opportunities

The article reviews the challenges and opportunities of detecting social bots amid the rise of advanced AI chatbots. It highlights gaps in current detection methods, especially regarding AI-generated conversations, and identifies emerging trends such as synthetic data generation, multimodal cross‑platform detection, low‑resource language support, and federated learning approaches. The authors propose these directions as promising avenues for future research.

By Emilio Ferrara
arXiv Computation and Language
3d ago

World Embedding Benchmark

The World Embedding Benchmark introduces 8,000 simulation-based video cases covering fluid, solid, dynamic, and optical physics, each paired with physical annotations. It supports three tasks—text‑video retrieval, physical‑property regression, and multiple‑choice classification—to assess how well video embeddings capture physical alignment versus quantitative information. Experiments show that while pre‑trained models perform poorly on retrieval and classification, lightweight probes can extract useful physical data, and physics‑specific contrastive training improves alignment but harms regression, highlighting a trade‑off. Retrieval‑augmented generation using these embeddings further enhances the physical fidelity of generated videos.

By Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li, Fengyu Cai, Hou Pong Chan, Jialin Yu, Hao Zhang, Chenghua Lin, Chenghao Xiao
arXiv Computation and Language
3d ago

BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions

BioMol-MQA is a new question‑answering dataset focused on polypharmacy that combines a multimodal knowledge graph—containing both text and molecular structure—with challenging questions designed to test large language models’ ability to retrieve and reason over this diverse information. The dataset highlights the limitations of current retrieval‑augmented generation systems, which typically handle only single‑modality text, by demonstrating that existing LLMs perform poorly unless provided with the necessary multimodal background data. This underscores the need for more robust RAG frameworks capable of integrating multiple data types for accurate responses.

By Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu, Fali Wang, Suhang Wang
arXiv Computation and Language
3d ago

The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?

The paper introduces Percept-V, a dataset of 6,000 program-generated images across 30 domains designed to test simple visual perception skills from the TVPS-4 framework. Experiments show that state‑of‑the‑art multimodal large language models perform poorly compared to humans, especially as image complexity increases, and that fine‑tuning yields only limited generalization to related datasets. The study highlights specific perception skills that remain challenging for current models.

By Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla
arXiv AI
3d ago

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero is presented as the first benchmark for long-horizon, part-level 3D editing, featuring natural-language instructions and target images for both geometry and texture. The benchmark uses a deterministic assembly engine that produces the exact target after each edit, and every sequence is manually reviewed. It compares non-agentic top‑down methods with LLM/VLM agent bottom‑up approaches, finding that the latter follow instructions more closely and preserve unedited parts better, though each edit takes minutes.

By Ruihan Yu, Yu-Ju Tsai, Muyao Niu, Runyi Li, Lian Fu, Hanqing Liu, Zheng-Hui Huang, Yonghao Yu, Sho Kuno, Ming-Hsuan Yang, Kaipeng Zhang, Zhixiang Wang
arXiv AI
3d ago

Toward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells, Organoids, and Biobots

The paper demonstrates that a natural-language interface can be trained offline to control a xenobot—a synthetic multicellular construct—by using a vision‑language model to judge whether archived intervention outcomes match language descriptions. By treating existing intervention–outcome data as a fixed dataset, the authors train a language‑to‑intervention mapping without new experiments, achieving 80% accuracy on held‑out data compared to a 66.7% baseline. This approach shows that language‑driven control of living systems can be learned purely from archival data and automated visual assessment.

By Nam H. Le, Douglas Blackiston, Michael Levin, Josh Bongard
arXiv AI
3d ago

FastOPD: On-Policy Distillation for Lightweight VLA Deployment

FastOPD is a framework that distills large Vision‑Language‑Action (VLA) models into lightweight versions by using on‑policy distillation with a flow map and a self‑consistency objective. The method trains a compact student to mimic the teacher’s dynamics, achieving performance close to the teacher while drastically reducing inference steps. Experiments on LIBERO, RoboTwin 2.0, and real‑robot deployments show significant latency reductions and higher success rates compared to existing few‑step distillation baselines.

By Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye
arXiv Machine Learning
3d ago

Toward Omni Multimodal Graph Foundation Model: A Topology-Driven Binding Approach

The paper introduces GraphBind, a topology-driven method for multimodal graph foundation models that binds heterogeneous node modalities into a unified shared space using graph topology. By leveraging stable graph structure to organize self and neighborhood semantics, GraphBind adapts this integrated space for both discriminative and generative tasks. Experiments against 11 baselines show that GraphBind outperforms them, achieving up to 28.1% relative improvement on key tasks.

By Xunkai Li, Chenxi Wan, Yinlin Zhu, Wang Luo, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
3d ago

MACTS-EM: Multi-Agent Collaborative Time Series Forecasting with Emergent Memory

MACTS-EM is a new multi‑agent framework for time series forecasting that combines domain‑specialised agents, a meta‑cognitive allocation layer, emergent memory for cross‑domain transfer, multimodal context integration, and adversarial robustness. The authors evaluate the system on financial, climate, energy, and pandemic data, reporting 8‑12% higher accuracy, 22‑27% better zero‑shot transfer, 16‑21% greater resilience to regime shifts, and 15‑18% faster recovery from distribution changes compared to existing methods. These results suggest that collaborative, agent‑based approaches can outperform traditional architectures in complex, real‑world forecasting tasks.

By Ahmad Shahi, Mamehgol Yousefi
arXiv Machine Learning
3d ago

SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents

SCAD (Structured Credit Assignment and Distillation) is a new method for training long‑horizon agents that separates planning from bounded subtask execution, distills execution locally, and refines planning credit using cross‑rollout subtask prefix trees. It gives planning full terminal credit while providing execution with positive terminal credit and teacher guidance. On all tested benchmarks, SCAD improves macro‑average accuracy by 4.48 percentage points on text tasks and 4.19 points on multimodal tasks compared to the strongest baseline.

By Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo
arXiv Machine Learning
3d ago

Min-Cost Flow Routing for Evidence Assembly in Long Multimodal Documents

The paper introduces <flowreader>, a method that casts evidence selection for long multimodal documents as a minimum‑cost flow problem over a multimodal content graph. It uses spectral decomposition to identify latent query‑relevant aspects and allocates a fixed evidence budget proportionally to their spectral energy, ensuring aspect coverage without a language‑model planning call. On the VisDoMBench benchmark with Qwen3‑VL‑32B, <flowreader> achieves the highest macro accuracy (68.9%) and outperforms prior systems by 2.7 points, while using fewer content nodes and reader tokens.

By Ambuj Mehrish, Sebastiano Vascon
arXiv Computer Vision
3d ago

A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

The paper introduces a simulation‑grounded vision‑language model (VLM) framework for wildfire monitoring that converts 2D wildfire simulations into labeled video episodes using a fixed Blender mapping to create low‑detail 3D proxies. These proxies, along with controllable video generation, provide a multimodal memory that a training‑free multi‑agent VLM system uses to retrieve reference episodes, reconcile visual and memory‑based predictions, and generate structured wildfire reports. The system achieves 77.3% accuracy on six simulator‑derived report fields, outperforming direct VLM querying and text‑only memory baselines.

By Duowen Chen, Yuchen Sun, Zhiqi Li, Yuxuan Liao, Sinan Wang, Bart van Bloemen Waanders, Bo Zhu
arXiv Computer Vision
3d ago

Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models

The paper introduces “SpaceConflict”, a benchmark of 23,196 multimodal inputs designed to test whether large language models can not only report spatial facts but also use them in reasoning. Experiments show a gap between a model’s ability to recover an initial spatial state from visual evidence and its ability to apply that state to perform transformations, with the gap narrowing but not closing as model size increases. To address this, the authors propose Operational State Supervision (OSS), which supervises task‑relevant spatial states and their transformation trajectories, improving accuracy on judgments that require state organization and use.

By Jinchang Zhang, Guoyu Lu
arXiv Computer Vision
3d ago

Behavior Pack Optimization for Video MLLM Post-Training

The paper introduces Behavior Pack Optimization (BPO), a post‑training method for video multimodal large language models that replaces single‑response rewards with a set of outputs across counterfactual views. BPO enforces stability when interventions are irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains, using an anchor‑relative advantage to keep the objective stable with small pack sizes. Experiments on datasets such as TempCompass, MVBench, and NExT‑QA show that BPO improves macro accuracy and abstention metrics for models like Qwen2.5‑VL‑7B‑Instruct, with gains that transfer to other benchmarks and models.

By Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou
arXiv Computer Vision
3d ago

Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

The paper introduces DyCPO, a co‑evolutionary framework that jointly optimizes token selection and adaptive counterfactual intervention for video reasoning. It builds a multi‑role dependence metric to balance visual exploration with answer‑relevance mining, suppressing filler tokens and spurious visual noise. By deriving counterfactual signals from the model’s own successful and failed rollouts, DyCPO enables self‑diagnostic analysis and co‑evolution of the optimization objective with the policy, leading to consistent performance gains on complex video reasoning benchmarks.

By Yudong Han, Yong Wang, Zaiquan Yang, Liang Lin, Chongyang Tao, Xiangxiang Chu, Liyuan Pan
arXiv Computer Vision
3d ago

Event-guided Neural Video Compression

The paper introduces Event-guided Neural Video Codec (ENVC), a neural video compression method that incorporates event streams—capturing brightness changes between frames—into both motion and frame coding stages. By using event-guided motion priors and event-conditioned predictors, ENVC improves RGB compression efficiency, achieving significant BD-rate savings across six benchmarks. The authors also synthesize paired RGB-event data for training and demonstrate that the gains persist on large-motion sequences, highlighting events as a valuable complementary modality for video coding.

By Jiyun Kong, Jungwoo Kim, Enes Eray Demirtas, Touradj Ebrahimi, Jong-Seok Lee
arXiv Computer Vision
3d ago

GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation

GeoScaffold introduces a geometric supervision framework that embeds 3D geometry into vision‑language navigation policies during training, eliminating the need for depth sensors or 3D encoders at inference. The method learns a compact depth tokenizer from training trajectories, then fine‑tunes the policy with geometry query tokens that reconstruct depth, connectivity, and traversability, turning these states into compact geometric latents for action decoding. After training, the tokenizer, generators, and reconstruction heads are discarded, leaving a lightweight backbone that outperforms vision‑only navigators on continuous VLN benchmarks.

By Yixuan Jiang, Wentong Li, An Liu, Zihao Xin, Fulin Tang, Cong Leng, Yang Gao, Jian Cheng