Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,568 stories · RSS feed

arXiv AI
2d ago

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero is presented as the first benchmark for long-horizon, part-level 3D editing, featuring natural-language instructions and target images for both geometry and texture. The benchmark uses a deterministic assembly engine that produces the exact target after each edit, and every sequence is manually reviewed. It compares non-agentic top‑down methods with LLM/VLM agent bottom‑up approaches, finding that the latter follow instructions more closely and preserve unedited parts better, though each edit takes minutes.

By Ruihan Yu, Yu-Ju Tsai, Muyao Niu, Runyi Li, Lian Fu, Hanqing Liu, Zheng-Hui Huang, Yonghao Yu, Sho Kuno, Ming-Hsuan Yang, Kaipeng Zhang, Zhixiang Wang
arXiv AI
2d ago

Toward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells, Organoids, and Biobots

The paper demonstrates that a natural-language interface can be trained offline to control a xenobot—a synthetic multicellular construct—by using a vision‑language model to judge whether archived intervention outcomes match language descriptions. By treating existing intervention–outcome data as a fixed dataset, the authors train a language‑to‑intervention mapping without new experiments, achieving 80% accuracy on held‑out data compared to a 66.7% baseline. This approach shows that language‑driven control of living systems can be learned purely from archival data and automated visual assessment.

By Nam H. Le, Douglas Blackiston, Michael Levin, Josh Bongard
arXiv AI
2d ago

FastOPD: On-Policy Distillation for Lightweight VLA Deployment

FastOPD is a framework that distills large Vision‑Language‑Action (VLA) models into lightweight versions by using on‑policy distillation with a flow map and a self‑consistency objective. The method trains a compact student to mimic the teacher’s dynamics, achieving performance close to the teacher while drastically reducing inference steps. Experiments on LIBERO, RoboTwin 2.0, and real‑robot deployments show significant latency reductions and higher success rates compared to existing few‑step distillation baselines.

By Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye
arXiv Machine Learning
2d ago

Toward Omni Multimodal Graph Foundation Model: A Topology-Driven Binding Approach

The paper introduces GraphBind, a topology-driven method for multimodal graph foundation models that binds heterogeneous node modalities into a unified shared space using graph topology. By leveraging stable graph structure to organize self and neighborhood semantics, GraphBind adapts this integrated space for both discriminative and generative tasks. Experiments against 11 baselines show that GraphBind outperforms them, achieving up to 28.1% relative improvement on key tasks.

By Xunkai Li, Chenxi Wan, Yinlin Zhu, Wang Luo, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
2d ago

MACTS-EM: Multi-Agent Collaborative Time Series Forecasting with Emergent Memory

MACTS-EM is a new multi‑agent framework for time series forecasting that combines domain‑specialised agents, a meta‑cognitive allocation layer, emergent memory for cross‑domain transfer, multimodal context integration, and adversarial robustness. The authors evaluate the system on financial, climate, energy, and pandemic data, reporting 8‑12% higher accuracy, 22‑27% better zero‑shot transfer, 16‑21% greater resilience to regime shifts, and 15‑18% faster recovery from distribution changes compared to existing methods. These results suggest that collaborative, agent‑based approaches can outperform traditional architectures in complex, real‑world forecasting tasks.

By Ahmad Shahi, Mamehgol Yousefi
arXiv Machine Learning
2d ago

SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents

SCAD (Structured Credit Assignment and Distillation) is a new method for training long‑horizon agents that separates planning from bounded subtask execution, distills execution locally, and refines planning credit using cross‑rollout subtask prefix trees. It gives planning full terminal credit while providing execution with positive terminal credit and teacher guidance. On all tested benchmarks, SCAD improves macro‑average accuracy by 4.48 percentage points on text tasks and 4.19 points on multimodal tasks compared to the strongest baseline.

By Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo
arXiv Machine Learning
2d ago

Min-Cost Flow Routing for Evidence Assembly in Long Multimodal Documents

The paper introduces <flowreader>, a method that casts evidence selection for long multimodal documents as a minimum‑cost flow problem over a multimodal content graph. It uses spectral decomposition to identify latent query‑relevant aspects and allocates a fixed evidence budget proportionally to their spectral energy, ensuring aspect coverage without a language‑model planning call. On the VisDoMBench benchmark with Qwen3‑VL‑32B, <flowreader> achieves the highest macro accuracy (68.9%) and outperforms prior systems by 2.7 points, while using fewer content nodes and reader tokens.

By Ambuj Mehrish, Sebastiano Vascon
arXiv Computer Vision
2d ago

A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

The paper introduces a simulation‑grounded vision‑language model (VLM) framework for wildfire monitoring that converts 2D wildfire simulations into labeled video episodes using a fixed Blender mapping to create low‑detail 3D proxies. These proxies, along with controllable video generation, provide a multimodal memory that a training‑free multi‑agent VLM system uses to retrieve reference episodes, reconcile visual and memory‑based predictions, and generate structured wildfire reports. The system achieves 77.3% accuracy on six simulator‑derived report fields, outperforming direct VLM querying and text‑only memory baselines.

By Duowen Chen, Yuchen Sun, Zhiqi Li, Yuxuan Liao, Sinan Wang, Bart van Bloemen Waanders, Bo Zhu
arXiv Computer Vision
2d ago

Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models

The paper introduces “SpaceConflict”, a benchmark of 23,196 multimodal inputs designed to test whether large language models can not only report spatial facts but also use them in reasoning. Experiments show a gap between a model’s ability to recover an initial spatial state from visual evidence and its ability to apply that state to perform transformations, with the gap narrowing but not closing as model size increases. To address this, the authors propose Operational State Supervision (OSS), which supervises task‑relevant spatial states and their transformation trajectories, improving accuracy on judgments that require state organization and use.

By Jinchang Zhang, Guoyu Lu
arXiv Computer Vision
2d ago

Behavior Pack Optimization for Video MLLM Post-Training

The paper introduces Behavior Pack Optimization (BPO), a post‑training method for video multimodal large language models that replaces single‑response rewards with a set of outputs across counterfactual views. BPO enforces stability when interventions are irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains, using an anchor‑relative advantage to keep the objective stable with small pack sizes. Experiments on datasets such as TempCompass, MVBench, and NExT‑QA show that BPO improves macro accuracy and abstention metrics for models like Qwen2.5‑VL‑7B‑Instruct, with gains that transfer to other benchmarks and models.

By Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou
arXiv Computer Vision
2d ago

Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

The paper introduces DyCPO, a co‑evolutionary framework that jointly optimizes token selection and adaptive counterfactual intervention for video reasoning. It builds a multi‑role dependence metric to balance visual exploration with answer‑relevance mining, suppressing filler tokens and spurious visual noise. By deriving counterfactual signals from the model’s own successful and failed rollouts, DyCPO enables self‑diagnostic analysis and co‑evolution of the optimization objective with the policy, leading to consistent performance gains on complex video reasoning benchmarks.

By Yudong Han, Yong Wang, Zaiquan Yang, Liang Lin, Chongyang Tao, Xiangxiang Chu, Liyuan Pan
arXiv Computer Vision
2d ago

Event-guided Neural Video Compression

The paper introduces Event-guided Neural Video Codec (ENVC), a neural video compression method that incorporates event streams—capturing brightness changes between frames—into both motion and frame coding stages. By using event-guided motion priors and event-conditioned predictors, ENVC improves RGB compression efficiency, achieving significant BD-rate savings across six benchmarks. The authors also synthesize paired RGB-event data for training and demonstrate that the gains persist on large-motion sequences, highlighting events as a valuable complementary modality for video coding.

By Jiyun Kong, Jungwoo Kim, Enes Eray Demirtas, Touradj Ebrahimi, Jong-Seok Lee
arXiv Computer Vision
2d ago

GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation

GeoScaffold introduces a geometric supervision framework that embeds 3D geometry into vision‑language navigation policies during training, eliminating the need for depth sensors or 3D encoders at inference. The method learns a compact depth tokenizer from training trajectories, then fine‑tunes the policy with geometry query tokens that reconstruct depth, connectivity, and traversability, turning these states into compact geometric latents for action decoding. After training, the tokenizer, generators, and reconstruction heads are discarded, leaving a lightweight backbone that outperforms vision‑only navigators on continuous VLN benchmarks.

By Yixuan Jiang, Wentong Li, An Liu, Zihao Xin, Fulin Tang, Cong Leng, Yang Gao, Jian Cheng
arXiv Computer Vision
2d ago

Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization

The paper investigates how affective fine‑tuning shapes the internal architecture of multimodal foundation models. By analyzing 13 model instances across nine designs, it finds that adapting the feed‑forward network (FFN) consistently outperforms attention‑only adaptation and nearly matches full‑model tuning, revealing the FFN as an efficient adaptation substrate. Moreover, joint optimization leads to emergent functional specialization, notably a prominent gate projection pathway, which the authors exploit in Gate‑Focused Efficient Tuning (GET) to achieve 96.2–98.0% of full‑model performance with only 19.3–24.5% of the trainable parameters.

By Zhen Zhang, Runhao Zeng, Sicheng Zhao, Xiping Hu
arXiv Machine Learning
2d ago

Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

This thesis introduces methods for assessing and enhancing the robustness of large language models (LLMs) against adversarial input variations. It proposes a generative robustness metric, R_stab(f), based on Jensen‑Shannon divergence, and proves bounds for localized attacks. The work also presents adaptive attacks, Trojan detection techniques, defense strategies for multi‑layer systems, and lightweight attestation for Model Context Protocol (MCP) agentic systems, all implemented in the JudgeGuard, TrojanArmor, and MCPSec suites.

By Narek Maloyan
arXiv AI
2d ago

Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS

The study investigates how two computational dimensions—model depth and refinement steps—affect intelligibility and speaker identity in masked-diffusion text‑to‑speech systems. Experiments with 15 models (19–133 M parameters) and up to 16 refinement steps show that refinement improves intelligibility more than identity, with a 1.86× asymmetry that persists even after retraining. Best‑of‑K search can recover identity when refinement fails, and analysis indicates that depth and steps target distinct bottlenecks, requiring separate optimization.

By Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh
arXiv AI
2d ago

Multimodal reasoning for broadly neutralizing antibody discovery from label-free human B cell repertoires across virus families

The paper introduces ImmuneAgent, a closed‑loop AI system that combines multimodal reasoning, continual meta‑learning, and wet‑lab feedback to identify broadly neutralizing antibodies (bnAbs) from human B cell repertoires. Applied to vaccinated or infected cohorts, ImmuneAgent achieved a 55% neutralization discovery rate and an 11% bnAb yield, outperforming existing sequence‑based predictors and co‑folding models. Five discovered antibodies provided full in vivo protection against lethal influenza, and the system uncovered conserved bnAb reservoirs and structural signatures that enabled cross‑viral antibody discovery without antigen‑specific sorting.

By Hantao Lou, Jianqing Zheng, Can Yue, Meihan Zhang, Yuanchao Bao, Yu Chen, Mengting Huang, Yupeng Yang, Qianyu Pan, Nana Fu, Yansong Shi, Hongli Li, Yangyang Chai, Ruyi Chen, Wansheng Li, Zhu Liang, Rongmei Yao, Yuanhan Mo, Lei Wang, Chunmei Wang, Yun Quan, Qiong Zhang, Xiangxi Wang, Xuetao Cao