Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,971 stories · RSS feed

arXiv AI
Sep 30

GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis

arXiv:2609.36082v1 Announce Type: new Abstract: We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing...

By Ethan D. Frakes, Amy Kvien, Rishabh Kundu, Redad Mehdi, Van D. Tran, Vibha S. Mandayam, Kristopher O. Davis, Erika I. Barcelos, Roger H. French, Yinghui Wu, Mengjie Li
arXiv AI
Sep 30

UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

arXiv:2609.37264v1 Announce Type: cross Abstract: Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separat...

By Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang, Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu
arXiv Computer Vision
Sep 30

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

arXiv:2604.15086v3 Announce Type: replace-cross Abstract: Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-gra...

By Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang, Jinjie Hu, Qiang Ji, Yihua Cao, Yihao Meng, Zhaoyue Cui, Mengmei Liu, Meng Meng, Jian Luan
arXiv AI
Sep 30

PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

PreviewDiff is a test‑time search method that uses multimodal critics to guide diffusion model sampling. By decoding partial previews at selected denoising checkpoints, scoring them with a multimodal judge, and branching over semantic prompt edits, it allows the generation process to be edited and rerouted before completion. The approach consistently outperforms budget‑matched Best‑of‑N sampling and scalar‑search baselines on image and video benchmarks, with early interventions and wider search yielding the biggest gains.

By Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister, Yale Song