arXiv:2609.36082v1 Announce Type: new
Abstract: We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing...
By Ethan D. Frakes, Amy Kvien, Rishabh Kundu, Redad Mehdi, Van D. Tran, Vibha S. Mandayam, Kristopher O. Davis, Erika I. Barcelos, Roger H. French, Yinghui Wu, Mengjie Li
arXiv:2609.36503v1 Announce Type: new
Abstract: Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired superv...
By Weihan Xu, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li, Yisi Liu, Liming Wang, Masato Ishii, Takashi Shibuya, Gopala Anumanchipalli, Paul Pu Liang
arXiv:2609.36828v1 Announce Type: new
Abstract: Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences wi...
By Wenxiao Fan, Jingling Fu, Lichen Ma, Yu He, Luohang Liu, Jinbao Xue, Ke Zhang, Junshi Huang, Kan Li
arXiv:2609.35912v1 Announce Type: cross
Abstract: Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that ag...
By Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du, Sibo Yi, Yuqi Qing, Zhenguang Liu, Qinming He, Shiwen Cui, Changhua Men
arXiv:2609.36838v1 Announce Type: cross
Abstract: Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teache...
By Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu, Wei Li, Wen Luo, Yang Xu, Yufan Shen, Luke Mao, Yang Du, Asher Qin, Houfeng Wang
arXiv:2609.37011v1 Announce Type: cross
Abstract: Federated learning enables multiple clients to collaboratively train models without sharing their private data. However, the lack of visibility into...
By Hongxu Su, Jianzhu Yao, Xuechao Wang, Pramod Viswanath
arXiv:2609.37196v1 Announce Type: cross
Abstract: Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allow...
By Yanjie Li, Xiangyu He, Xuelong Dai, Bin Xiao
arXiv:2609.37264v1 Announce Type: cross
Abstract: Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separat...
By Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang, Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu
arXiv:2609.37810v1 Announce Type: cross
Abstract: Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains chal...
By Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu, Yu-Gang Jiang
arXiv:2609.37195v1 Announce Type: new
Abstract: Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down ti...
By Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza
arXiv:2604.15086v3 Announce Type: replace-cross
Abstract: Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-gra...
By Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang, Jinjie Hu, Qiang Ji, Yihua Cao, Yihao Meng, Zhaoyue Cui, Mengmei Liu, Meng Meng, Jian Luan
arXiv:2609.36593v1 Announce Type: cross
Abstract: Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be des...
By Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li
arXiv:2609.35800v1 Announce Type: new
Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composa...
By Nenad Banfic
arXiv:2609.36215v1 Announce Type: new
Abstract: Power-grid digital twins must combine data-driven prediction with physically meaningful state evolution while preserving the provenance of environmenta...
By Huy Trinh, Michael Mai, Yu Nong
arXiv:2609.37148v1 Announce Type: new
Abstract: Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of oth...
By Siddhant Jain, Dimitra Tsovaltzi
arXiv:2609.37384v1 Announce Type: new
Abstract: Molecular representation learning is central to computer-aided drug discovery. Molecular graphs, SMILES strings, and 3D conformations provide complemen...
By Linqing Mo, Jiayu Zhou, Bin Chen
arXiv:2609.37515v1 Announce Type: cross
Abstract: Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities...
By Hyunjong Ok, Seunggu Kang, Jaeho Lee
arXiv:2609.35898v1 Announce Type: cross
Abstract: This paper proposes Wasserstein Causal Forests (WCF) for settings in which each unit's outcome is itself a probability distribution. This study also...
By Hugo Gobato Souto
arXiv:2609.35922v1 Announce Type: cross
Abstract: A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, f...
By Bhavik Mangla
PreviewDiff is a test‑time search method that uses multimodal critics to guide diffusion model sampling. By decoding partial previews at selected denoising checkpoints, scoring them with a multimodal judge, and branching over semantic prompt edits, it allows the generation process to be edited and rerouted before completion. The approach consistently outperforms budget‑matched Best‑of‑N sampling and scalar‑search baselines on image and video benchmarks, with early interventions and wider search yielding the biggest gains.
By Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister, Yale Song