arXiv:2610.02976v1 Announce Type: new
Abstract: Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about anoth...
By Hyunjae Ra, Aecheon Jung, Jungin Park, Sungeun Hong
arXiv:2610.02324v1 Announce Type: cross
Abstract: Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy...
By Xiaofei Yin, Tong Chu, Jiyuan Fu, Jun Lan, Shuheng Zhou, Huijia Zhu
arXiv:2610.02320v1 Announce Type: cross
Abstract: Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually...
By A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar
arXiv:2608.23758v3 Announce Type: replace-cross
Abstract: Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain lar...
By Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
arXiv:2610.03693v1 Announce Type: new
Abstract: Scarcity of labeled data limits development of deep learning biomarkers in oncology. We develop a two-stage AI model predicting pathological complete r...
By Jungkyu Park, Dhruva Biswas, Joseph Cappadona, Cerise Tang, Ken G. Zeng, Bartosz Machura, Chuwen Liu, Paolo Tarantino, Coral Omene, Francisco J. Esteva, Rohit Bhargava, Marcin Braun, Kamila Pa\'zdzierz, Jakub Czerwi\'nski, Hanna Roma\'nska-Knight, Albert Grinshpun, Bareket Daniel, Michele Buchinger, Frederick Howard, Piotr Wysocki, Brie Chun, Freya Schnabel, Rich Caruana, Jan Witowski, Krzysztof J. Geras
arXiv:2610.02349v1 Announce Type: cross
Abstract: Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-i...
By Ryuichi Yamafuji Lun, Jingzhen Wang, Shreyas Kolte, Ruiteng Li
arXiv:2610.02718v1 Announce Type: cross
Abstract: Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named...
By Peilin Yang, Xiaoyu Liu, Jian Sun, Qinghua Tao
arXiv:2606.20627v2 Announce Type: replace
Abstract: Joint-Embedding Predictive Architectures (JEPAs) enable agents to plan in latent space by imagining the outcomes of candidate actions, yet task spe...
By Samuel Barbeau, Simon Roy, Giovanni Beltrame, Christian Desrosiers, Nicolas Thome
arXiv:2610.03389v1 Announce Type: cross
Abstract: Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every...
By Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Shaoyi Huang, Pu Zhao, Yanzhi Wang
arXiv:2610.02765v1 Announce Type: cross
Abstract: Monitoring planned driving trajectories requires accurately estimating the collision likelihood with actors whose motion is itself impacted by the eg...
By Lu\'is Marques, Rong Fang, Disha Kamale, Dmitry Berenson
arXiv:2606.19617v2 Announce Type: replace-cross
Abstract: We present GB-LSR (Global-Bandwidth Local Spectral Representation), a fixed-grid local spectral representation for continuous image decoding....
By Max Shad, Naeem Khoshnevis
arXiv:2610.02666v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert...
By Jin Hyun, Jung Gyu Min, Gyuhyun Jung, Youngjoo Lee
arXiv:2610.02388v1 Announce Type: new
Abstract: Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes...
By Ran Dan, Si-Tong Wei, Pengfei Xiong, Wei Zhang, Yadong Mu, Peng-Shuai Wang
arXiv:2510.22276v4 Announce Type: replace-cross
Abstract: Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing Eng...
By Issa Sugiura, Shuhei Kurita, Yusuke Oda, Daisuke Kawahara, Yasuo Okabe, Naoaki Okazaki
arXiv:2610.02323v1 Announce Type: cross
Abstract: Flow-based Vision-Language-Action (VLA) policies generate action chunks by transporting samples from a task-agnostic isotropic Gaussian source. As th...
By Jie He, Wei Li, Junwen Tong, Rui Shao, Wei-Shi Zheng, Liqiang Nie
arXiv:2610.02270v1 Announce Type: cross
Abstract: Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies trea...
By Xinye Yang, Zhusi Zhong, Scott Collins, Grayson Baird, Xuyu Wang, Zhicheng Jiao
The paper introduces MOVE, a framework for multimodal open‑world verification and expansion in graph learning. MOVE jointly uses visual tokens, textual attributes, and graph context to identify nodes that cannot be assigned to existing classes, then employs a multimodal LLM to generate candidate class descriptions. It selectively expands the class space only when multimodal evidence consistently supports the new classes, avoiding redundancy, and reports an average 11.87% improvement across unknown recognition, open‑domain annotation, and downstream graph learning tasks.
By Zekai Chen, Jiayang Xing, Xun Wu, Miao Zhang, Xunkai Li, Kairui Yang, Zhengyu Wu, Xu Wang, Rong-Hua Li, Guoren Wang
The paper introduces a pullback Riemannian geometry tailored for multimodal data by employing a latent Gaussian mixture model. It defines a smooth, positive‑definite metric based on responsibility‑weighted component precision, extending the standard single‑Gaussian construction. Experiments on synthetic, multi‑view image, and MNIST datasets demonstrate reduced transport distortion, accurate trajectory recovery, and more realistic interpolation.
By Honglei Brinkmann, Lucas Ng, Georgios Batzolis, Mark Girolami, Carola-Bibiane Sch\"onlieb, Willem Diepeveen
HADRec is a Hierarchy-Aware Drug Recommendation framework that fuses molecular knowledge and electronic health records to improve medication recommendation. It uses LLaMA-7B to encode clinical notes, ChemBERTa to encode drug SMILES strings, and a cross‑attention mechanism for multimodal fusion, while a hierarchical predictor and consistency constraint loss enforce adherence to the ATC classification system. Experiments on MIMIC‑III and MIMIC‑IV show state‑of‑the‑art performance, strong generalization, and well‑calibrated predictions, with counterfactual evaluation indicating clinically aligned reasoning.
By Junke Wang, Hongshun Ling, Li Zhang, Jinjing Wu, Tong Shao, Fang Wang, Yuan Gao
The paper introduces a method called Debias Anything that jointly addresses fairness and diversity in diffusion models without requiring sensitive-attribute annotations. By connecting a frozen diffusion model to a pretrained vision-language embedding space via an adapter, the approach uses pairs of text prompts to guide batch composition toward desired attribute proportions and employs a disagreement score to promote diversity. The method is applicable to both unconditional and text-conditional diffusion models and demonstrates improved quality and diversity while maintaining comparable fairness levels in experiments.
By Th\'eau d'Audiffret, Mariia Vladimirova, Jean-Yves Franceschi