arXiv:2609.37837v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly el...
By Haotian Deng, Wenbin Xing, Gang Xu, Tao He, Jinkai Zheng, Chun Li, Zheng Zhu, Ming Li
arXiv:2609.38111v1 Announce Type: new
Abstract: Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when...
By Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng
arXiv:2609.35150v1 Announce Type: cross
Abstract: Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose vi...
By Siddhant Jain, Anna Lea Reinwarth, Dimitra Tsovaltzi, Rafael Math, Julia Renner
arXiv:2609.38106v1 Announce Type: cross
Abstract: Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggre...
By Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar
arXiv:2609.38177v1 Announce Type: cross
Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...
By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv:2510.12041v3 Announce Type: replace
Abstract: Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecifie...
By Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang
MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.
By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
arXiv:2608.00677v2 Announce Type: replace
Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-mo...
By Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu, Jie Li, Yan Teng, Xingjun Ma, Xia Hu, Yu-Gang Jiang
arXiv:2608.08477v5 Announce Type: replace
Abstract: We build VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model for offline use, and measure what limits its visual groundi...
By Juan S. Santillana
arXiv:2603.08249v2 Announce Type: replace-cross
Abstract: Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but...
By Pol Buitrago, Javier Hernando
arXiv:2606.28249v2 Announce Type: replace-cross
Abstract: Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised...
By Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao, Jingyuan Xing, Baiji Liu, Xiangmin Xu
arXiv:2609.35823v1 Announce Type: new
Abstract: Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes....
By Kang Yang, Shuai Liu, Hang Li, Yance Fang, Deying Li, Yongcai Wang
arXiv:2609.36136v1 Announce Type: new
Abstract: Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on vi...
By Xin Chen, Anan Du, Feng Feng, Pei Fu, Jian Luan, Longwei Xu, Shaojie Zhang, Hang Li, Heng Qu, Cheng Tan
arXiv:2609.36219v1 Announce Type: new
Abstract: Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as...
By Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang, Bolin Lai, Junho Kim, James Matthew Rehg
arXiv:2609.36386v1 Announce Type: new
Abstract: An event camera produces an asynchronous stream, but what is visible in that stream depends on how a downstream consumer, such as a model or detector,...
By Shoaib Ahmed Dipu, Md. Shaown Miah, Kamrul Hasan, Sayeed Shafayet Chowdhury
arXiv:2609.36628v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, partic...
By Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong, Chenwei Xoe, Zhaohe Liao
arXiv:2609.36680v1 Announce Type: new
Abstract: Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In v...
By Zizhao Li, Chengyi Cai, Mohammed Yaqoob Ansari, Feng Liu, Joseph West, Kourosh Khoshelham
arXiv:2609.36685v1 Announce Type: new
Abstract: Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the...
By Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu, Ming Meng
arXiv:2609.36776v1 Announce Type: new
Abstract: Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the import...
By Cheng Ye, Weidong Chen, Peipei Song, Zhendong Mao
arXiv:2609.36815v1 Announce Type: new
Abstract: Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotatio...
By Zhen Liu, Letian Li, Jinpeng Wang, Shuzhao Xie, Yuzhi Huang, Jingyan Jiang, Zhi Wang