arXiv:2604.03114v2 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy...
By Zhangyun Tan, Zeliang Zhang, Jiani Liu, Susan Liang, Yolo Y. Tang, Lisha Chen, Chenliang Xu
arXiv:2605.01970v4 Announce Type: replace-cross
Abstract: Memory systems enable otherwise stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. Th...
By Debeshee Das, Julien Piet, Darya Kaviani, Luca Beurer-Kellner, Florian Tram\`er, David Wagner
arXiv:2609.32013v2 Announce Type: replace-cross
Abstract: We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to pri...
By Quinlan Sykora, Sourav Biswas, Christopher Diehl, Andrew Cunningham, Thomas Gilles, Raquel Urtasun
arXiv:2609.32876v2 Announce Type: replace-cross
Abstract: State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons...
By Yishu Zhang, Yun Li, Daiwei Zhang
arXiv:2609.35575v2 Announce Type: replace-cross
Abstract: The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrat...
By Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang
arXiv:2609.36115v1 Announce Type: new
Abstract: Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-criti...
By Shenghong Dai, Shiva Kumar Pentyala, Yingchi Liu, Shubham Mehrotra, Suman Banerjee, James Zhu, Bin Bi, Sitaram Asur, Phil Mui
arXiv:2609.38133v1 Announce Type: new
Abstract: Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general...
By Ruoyu Lin, Magnus Egerstedt, Fabio Pasqualetti
arXiv:2609.36324v1 Announce Type: cross
Abstract: Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their gener...
By Yentl Collin, Evan Dufraisse, Amr Mohamed, Amine Khelif Khelif, Dani Bouch, Guokan Shang
arXiv:2609.37165v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific f...
By Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee, Woo-Jeong Baek, Adam Zhou, Chip Huyen, Jun-Ki Lee, Gi-Cheon Kang, Byoung-Tak Zhang
arXiv:2609.37230v1 Announce Type: cross
Abstract: Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a disti...
By Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim
arXiv:2609.37367v1 Announce Type: cross
Abstract: Decentralized large language model (LLM) fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a centra...
By Sayan Biswas, Jade Garcia Bourr\'ee, Rachid Guerraoui, Maxime Jacovella, Anne-Marie Kermarrec, Sathwika Peechara, Martijn de Vos, Milos Vujasinovic
arXiv:2609.37485v1 Announce Type: cross
Abstract: Bi-temporal change understanding, which localizes and characterizes what changed between two satellite images, is central to disaster response and en...
By Haruki Watase, Shunya Nagashima, Takayuki Nishimura
arXiv:2609.37488v1 Announce Type: cross
Abstract: Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial...
By Taiyo Sato, Takamasa Sanda, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama, Shunya Nagashima
arXiv:2609.37918v1 Announce Type: cross
Abstract: Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capab...
By Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami
arXiv:2602.15382v3 Announce Type: replace-cross
Abstract: Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal st...
By Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, Xiaoqian Wang, Jing Gao
arXiv:2609.32810v2 Announce Type: replace-cross
Abstract: Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions,...
By Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen, MingYu Lu, Huan-Yu Hsu, Yu Gu, Yue Guo, Sheng Wang, Wei Qiu, Hanwen Xu
arXiv:2609.35791v1 Announce Type: new
Abstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conver...
By Puneet Mathur, Dinesh Manocha
arXiv:2609.35821v1 Announce Type: new
Abstract: Disaster social sensing converts public social-media posts into evidence for situational awareness and humanitarian needs, but generative artificial in...
By Xiaoshan Zhou, Zaifu Zhan
arXiv:2609.36691v1 Announce Type: new
Abstract: Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embod...
By Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu, Haoran Lu, Yijiang Li, Yaqi Xie, Katia P. Sycara, Han Liu
arXiv:2609.36850v1 Announce Type: new
Abstract: Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipul...
By Wenbin Shen, Guoxuan Qin, Guangxu Yao, Baodong Wang, Yuanbo Rui, Zhichao Lian