arXiv:2606. 08094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware on which robots actually run.
By Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen, Thanh Q. Duong, Linh D. Le, Duy M. H. Nguyen, Vien A. Ngo, An T. Le
arXiv:2606. 08034v1 Announce Type: cross Abstract: Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions.
By Muhammad Falensi Azmi, Ikhlasul Akmal Hanif, Vallerie Alexandra Putra, Adi Yeltay, Abdullah Mubarak, Fajri Koto
arXiv:2606. 07924v1 Announce Type: cross Abstract: This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via MultimodAl Retrieval (MAGMaR).
By Jiaxin Dai, Zehang Wei, Jiamin Yan, Xiang Xiang
arXiv:2606. 07943v1 Announce Type: cross Abstract: Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks.
By Haochang Hao, Dehai Min, Zhifang Zhang, Yunbei Zhang, Miao Xu, Yingqiang Ge, Lu Cheng
arXiv:2606. 07613v1 Announce Type: cross Abstract: Visual evidence has long been treated as a reliable form of legal proof, but advances in artificial intelligence (AI) are undermining that assumption.
By Jinzhe Tan, Ali Ekber Cinar, Karim Benyekhlef
arXiv:2606. 07542v1 Announce Type: cross Abstract: Generative AI is reshaping healthcare, yet most existing advances rely on hospital-grade devices, which limits their accessibility and potential for health management outside clinical settings.
By Changshuo Liu, Junran Wu, Zhongle Xie, Wenqiao Zhang, Kaiping Zheng, Jiaqi Zhu, Qingpeng Cai, Ooi Gene Anne, Marcus Chun Jin Tan, Jianwei Yin, James Wei Luen Yip, Beng Chin Ooi
arXiv:2606. 07533v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues.
By Pawe{\l} Pozorski, Jakub Muszy\'nski, Maria Ganzha
arXiv:2606. 08405v1 Announce Type: new Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific discovery in physical systems fundamentally requires an interpretable chain of reasoning that connects physical evidence to structured control architectures.
By Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang
arXiv:2606. 09648v1 Announce Type: cross Abstract: Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment.
By Luciano Duarte, Olga Ovcharenko, Sebastian Schelter
arXiv:2606. 07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.
By Lujun Li, Lama Sleem, Niccolo Gentile, Yangjie Xu, Yewei Song, Wenbo Wu, Radu State
arXiv:2606. 09749v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks.
By Seongbin Park, Fan Zhang, Baharan Mirzasoleiman, Shahriar Talebi, Nader Sehatbakhsh
arXiv:2606. 08314v1 Announce Type: new Abstract: The coffee supply chain is one of the most complex agri-food networks, marked by geographically dispersed production, multi-tier coordination, and high sensitivity to quality and freshness.
By Ger\c{c}ek Budak (Department of Industrial Engineering, Ankara Y{\i}ld{\i}r{\i}m Beyaz{\i}t University, Ke\c{c}i\"oren, Ankara 06010, T\"urkiye), Faraz Gholamzadeh Gharehgheshlaghi (Department of Industrial Engineering, Ankara Y{\i}ld{\i}r{\i}m Beyaz{\i}t University, Ke\c{c}i\"oren, Ankara 06010, T\"urkiye), Melika Barjesteh Vaezi (Department of Kinesiology and Sport Management, Texas Tech University, Lubbock, TX, United States), Ahmad Gholizadeh Lonbar (Department of Civil, Construction, and Environmental Engineering, University of Alabama, Tuscaloosa, AL, USA)
arXiv:2606. 07646v1 Announce Type: cross Abstract: Test-time adaptation (TTA) aims to align a model to shifting test domains using only unlabeled streaming data.
By Xiaoran Xu, Yifan Xu, Yupeng Wu, Xiaoshan Yang, Changsheng Xu
arXiv:2606. 09605v1 Announce Type: new Abstract: Foundation models offer a promising route to compress multi-modal physiological signals into compact representations of human health, with broad applications across sleep medicine, cardiology, neurology and other healthcare domains.
By Jonathan F. Carter, Lionel Tarassenko
arXiv:2606. 07639v1 Announce Type: cross Abstract: Video understanding is shifting from the offline paradigm -- taking a fully recorded video as input and producing a single answer after it ends -- toward real-time interaction, in which the model perceives new frames while still replying, revises its answer as new evidence appears, and remains silent when there is nothing to say.
By Pengyu Wang, Chenkun Tan, Shaojun Zhou, Wei Huang, Qirui Zhou, Zhan Huang, Zhen Ye, Jijun Cheng, Xiaomeng Qian, Yanxin Chen, Xingyang He, Huazheng Zeng, Chenghao Wang, Pengfei Wang, Hongkai Wang, Shanqing Gao, Yixian Tian, Chenghao Liu, Xinghao Wang, Botian Jiang, Xipeng Qiu
arXiv:2606. 08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large multimodal datasets linking diverse foods to complete nutrient profiles.
By Runze Yan, Minxiao Wang, Jiaying Lu, Darren Liu, Xiao Hu, Hanqi Luo
arXiv:2511. 17855v5 Announce Type: replace Abstract: Robots must learn from both what people do and what they say, but either modality alone is often incomplete: physical corrections are grounded but ambiguous in intent, while language expresses high-level goals but lacks physical grounding.
By Jordan Abi Nader, David Lee, Nathaniel Dennler, Andreea Bobu
arXiv:2606. 07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking external tools.
By Youting Wang, Yuan Tang, Yitian Qian, Chen Zhao
arXiv:2606. 09362v1 Announce Type: cross Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues.
By Eduardo Borges, Manuel Abreu, Lu\'is Garrote, Urbano J. Nunes
arXiv:2606. 08849v1 Announce Type: new Abstract: Urban public transport disruptions require rapid response strategies, yet existing studies rarely provide a decision support framework to compare alternative disruption response solutions using a common set of dynamic, passenger, operator, and environment oriented indicators.
By Sara Jaber, S. M. Hassan Mahdavi, Neila Bhouri, Mostafa Ameli