Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,626 stories · RSS feed

arXiv AI
Jun 9

DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home

arXiv:2606. 07542v1 Announce Type: cross Abstract: Generative AI is reshaping healthcare, yet most existing advances rely on hospital-grade devices, which limits their accessibility and potential for health management outside clinical settings.

By Changshuo Liu, Junran Wu, Zhongle Xie, Wenqiao Zhang, Kaiping Zheng, Jiaqi Zhu, Qingpeng Cai, Ooi Gene Anne, Marcus Chun Jin Tan, Jianwei Yin, James Wei Luen Yip, Beng Chin Ooi
arXiv AI
Jun 9

Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control

arXiv:2606. 08405v1 Announce Type: new Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific discovery in physical systems fundamentally requires an interpretable chain of reasoning that connects physical evidence to structured control architectures.

By Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang
arXiv AI
Jun 9

Integrating Deep Learning Demand Forecasting with Multi-Objective Optimization for Circular Coffee Supply Chains: A Data-Driven Framework for Cost, Emissions, and Freshness Management

arXiv:2606. 08314v1 Announce Type: new Abstract: The coffee supply chain is one of the most complex agri-food networks, marked by geographically dispersed production, multi-tier coordination, and high sensitivity to quality and freshness.

By Ger\c{c}ek Budak (Department of Industrial Engineering, Ankara Y{\i}ld{\i}r{\i}m Beyaz{\i}t University, Ke\c{c}i\"oren, Ankara 06010, T\"urkiye), Faraz Gholamzadeh Gharehgheshlaghi (Department of Industrial Engineering, Ankara Y{\i}ld{\i}r{\i}m Beyaz{\i}t University, Ke\c{c}i\"oren, Ankara 06010, T\"urkiye), Melika Barjesteh Vaezi (Department of Kinesiology and Sport Management, Texas Tech University, Lubbock, TX, United States), Ahmad Gholizadeh Lonbar (Department of Civil, Construction, and Environmental Engineering, University of Alabama, Tuscaloosa, AL, USA)
arXiv AI
Jun 9

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention

arXiv:2606. 07639v1 Announce Type: cross Abstract: Video understanding is shifting from the offline paradigm -- taking a fully recorded video as input and producing a single answer after it ends -- toward real-time interaction, in which the model perceives new frames while still replying, revises its answer as new evidence appears, and remains silent when there is nothing to say.

By Pengyu Wang, Chenkun Tan, Shaojun Zhou, Wei Huang, Qirui Zhou, Zhan Huang, Zhen Ye, Jijun Cheng, Xiaomeng Qian, Yanxin Chen, Xingyang He, Huazheng Zeng, Chenghao Wang, Pengfei Wang, Hongkai Wang, Shanqing Gao, Yixian Tian, Chenghao Liu, Xinghao Wang, Botian Jiang, Xipeng Qiu
arXiv AI
Jun 9

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis

arXiv:2606. 08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large multimodal datasets linking diverse foods to complete nutrient profiles.

By Runze Yan, Minxiao Wang, Jiaying Lu, Darren Liu, Xiao Hu, Hanqi Luo
arXiv Machine Learning
Jun 9

Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study

arXiv:2606. 09362v1 Announce Type: cross Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues.

By Eduardo Borges, Manuel Abreu, Lu\'is Garrote, Urbano J. Nunes
arXiv AI
Jun 9

A Resilience-as-a-Service assessment framework for coordinated disruption response in interdependent urban transit systems

arXiv:2606. 08849v1 Announce Type: new Abstract: Urban public transport disruptions require rapid response strategies, yet existing studies rarely provide a decision support framework to compare alternative disruption response solutions using a common set of dynamic, passenger, operator, and environment oriented indicators.

By Sara Jaber, S. M. Hassan Mahdavi, Neila Bhouri, Mostafa Ameli