Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,018 stories · RSS feed

arXiv Machine Learning
Jul 28

SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception

arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.

By Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian
arXiv AI
Jul 28

CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics

arXiv:2607. 23258v1 Announce Type: new Abstract: Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to new datasets, experimental paradigms, and even species.

By Xinhong Xu, Yimeng Zhang, Yuanlong Zhang
arXiv Machine Learning
Jul 28

Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs

arXiv:2506. 03168v2 Announce Type: replace-cross Abstract: Amid the challenges posed by global population growth and climate change, traditional agricultural Internet of Things (IoT) systems is currently undergoing a significant digital transformation to facilitate efficient big data processing.

By Dawen Jiang, Zhishu Shen, Qiushi Zheng, Tiehua Zhang, Wei Xiang, Jiong Jin
arXiv AI
Jul 28

Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing

arXiv:2607. 23368v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful and interpretable explanations remains a key consideration for their responsible deployment in clinical settings.

By Jakub Rymarski (University of Warsaw, Poland), Adam Rempa{\l}a (University of Warsaw, Poland), Bart{\l}omiej Sobieski (University of Warsaw, Poland), Przemys{\l}aw Biecek (University of Warsaw, Poland)