Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,022 stories · RSS feed

arXiv Machine Learning
Jul 16

GFlowRL: Scaling Distribution-Matching RL to Large Language Models

arXiv:2607. 13394v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes.

By Xiaodong Liu, Michael Xu, Jack W. Stokes, Paul Smolensky, Doug Burger, Jianfeng Gao
arXiv Machine Learning
Jul 16

BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping

arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.

By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
arXiv Machine Learning
Jul 16

RF-Informed Graph Neural Networks for Accurate and Data-Efficient Circuit Performance Prediction

arXiv:2508. 16403v3 Announce Type: replace Abstract: Accurately predicting the performance of active radio frequency (RF) circuits is essential for modern wireless systems but remains challenging due to highly nonlinear behavior and the high computational cost of traditional simulation tools.

By Anahita Asadi, Leonid Popryho, Inna Partin-Vaisband
arXiv Machine Learning
Jul 16

Multimodal Empirical Bayes Variational Autoencoders for Joint Longitudinal and Time-to-Event Modeling

arXiv:2607. 13984v1 Announce Type: cross Abstract: Longitudinal tumor measurements, dropout information, and genetic covariates provide complementary information about treatment response, but integrating these data sources within a single population modeling framework remains challenging.

By Anders Sj\"oberg, Nils Olsson, Marcus Baaz, Mats Jirstrand
arXiv AI
Jul 16

A Hybrid Mamba for Audio-Visual Navigation

arXiv:2607. 13110v1 Announce Type: cross Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences.

By Yi Wang, Yinfeng Yu
arXiv AI
Jul 16

Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning

arXiv:2607. 13826v1 Announce Type: cross Abstract: Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major peripancreatic vessels on CT imaging, yet expert assessment often shows substantial variability.

By Vincent Ochs, Christoph Kuemmerli, Florentin Bieder, Julia Wolleb, Joel L. Lavanchy, Julia Ruppel, Jan Liechti, Stephanie Taha-Mehlitz, Christian Andreas Nebiker, Beat Mueller, Giuseppe Kito Fusai, Joerg-Matthias Pollok, Anas Taha, Philippe C. Cattin, Sebastian Staubli