Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,303 stories · RSS feed

arXiv Machine Learning
Jul 16

GFlowRL: Scaling Distribution-Matching RL to Large Language Models

arXiv:2607. 13394v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes.

By Xiaodong Liu, Michael Xu, Jack W. Stokes, Paul Smolensky, Doug Burger, Jianfeng Gao
arXiv AI
Jul 16

LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

arXiv:2603. 13952v3 Announce Type: replace-cross Abstract: In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, their correlation with perceived speech quality is often suboptimal and provides limited interpretability for optimization.

By Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang, Shao-Yi Chien, Yu Tsao, Fan-Gang Zeng
arXiv AI
Jul 16

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

arXiv:2607. 13125v1 Announce Type: cross Abstract: We introduce Boogu-Image-0.

By Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, Binxiao Huang, Ming Zhang, Cheuk Him Chau, Xinyu Fu, Yingzhao Lian, Tom S. Y. Li, Jintao Lin, Bowen Dong, Zian Qian, Yuhao Liu, Yuxuan Hu, Weikang Shi, Bin Zou, Bowen Zheng, Haoxuan Che, Chang Chen, Yuyang He, Heyang Sun, Tianyu Huang, Chong Hou Choi, Cheng Gong, Han Shi, Haoli Bai, Xihui Liu, Hongsheng Li, Qifeng Chen, Chao Huang, Rui Liu, Chenyang Lei
arXiv Machine Learning
Jul 16

BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping

arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.

By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
arXiv AI
Jul 16

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

arXiv:2607. 13621v1 Announce Type: new Abstract: Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode.

By Kun Yu, Jianhua Yang, Yixiang Chen, Changwei Wang, Hongyuan Yu, Yan Huang, Fushuo Huo, Ya Jing, Zhumin Chen, Keji He
arXiv Machine Learning
Jul 16

A novel unsupervised machine learning strategy to handle multimodal cardiac PET/MRI data

arXiv:2607. 13936v1 Announce Type: cross Abstract: Arrhythmogenic left ventricular cardiomyopathy is a genetic myocardial disease difficult to diagnose due to the lack of gold standard criteria.

By Brunnhilde Ponsi (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France), Thomas Carlier (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France), Lara Marteau (Nantes Universit\'e, CHU Nantes, Nantes, France, Cardiology Department, INSERM UMR 1307, CIC 1413, l'institut du Thorax, Nantes, France), Aur\'elien Monnet (Siemens Healthineers France, Courbevoie, France), Thomas Eug\`ene (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France), Jean-Michel Serfaty (Nantes Universit\'e, CHU Nantes, Nantes, France, Radiology Department, l'institut du Thorax, Nantes, France), Nicolas Piriou (Nantes Universit\'e, CHU Nantes, Nantes, France, Cardiology Department, INSERM UMR 1307, CIC 1413, l'institut du Thorax, Nantes, France), Hatem Necib (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France)
arXiv AI
Jul 16

Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit

arXiv:2607. 13095v1 Announce Type: cross Abstract: We present a full-pipeline inference optimization for the MiMo-V2.

By Xiaomi MiMo Team, Anqi Liu, Aoxin Ma, Bo Chen, Bo Yang, Chen Wang, Chen Zhang, Chengda Tang, Chengwei Wang, Chiheng Lou, Depeng Yan, Fuli Luo, Gang Wang, Hailin Zhang, Jiale Sun, Kang Zhou, Rui Huang, Shaohui Liu, Shen Huang, Shijie Cao, Shuaishuai Fan, Tianling Zhou, Xiangwei Deng, Xueyang Xie, Xuli Wang, Yingchun Lai, Yu Yang, Yuan Zhang, Zhen Tang, Zhonghua Deng, Zihan Jiang
arXiv AI
Jul 16

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

arXiv:2607. 13712v1 Announce Type: cross Abstract: Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfaithful reasoning, which substantially undermine their faithfulness and practical utility.

By Zhixiao Zheng, Zheren Fu, Zhiyuan Yao, Chunxiao Liu, Dongming Zhang, Zhendong Mao