Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,022 stories · RSS feed

arXiv AI
Jul 16

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset

arXiv:2603. 11811v2 Announce Type: replace-cross Abstract: The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the prohibitive cost and scalability limits of human-in-the-loop collection paradigms.

By Yongzhong Wang, Keyu Zhu, Yong Zhong, Liqiong Wang, Jinyu Yang, Feng Zheng
arXiv Machine Learning
Jul 16

A novel unsupervised machine learning strategy to handle multimodal cardiac PET/MRI data

arXiv:2607. 13936v1 Announce Type: cross Abstract: Arrhythmogenic left ventricular cardiomyopathy is a genetic myocardial disease difficult to diagnose due to the lack of gold standard criteria.

By Brunnhilde Ponsi (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France), Thomas Carlier (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France), Lara Marteau (Nantes Universit\'e, CHU Nantes, Nantes, France, Cardiology Department, INSERM UMR 1307, CIC 1413, l'institut du Thorax, Nantes, France), Aur\'elien Monnet (Siemens Healthineers France, Courbevoie, France), Thomas Eug\`ene (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France), Jean-Michel Serfaty (Nantes Universit\'e, CHU Nantes, Nantes, France, Radiology Department, l'institut du Thorax, Nantes, France), Nicolas Piriou (Nantes Universit\'e, CHU Nantes, Nantes, France, Cardiology Department, INSERM UMR 1307, CIC 1413, l'institut du Thorax, Nantes, France), Hatem Necib (Nantes Universit\'e, CHU Nantes, Nantes, France, CRCI2NA, INSERM UMR 1307, Nantes, France)
Hugging Face Trending Papers
Jul 15

SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning

Vision Language Models (VLMs) demonstrate strong perceptual abilities but remain limited in tasks requiring analytical reasoning across multiple visual states, such as multi-image comparison, change detection, and multi-step visual inference. These capabilities are critical for real-world multimodal applications where reasoning must be grounded in systematic differences between visual contexts.

Hugging Face Trending Papers
Jul 15

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-language model grounds its predictions in visual evidence. Existing visual-intervention methods contrast policy behavior on original and modified images, yet assign supervision by the type of intervention rather than its observed effect.

Hugging Face Trending Papers
Jul 15

Fine-grained CLIP fine-tuning with self-annotated region alignment

Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the whole image to a text description. Considering the large data and computational burden in pre-training a vision-language model from scratch, a series of works aim to enhance the fine-grained ability of CLIP through a fine-tuning scheme.

Hugging Face Trending Papers
Jul 15

From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception

Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of language understanding, environment perception, and autonomous navigation.

Hugging Face Trending Papers
Jul 15

How Far Can Root Cause Analysis Go on Real-World Telemetry Data?

Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches. The OpenRCA dataset exemplifies these challenges: it is large-scale, multimodal, and lacks detailed domain knowledge, and yields consistently low accuracy across all existing methods.

Hugging Face Trending Papers
Jul 15

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.

Hugging Face Trending Papers
Jul 15

TRACE-PCa: Predicting Prostate Cancer Progression from Longitudinal MRI During Active Surveillance

Active surveillance (AS) is the preferred strategy for favorable-risk prostate cancer, yet current protocols rely on scheduled repeat biopsies, most of which reveal no progression and are unnecessary. Existing risk-stratification tools operate on single time-point imaging or depend on explicit lesion segmentation, limiting their ability to capture longitudinal change and excluding patients without an MRI-visible lesion.

arXiv AI
Jul 15

An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge

arXiv:2607. 12468v1 Announce Type: cross Abstract: We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time.

By Shuming Fang, Shuifei Zeng
arXiv AI
Jul 15

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

arXiv:2607. 09759v2 Announce Type: replace-cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest.

By Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu