Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Computation and Language
23h ago

Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech

Nord-Parl-TTS is an open text‑to‑speech dataset for Finnish and Swedish created from recordings of Nordic parliamentary proceedings. It contains 900 hours of Finnish and 5 090 hours of Swedish speech, processed with an adapted Emilia pipeline and accompanied by unified evaluation sets for model development and benchmarking. The dataset aims to reduce the resource gap in TTS between high‑ and lower‑resourced languages.

By Zirui Li, Jens Edlund, Yicheng Gu, Nhan Phan, Lauri Juvela, Mikko Kurimo
arXiv Computer Vision
23h ago

Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning

The paper investigates how vision‑language models acquire and specialize semantic capabilities during fine‑tuning, proposing a trajectory‑based framework that separates acquisition, optima, and specialization phases. It introduces Structured Semantic Routing (SSR) to analyze how supervision representation affects what is learned, showing that SSR improves name‑free attribute‑profile retrieval and that different capabilities peak at different training stages. The study reveals that continued optimization can preserve target‑class retrieval while diminishing transferable semantic knowledge, highlighting a trade‑off between specialization and generalization.

By Suguru Onda, Matthew Bailey, Ryan Farrell
arXiv Computer Vision
23h ago

Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition

The paper introduces Latent-Action-Guided Video-Language Feature Learning (LAG-VLFL) for recognizing surgical instrument–tissue interactions. By compressing frame-to-frame feature changes into latent actions and predicting next‑frame features, the method aligns video and textual action descriptions without requiring extra spatial or motion annotations. Experiments show that LAG-VLFL improves interaction grounding, temporal‑direction sensitivity, and achieves competitive recognition with faster inference and lower INT4 accuracy loss compared to V‑JEPA2/2.1.

By Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin
arXiv AI
23h ago

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

Rationale-Guided Policy Optimization (RGPO) is a reinforcement‑learning framework that adaptively uses ground‑truth rationale information to scaffold a language model’s reasoning process. Instead of treating reference solutions as fixed imitation targets, RGPO temporarily incorporates rationales to help the model generate better responses, then reverts to unguided learning with higher‑reward, model‑generated solutions. Experiments in both language‑only and vision‑language tasks show that RGPO consistently outperforms RLVR baselines, with ablation studies confirming that adaptive rationale guidance is a key factor in its success.

By Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei
arXiv Computer Vision
23h ago

Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update

arXiv:2610.07031v1 Announce Type: new Abstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics a...

By Sitian Shen, Jiuming Liu, Mengmeng Liu, Yian Wang, Michael Ying Yang, Francesco Nex, Hao Cheng, Daniele De Martini, Ayush Tewari, Per Ola Kristensson
arXiv Computer Vision
23h ago

World Models' Last Exam in Physics

arXiv:2610.08791v1 Announce Type: new Abstract: Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and p...

By Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin, Ziming Qin, Zheng Jiang, Wenyi Li, Calvin Xiao, Youjie Zheng, Kaisen Yang, Qinhuai Na
arXiv AI
23h ago

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

The paper introduces Physics‑Guided Visual Prompting (PG‑VP), a plug‑and‑play module that overlays a virtual obstacle onto the input of a frozen Vision‑Language‑Action model to guide navigation around invisible hazards such as radiation or temperature spikes. PG‑VP performs a physics‑based risk assessment to determine the avoidance direction and dynamically renders the same virtual obstacle across frames, allowing the existing navigation policy to detour without retraining. Experiments on OmniNav with R2R‑CE and RxR‑CE datasets show that PG‑VP steers the policy toward low‑risk actions in 84.9% and 83.2% of cases, while real‑world tests on a robot demonstrate significant safety improvements against thermal and radiation sources.

By Hojoon Son, Fan Zhang