Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,971 stories · RSS feed

arXiv Computation and Language
Sep 30

Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment

arXiv:2609.37837v1 Announce Type: new Abstract: Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly el...

By Haotian Deng, Wenbin Xing, Gang Xu, Tao He, Jinkai Zheng, Chun Li, Zheng Zhu, Ming Li
arXiv Computation and Language
Sep 30

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv AI
Sep 30

MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis

MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.

By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
arXiv Computation and Language
Sep 30

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

arXiv:2606.28249v2 Announce Type: replace-cross Abstract: Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised...

By Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao, Jingyuan Xing, Baiji Liu, Xiangmin Xu