Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv Computer Vision
Sep 23

When Point Clouds Outperform Pixels: Rethinking Zero-Shot Multimodal Anomaly Detection

The paper challenges the common assumption that RGB and point cloud data contribute equally to zero‑shot multimodal anomaly detection. It shows that point clouds are more reliable under category shift and introduces WOOPS, a framework that enhances point cloud features with a Multi‑view Information Decoupling module and calibrates modality contributions via a Modality Reliability Calibration module. Experiments demonstrate that WOOPS achieves top performance on new stringent metrics in both unimodal and multimodal settings, and that point cloud information also benefits RGB‑only inference.

By Chenglin Ye, Lupeng Liu, Dongbo Yu, Jun Xiao, Yunbiao Wang
arXiv Computer Vision
Sep 23

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.

By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
arXiv Computer Vision
Sep 23

LLaVA-Assessor: Building the Foundation LMM For Visual Quality Assessment

LLaVA‑Assessor is a unified large multi‑modal model (LMM) designed for visual quality assessment, combining image and video inputs. It introduces a two‑task framework—quality interpretation and quality scoring—supported by an adaptive architecture, a rigorous human‑annotated dataset, and a machine‑synthesized data expansion pipeline. The model employs a prompt‑disentanglement strategy to stabilize multi‑task training and achieves strong performance across 11 quality scoring test sets and 4 interpretation benchmarks.

By Ziheng Jia, Zicheng Zhang, Jiaying Qian, Guangtao Zhai, Xiongkuo Min
arXiv Computer Vision
Sep 23

From Token Importance to Conditional Removability: Rethinking Visual Token Pruning in Multimodal Large Language Models

The paper argues that token importance alone is insufficient to determine safe removal of visual tokens in multimodal large language models, because removability depends on representation depth and the surrounding deletion set. Through controlled experiments, the authors show that the same tokens can have different effects when removed at different depths or contexts. They introduce CoRePrune, a training‑free two‑stage pruning framework that refreshes deletion effects as visual representations evolve and refines candidate tokens based on the current deletion set, achieving high performance retention across multiple backbones and reducing prefill time significantly.

By Shengli He, Yongchao Liang, Roumeng He, Junjie Zeng, Jiyuan He, Xin Fang, Can Wu, Li Zheng
arXiv Computer Vision
Sep 23

Semantically-Guided Domain Randomization for Industrial Object Detection in Low-Image-Budget Regimes

Semantically-Guided Domain Randomization (S‑GDR) is an annotation‑free pipeline that uses vision‑language model captioning of a small real reference set, diffusion‑based background synthesis, and mask‑based object composition to generate synthetic training data. In a high‑mix, low‑volume automotive detection benchmark, S‑GDR achieves a mAP50‑95 of 0.739 with only 200 synthetic images, outperforming a domain‑randomized render baseline and several other synthetic data methods under the same budget. These results suggest S‑GDR is a viable alternative for training visual perception systems when annotation, energy, and time resources are severely limited.

By Jose Moises Araya-Martinez, Gautham Mohan, Jens Lambrecht
arXiv Computer Vision
Sep 23

Virtual Encoders in Multimodal Transformers

The paper investigates how multimodal language models can generate perceptual representations without dedicated encoders. It shows that the shared transformer can internally create these representations in its early-to-middle layers, a structure termed a Virtual Encoder. Experiments with linear probing, similarity metrics, and causal analysis reveal that this encoder-like computation emerges even when models receive only perceptual tokens, indicating that perception and language processing can be decoupled within a single architecture.

By Katsuya Ogata, Yuta Nakashima
arXiv Computer Vision
Sep 23

GAD-MambaUNet: Direction-Group Mamba with Gradient-Adaptive DINOv3 Distillation for Lightweight Medical Image Segmentation

The paper introduces GAD-MambaUNet, a lightweight medical image segmentation network that integrates efficient local modeling, Direction-Group Graph Selective Scan (DG‑GSS) for structured information exchange, and training‑time supervision from a frozen DINOv3 teacher with Gradient‑Adaptive Distillation. GAD‑MambaUNet demonstrates a strong accuracy‑efficiency trade‑off compared to other lightweight and general segmentation methods, and ablation studies confirm the benefits of DG‑GSS and DINOv3‑GAD supervision. Future work aims to refine teacher‑student alignment and apply the framework to multi‑class and multi‑modal medical segmentation tasks.

By Fang Wang, Huitao Li, Wenhan Chao, Zheng Zhuo, Xinxin Yang
arXiv Computer Vision
Sep 23

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

HARMONY is a hierarchical chain-of-thought framework that reconstructs complete 3D indoor scenes from a single monocular image. It combines agentic reasoning with visual geometry foundation models, starting with camera calibration and semantic layout recovery, then placing objects hierarchically while refining geometry with point cloud estimations. The method achieves semantically consistent scenes that align perceptually with the input image, outperforming existing baselines on synthetic and real-world data.

By Shufan Sun, Chen Wang, Enxin Song, Jiatao Gu, Lingjie Liu
arXiv Computer Vision
Sep 23

Neoadjuvant chemotherapy response prediction using pretreatment diffusion and contrast-enhanced magnetic resonance imaging with clinical variables

arXiv:2609.26105v1 Announce Type: cross Abstract: Prediction of pathological complete response before neoadjuvant chemotherapy may facilitate more tailored therapeutic planning for breast cancer pati...

By Pablo Garc\'ia Marcos, Paula Puerta Gonz\'alez, Guillermo Lorenzo, H\'ector G\'omez, Covadonga del Camino, Ad\'an Rodr\'iguez, Ignacio Pel\'aez, Angel Rio-Alvarez, V\'ictor M. Gonz\'alez
arXiv Computer Vision
Sep 23

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

arXiv:2505.01583v2 Announce Type: replace Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...

By Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou, Vivian Wang, Huayu Wang, Hsiang-Wei Huang, Wenhao Chai, Hou-I Liu, Kuang-Ming Chen, Cheng-Yen Yang, Yi-Ling Chen, Vibhav Vineet, Qin Cai, Jenq-Neng Hwang
arXiv Computation and Language
Sep 23

MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis

MultiViewDx is a physician‑validated multimodal instruction dataset that links medical imaging studies with patient context and normalizes heterogeneous reports into an evidence‑linked workflow (evidence → findings → differential discussion → diagnosis). The dataset covers a wide range of imaging modalities and uses a unified image‑text retriever to ensure that instruction synthesis is grounded in source‑supported evidence. Fine‑tuned models on MultiViewDx achieve the highest average accuracy on four MedVQA benchmarks and receive the strongest overall rating on JAMA Clinical Challenge cases, with ablations confirming the importance of case‑level multi‑view organization and evidence‑linked reasoning.

By Junda Wang, Zonghai Yao, Yujan Ting, Eric Z. Chen, Hieu Tran, Hong Yu, Weijing Huang, Terrence Chen