Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Computer Vision
Oct 1

Learning to Reason with Compressed Context: Ground-Truth-Free Adaptation of OmniLLMs via Self-Distillation

arXiv:2609.39953v1 Announce Type: new Abstract: Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computa...

By Jianghao Wang, Ke Meng, Jian Li, Chi Cheng, Longyu Qi, Liyin Liang, Yifeng Qian, Chunbo Lai, Yutian Lin, Zeyu Wang
arXiv Computer Vision
Oct 1

What to Attend, What to Keep: Skill-Conditioned Visuotactile Representation with Progress-Guided Event Memory

arXiv:2609.38494v1 Announce Type: cross Abstract: Robotic manipulation integrates vision, touch, and language, whose importance shifts across stages: vision guides reaching, while touch, through its...

By Amir-Hossein Shahidzadeh, Seungjae Lee, Eadom Dessalene, Shanthosh Raaj Mohanram Mageswari, Soroush Etemad, Furong Huang, Cornelia Ferm\"uller, Yiannis Aloimonos
arXiv Computer Vision
Oct 1

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...

By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin
arXiv Computer Vision
Oct 1

Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification

arXiv:2603.24528v2 Announce Type: replace Abstract: Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that...

By Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bart{\l}omiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
arXiv Computer Vision
Oct 1

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

arXiv:2604.18484v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from co...

By Kangan Qian, ChuChu Xie, Yang Zhong, Jingrui Pang, Siwen Jiao, Sicong Jiang, Zilin Huang, Yunlong Wang, Kun Jiang, Mengmeng Yang, Hao Ye, Guanghao Zhang, Hangjun Ye, Guang Chen, Long Chen, Diange Yang