Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,971 stories · RSS feed

arXiv Computer Vision
Oct 1

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

arXiv:2609.38541v1 Announce Type: new Abstract: Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic...

By Donghao Zhou, Haoyang He, Fan Zhang, Hao Yang, Guisheng Liu, Xin Gao, Zhongwei Wan, Xingyuan Bu, Jie Wang, Qiangpeng Yang, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng
arXiv Computer Vision
Oct 1

Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models

arXiv:2609.39899v1 Announce Type: new Abstract: Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians...

By Bangwei Guo, Xiao Chen, Boris Mailhe, Jia Yao, Yiqing Wang, Ankush Mukherjee, Yikang Liu, Zheyuan Zhang, Hang Yu, Terrence Chen, Shanhui Sun
arXiv Computer Vision
Oct 1

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

arXiv:2609.34330v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visu...

By Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang
arXiv Computer Vision
Oct 1

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

arXiv:2609.34792v2 Announce Type: replace Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...

By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
arXiv AI
Oct 1

Tactile Curiosity Drives Robot Interaction

The paper introduces TacEx, a tactile‑curiosity framework that guides reinforcement learning agents to explore contact dynamics by focusing epistemic uncertainty on the tactile channel. By anchoring curiosity to touch, robots learn to manipulate and grasp objects without task rewards or demonstrations, generating an interaction‑dense dataset that supports offline pick‑and‑place policy learning. TacEx also enhances vision‑language‑action models through post‑training, improving downstream performance while remaining sample‑efficient.

By Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza
arXiv Computation and Language
Oct 1

MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs

arXiv:2609.40041v1 Announce Type: new Abstract: We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGh...

By Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson, Kishankumar Bhimani, Jochebed Afua Basil