arXiv:2603.02767v4 Announce Type: replace-cross
Abstract: Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield repre...
By Hanpeng Liu, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, Kaiwen Long, Yaqian Li, Kun He
arXiv:2603.26741v2 Announce Type: replace-cross
Abstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, languag...
By Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Yetong Sha, Qiyu Hu, Feng Liu, Siyu Huang, Qi Dai, Zhi-Qi Cheng
arXiv:2604.11751v2 Announce Type: replace-cross
Abstract: World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the...
By Quanyi Li, Lan Feng, Haonan Zhang, Wuyang Li, Letian Wang, Alexandre Alahi, Harold Soh
arXiv:2604.09531v2 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely be...
By Guanyu Zhou, Yida Yin, Wenhao Chai, Shengbang Tong, Xingyu Fu, Zhuang Liu
arXiv:2609.38413v1 Announce Type: new
Abstract: Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evide...
By Susan Liang, Jianmin Wu, Daxiang Dong
arXiv:2609.38444v1 Announce Type: new
Abstract: Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synt...
By Duowen Chen, Jinjin He, Gouthaman KV, Sandeep Bangalore Venkatesh, Bo Zhu
arXiv:2609.38541v1 Announce Type: new
Abstract: Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic...
By Donghao Zhou, Haoyang He, Fan Zhang, Hao Yang, Guisheng Liu, Xin Gao, Zhongwei Wan, Xingyuan Bu, Jie Wang, Qiangpeng Yang, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng
arXiv:2609.38777v1 Announce Type: new
Abstract: A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. Howeve...
By Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang
arXiv:2609.38968v1 Announce Type: new
Abstract: Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-mo...
By Zeyu Wang, Mingyu Ge, Haiyu Song, Haoran Duan
arXiv:2609.38979v1 Announce Type: new
Abstract: Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge...
By Chang Liu, Yu Tian, Rui Xie
arXiv:2609.39335v1 Announce Type: new
Abstract: Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing met...
By Zijing Qin, Jun Zhou, Ruicheng Zhang, Jiaqi Hou, Zunnan Xu, Ronghui Li, Zhenyu Xie, Xiu Li
arXiv:2609.39492v1 Announce Type: new
Abstract: Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes subst...
By Moseli Mots'oehli, Thulani Babeli
arXiv:2609.39899v1 Announce Type: new
Abstract: Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians...
By Bangwei Guo, Xiao Chen, Boris Mailhe, Jia Yao, Yiqing Wang, Ankush Mukherjee, Yikang Liu, Zheyuan Zhang, Hang Yu, Terrence Chen, Shanhui Sun
arXiv:2609.40362v1 Announce Type: new
Abstract: We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and q...
By Hongyuan Tao, Xinggang Wang, Lianghui Zhu, Yongkang Li, Yunchao Wei, Bin Feng, Shaoyu Chen, Qian Zhang, Chang Huang, Kai Yu
arXiv:2609.38182v1 Announce Type: cross
Abstract: Avatar-based multimodal empathetic response generation has emerged as a pivotal capability in human-centric systems, aiming to recognize user emotion...
By Xiaolin Chen, Xuemeng Song, Jinlan Fu, Weili Guan, Mong-Li Lee, Wynne Hsu
arXiv:2609.40007v1 Announce Type: cross
Abstract: A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show...
By Yatharth Agarwal, Vijay Raghunathan
arXiv:2505.05855v4 Announce Type: replace
Abstract: Multi-contrast super-resolution (MCSR) is crucial for enhancing MRI but current deep learning methods are limited. They typically require large, pa...
By Yinzhe Wu, Hongyu Rui, Fanwen Wang, Jiahao Huang, Zhenxuan Zhang, Haosen Zhang, Zi Wang, Guang Yang
arXiv:2607.09351v2 Announce Type: replace
Abstract: Severe downsampling makes single image super-resolution (SISR) an ill-posed problem, in which the language-guided multi-modal methods are especiall...
By Haotong Cheng, Yuxuan Li, Zijie Cui
arXiv:2609.34330v2 Announce Type: replace
Abstract: Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visu...
By Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang
arXiv:2609.34792v2 Announce Type: replace
Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...
By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi