arXiv:2605.01278v3 Announce Type: replace
Abstract: In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified unders...
By Zeyu Chen, Guanghao Zhou, Min Yang, Qixiang Yin, Ziwang Zhao, Huanjin Yao, Pengjiu Xia, Cen Chen, Minghui Qiu
The paper introduces GPUB, a large-scale benchmark for grounded product understanding in e‑commerce livestream videos, featuring 3,000 livestreams, 31K fashion products, and multi‑moment temporal annotations. It defines three evaluation tasks, with the main task (GPrU) requiring simultaneous product identification and moment localization. Existing multimodal models perform poorly on GPrU, prompting the authors to develop UniPro, which improves performance by learning product‑aligned, temporally structured representations.
By Xinyu Zhang, Junjie Chen, Jiawei Ge, Qianlong Li, Libin Ma, Baokun Pan, Yahui Luo
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.
By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
arXiv:2604. 00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention.
By Junxian Wu, Chenghan Fu, Zhanheng Nie, Daoze Zhang, Bowen Wan, Wanxian Guan, Chuan Yu, Jian Xu, Bo Zheng
arXiv:2606. 06991v1 Announce Type: cross Abstract: Online Video Large Language Models (Video-LLMs) have advanced toward seamless human-AI interaction through frame-by-frame processing and proactive responding.
By Zhenyu Yang, Kairui Zhang, Shengsheng Qian, Weiming Dong, Changsheng Xu
The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.
By Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan