OmniVChat defines a native audio‑visual dialogue task where models receive raw audio and video from a user and produce text responses, eliminating the need for separate text queries or speech recognition. To address data scarcity and evaluation challenges, the authors introduce OmniVChat‑Studio, a multi‑agent engine that synthesizes single‑ and multi‑turn dialogues, and OmniVChat‑Bench, a benchmark assessing models across five dialogue ability categories. They also propose OmniVChat‑RL, a reinforcement‑learning reward that balances reply correctness, efficiency, and style, and demonstrate that training Qwen3‑Omni‑Instruct with this reward on synthesized data improves performance on both synthetic and human‑recorded benchmarks.
By Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng, Muzhi Zhu, Zheqi Dai, Haoning Xu, Dongchao Yang, Chunyat Wu, Zining Liang, Zhengxi Liu, Xiquan Li, Xie Chen, Xize Cheng, Qize Yang, Jin Xu, Qiuqiang Kong
arXiv:2601. 11007v2 Announce Type: replace Abstract: LLM role-playing aims to portray arbitrary characters in interactive narratives, yet existing systems often suffer from limited immersion and adaptability.
By Zhenhua Xu, Dongsheng Chen, Shuo Wang, Jian Li, Chengjie Wang, Meng Han, Yabiao Wang
WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.
By Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
arXiv:2606. 00616v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos.
By Shivam Singh, Saptarshi Majumdar, Pratik Prabhanjan, Zicheng Liu, Emad Barsoum
arXiv:2609.22255v1 Announce Type: new
Abstract: Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent...
By Rotem Dror, Zohar Elyoseph, Yuval Haber, Elad Refoua, Oshrat Ayalon, Adir Solomon
arXiv:2602. 13602v2 Announce Type: replace-cross Abstract: We present \revise (\underline{Re}asoning with \underline{Vi}deo \underline{S}parsity), a multi-round agent for video question answering (VQA).
By Chenwei Xu, Zhen Ye, Shang Wu, Weijian Li, Zihan Wang, Zhuofan Xia, Lie Lu, Pranav Maneriker, Fan Du, Manling Li, Han Liu