arXiv:2606. 09169v1 Announce Type: new Abstract: In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.
By Lingyi Meng, Zecong Tang, Haoran Li, Tengju Ru, Zhejun Cui, Weitong Lian, Qi Kang, Hangshuo Cao, Yichen Zhu, Yechi Liu, Kaixuan Wang, Yu-Jie Yuan, Chunwei Wang, Yu Zhang, Bo Dai
Hy‑MultiTurn is a Chinese benchmark designed to evaluate deep multi‑turn dialogue understanding over long interactions. It introduces six controlled evaluation modes—constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution—across 209 tasks ranging from 12 to 76 turns, incorporating dialogue length, irrelevant distractions, and colloquial phrasing. Testing 22 state‑of‑the‑art models shows the benchmark is highly challenging, with even the best model meeting all criteria only 41.1% of the time and no model excelling in every mode.
By Eileen Ye, Jiawen Tao, Yaoming Li, Chenxu Liu, Wenhan Yu, Yaxin Fan, Xiaokun Yuan, Mengzhou Wu, Yanbing Jiang, Maxm Pan
TurnBench is a new multi‑domain benchmark for evaluating turn‑taking dynamics in spoken dialogue. It comprises a 30‑hour hand‑labeled corpus of dyadic human conversations, a standardized evaluation protocol for end‑of‑turn and interruption detection, and covers six distinct interaction styles with triple annotation. The benchmark also provides a 104‑hour training set, a public leaderboard, and an interactive dataset viewer at https://turnbench.sesame.com.
By Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe
The paper introduces Omni Demand Understanding (ODU), a benchmark designed to test whether multimodal models can infer a user's underlying demand from complex audio‑visual interactions. ODU requires models to detect the presence of a demand and infer intent using multimodal and conversational context, evaluated across single‑turn and multi‑turn scenarios. The authors built ODU‑Bench through a taxonomy‑guided approach, agentic video generation, and human‑recorded interactions, and found that even top models like Gemini 3.1 Pro recover only 44.7% of key information, with many models exhibiting high false‑trigger rates.
By Qi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu, Yuxuan Wang, Ziyang Ma, Ruiyang Xu, Meng Gao, Yinsong Yan, Ling Wang, Hui Wang, Wen Huang, Yiheng Chen, Guanrou Yang, Qiuqiang Kong, Jin Xu, Xie Chen
arXiv:2609.13076v1 Announce Type: cross
Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-en...
By Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath
arXiv:2609.00802v1 Announce Type: new
Abstract: Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate i...
By Taiga Mori, Koji Inoue, Mikey Elmers, Divesh Lala, Tatsuya Kawahara
arXiv:2609.09852v1 Announce Type: new
Abstract: The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech intera...
By Heyang Liu, Jiayi Huang, Wenyang Xiao, Ziyang Cheng, Lixin Zhang, Zhen Liu, Miao He, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang
arXiv:2609.08977v3 Announce Type: replace-cross
Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
By Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao
arXiv:2607.26178v2 Announce Type: replace
Abstract: Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current mo...
By Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-T\"ur
arXiv:2606.11167v2 Announce Type: replace
Abstract: Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, cur...
By Atsumoto Ohashi, Neil Zeghidour, Alexandre D\'efossez, Eugene Kharitonov
arXiv:2507. 19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework.
By Sara Papi, Maike Z\"ufle, Marco Gaido, Beatrice Savoldi, Danni Liu, Ioannis Douros, Luisa Bentivogli, Jan Niehues
The paper introduces a pipeline that generates intent‑labeled, two‑channel conversational speech from relational event lists, enabling controlled synthesis of full‑duplex dialogue with 42 phenomena across eight families in English and Mandarin. By having an LLM author each event’s speaker, text, conversational act, and attachment, and then aligning and timing these events independently, the system produces diverse, realistic turn‑taking signals. Experiments show that models trained on this synthetic corpus achieve higher floor‑occupancy accuracy and better start‑speaking/listening F1 scores compared to models trained on prior data.
By Matthew Sun, Vinay Kothapally, Meng Yu, Chao Huang, Hao Zhang, Yixuan Zhang, Steve Yves