IRWOZ 2.0 is a refined dialogue dataset for industrial human‑robot interaction, expanding to 390 dialogues across four domains—Assembly, Delivery, Position, and Relocation. The dataset was improved using large language models (Mistral and Claude‑3.5) for generation and quality refinement, including manual corrections and automated typo removal. Benchmark tests show a substantial boost in dialogue state‑tracking performance, with GPT‑2’s BLEU‑4 score rising from 0.1651 to 0.5604 compared to the original IRWOZ.
By Chen Li, Dimitrios Chrysostomou
arXiv:2606. 18747v1 Announce Type: cross Abstract: Expressive gestures are essential for natural and effective communication, complementing speech when verbal cues alone are insufficient (e.
By Chris Lee, Flora Salim, Benjamin Tag, Francisco Cruz
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution.
arXiv:2607. 11377v1 Announce Type: cross Abstract: Long-term physical coexistence with intelligent robots requires more than capable robot policies.
By Weiqi Jin, Peijun Tang, Kuncheng Luo, Baifu Huang, Binyan Sun, Haotian Yang, Shangjin Xie, Jianan Wang
arXiv:2607. 18985v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge.
By Jialian Li, Junhong Liu, Yuchen Cao, Weiran Guo, Jiaming Song, Xutao Wang, Yi Zhao, Jiangpin Liu, Jie Chen
The paper introduces Inverse Turing Bench, a benchmark designed to assess how well language models can distinguish between human-only and human-AI dialogues in multi-turn text. It provides paired dialogue transcripts and evaluates models on correctly identifying the type of conversation. Preliminary results show GPTZero, Claude Opus-4.6, and GPT-5.5 achieving the highest accuracies of 89.41%, 77.92%, and 75.94% respectively, highlighting both the strengths and limitations of statistical versus semantic detection approaches.
By William Hager, Ishika Rathi, Masum Hasan, Cameron Jones
arXiv:2606. 31158v1 Announce Type: cross Abstract: The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics.
By Snehasis Banerjee, Ranjan Dasgupta
arXiv:2607. 18985v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge.
By Jialian Li, Junhong Liu, Yuchen Cao, Weiran Guo, Jiaming Song, Xutao Wang, Yi Zhao, Jiangpin Liu, Jie Chen
arXiv:2609.37853v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or per...
By Wentao Liu, Xi Chen, Siyu Song, Biao Yuan, Yu Zhang, Zhou Zhuotong, Jingying Zhou, Guohao Feng, Shasha Hu, Tianfu Wang, Shangshang Yang, Haoyang Liu, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
arXiv:2608. 19794v1 Announce Type: new Abstract: The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI).
By Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao
Conversation Coach is a voice‑first AI system that lets managers rehearse difficult workplace conversations in a realistic spoken format. It tackles low‑latency interaction, adaptive bot personalities that simulate various employee types, and personalized feedback on content and policy compliance. The authors compare an end‑to‑end speech‑to‑speech model with a cascaded approach, finding the former offers lower latency and cost, while the cascaded model provides better reasoning for coaching quality, and they deployed the cascaded architecture to 40,000+ managers over six months.
By Fanyou Wu, Suraj Maharjan, Ainur Yessenalina, Dennis Xu Chen, Rahul Srivastava, Srinivasan H. Sengamedu
The paper reports the first Turing test for speech‑to‑speech systems, gathering 2,968 human judgments on conversations between nine state‑of‑the‑art S2S systems and 28 humans. None of the evaluated systems passed the test, highlighting a clear gap in human‑likeness. The authors diagnose the failure with an 18‑dimension taxonomy, finding that paralinguistic cues, emotional expressivity, and conversational persona—not semantic understanding—are the main bottlenecks, and they propose an interpretable model for automatic human‑vs‑machine discrimination.
By Xiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou, Chi Zhang, Bo Cheng, Jiale Han, Benyou Wang