arXiv AI

LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot

The paper introduces an LLM-based Conversational AI Knowledge Assistant for the Raspberry‑Pi‑powered 13‑Axis MyBuddy humanoid robot. It combines large language model-driven language understanding, real‑time speech recognition, internet‑based knowledge retrieval (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to support intelligent, multi‑turn conversations and emotional‑support interactions. This system aims to overcome the limitations of traditional rule‑based dialogue systems in humanoid robots.

arXiv AI
Sep 4

IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

IRWOZ 2.0 is a refined dialogue dataset for industrial human‑robot interaction, expanding to 390 dialogues across four domains—Assembly, Delivery, Position, and Relocation. The dataset was improved using large language models (Mistral and Claude‑3.5) for generation and quality refinement, including manual corrections and automated typo removal. Benchmark tests show a substantial boost in dialogue state‑tracking performance, with GPT‑2’s BLEU‑4 score rising from 0.1651 to 0.5604 compared to the original IRWOZ.

By Chen Li, Dimitrios Chrysostomou
arXiv Computation and Language
Sep 11

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

The paper introduces Inverse Turing Bench, a benchmark designed to assess how well language models can distinguish between human-only and human-AI dialogues in multi-turn text. It provides paired dialogue transcripts and evaluates models on correctly identifying the type of conversation. Preliminary results show GPTZero, Claude Opus-4.6, and GPT-5.5 achieving the highest accuracies of 89.41%, 77.92%, and 75.94% respectively, highlighting both the strengths and limitations of statistical versus semantic detection approaches.

By William Hager, Ishika Rathi, Masum Hasan, Cameron Jones
arXiv Computation and Language
4d ago

AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction

arXiv:2609.37853v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or per...

By Wentao Liu, Xi Chen, Siyu Song, Biao Yuan, Yu Zhang, Zhou Zhuotong, Jingying Zhou, Guohao Feng, Shasha Hu, Tianfu Wang, Shangshang Yang, Haoyang Liu, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
arXiv AI
Aug 21

Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents

arXiv:2608. 19794v1 Announce Type: new Abstract: The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI).

By Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao
arXiv AI
Sep 2

Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations

Conversation Coach is a voice‑first AI system that lets managers rehearse difficult workplace conversations in a realistic spoken format. It tackles low‑latency interaction, adaptive bot personalities that simulate various employee types, and personalized feedback on content and policy compliance. The authors compare an end‑to‑end speech‑to‑speech model with a cascaded approach, finding the former offers lower latency and cost, while the cascaded model provides better reasoning for coaching quality, and they deployed the cascaded architecture to 40,000+ managers over six months.

By Fanyou Wu, Suraj Maharjan, Ainur Yessenalina, Dennis Xu Chen, Rahul Srivastava, Srinivasan H. Sengamedu
arXiv AI
Sep 10

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

The paper reports the first Turing test for speech‑to‑speech systems, gathering 2,968 human judgments on conversations between nine state‑of‑the‑art S2S systems and 28 humans. None of the evaluated systems passed the test, highlighting a clear gap in human‑likeness. The authors diagnose the failure with an 18‑dimension taxonomy, finding that paralinguistic cues, emotional expressivity, and conversational persona—not semantic understanding—are the main bottlenecks, and they propose an interpretable model for automatic human‑vs‑machine discrimination.

By Xiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou, Chi Zhang, Bo Cheng, Jiale Han, Benyou Wang