arXiv Computation and Language

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

The paper introduces Inverse Turing Bench, a benchmark designed to assess how well language models can distinguish between human-only and human-AI dialogues in multi-turn text. It provides paired dialogue transcripts and evaluates models on correctly identifying the type of conversation. Preliminary results show GPTZero, Claude Opus-4.6, and GPT-5.5 achieving the highest accuracies of 89.41%, 77.92%, and 75.94% respectively, highlighting both the strengths and limitations of statistical versus semantic detection approaches.

arXiv AI
Sep 10

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

The paper reports the first Turing test for speech‑to‑speech systems, gathering 2,968 human judgments on conversations between nine state‑of‑the‑art S2S systems and 28 humans. None of the evaluated systems passed the test, highlighting a clear gap in human‑likeness. The authors diagnose the failure with an 18‑dimension taxonomy, finding that paralinguistic cues, emotional expressivity, and conversational persona—not semantic understanding—are the main bottlenecks, and they propose an interpretable model for automatic human‑vs‑machine discrimination.

By Xiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou, Chi Zhang, Bo Cheng, Jiale Han, Benyou Wang
arXiv Computation and Language
Aug 28

Evaluating Language Models in Realistic Conversational Contexts

The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.

By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv AI
Sep 4

IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

IRWOZ 2.0 is a refined dialogue dataset for industrial human‑robot interaction, expanding to 390 dialogues across four domains—Assembly, Delivery, Position, and Relocation. The dataset was improved using large language models (Mistral and Claude‑3.5) for generation and quality refinement, including manual corrections and automated typo removal. Benchmark tests show a substantial boost in dialogue state‑tracking performance, with GPT‑2’s BLEU‑4 score rising from 0.1651 to 0.5604 compared to the original IRWOZ.

By Chen Li, Dimitrios Chrysostomou
arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv Computation and Language
Aug 31

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

arXiv:2305.12474v4 Announce Type: replace Abstract: Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensiv...

By Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, Xipeng Qiu, Tianxiang Sun, Peng Li, Shiqiao Meng, Yanjun Zheng, Jun Zhan, Zhangyue Yin, Xiannian Hu, Guofeng Quan, Qixiang Wang
Hugging Face Trending Papers
Jul 6

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality.

arXiv AI
Jul 17

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.

By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis