The paper introduces Inverse Turing Bench, a benchmark designed to assess how well language models can distinguish between human-only and human-AI dialogues in multi-turn text. It provides paired dialogue transcripts and evaluates models on correctly identifying the type of conversation. Preliminary results show GPTZero, Claude Opus-4.6, and GPT-5.5 achieving the highest accuracies of 89.41%, 77.92%, and 75.94% respectively, highlighting both the strengths and limitations of statistical versus semantic detection approaches.
By William Hager, Ishika Rathi, Masum Hasan, Cameron Jones
The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.
By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv:2601. 02813v3 Announce Type: replace Abstract: Aligning language models to qualitative behavioral traits, such as human-likeness, remains difficult because they are hard to define, measure, and optimize.
By Masum Hasan, Junjie Zhao, Ehsan Hoque
arXiv:2605. 28882v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important.
By Yihang Lin, Yunze Gao, Zeyang Lin, Dongbo Li, Kun Peng, Yue Liu
arXiv:2606. 31729v1 Announce Type: cross Abstract: Text-to-speech (TTS) evaluation is an open challenge.
By Dominika Woszczyk, Andreas Triantafyllopoulos, Jura Miniota, \'Eva Sz\'ekely, Bjoern Schuller
arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.
By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran