arXiv Computation and Language

Evaluating Communicative Success in Machine-Translated Conversation

The paper introduces a three‑layer checklist-and-judge framework to evaluate interpreter agents that mediate live conversation across languages. It assesses semantic, pragmatic, and cultural‑social dimensions—naturalness, intent, and social appropriateness—rather than just fidelity, in both single‑turn and multi‑turn settings. Extensive validation shows that conventional MT metrics miss failures in stronger interpreters, and that context, structured instructions, and cultural cues influence communicative success.

arXiv AI
Sep 2

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench is a new multilingual benchmark that tests large language model agents on culturally grounded everyday workflows, offering 1,600 tasks in seven languages and eight cultures. The benchmark evaluates agents through structured sandbox actions and introduces Constrained Task Success (CTS), a metric that assesses task completion, minimal modification, and other complementary aspects via deterministic and LLM-as-a-Judge evaluations. Experiments show that even leading models achieve only 49.2% CTS, revealing significant gaps in correctness and state preservation across languages and cultures.

By Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
arXiv Computation and Language
Aug 28

Evaluating Language Models in Realistic Conversational Contexts

The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.

By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv AI
Jul 15

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.

By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
arXiv AI
Sep 4

Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

The paper proposes treating translation as a structured decision space explored by multiple autonomous agents, rather than producing a single output. Using Turkish–Syrian Arabic dialogue, three agents—zero‑shot, dialect‑stabilized, and pivot translation—are compared on 5,000 sentences, with stabilization nearly doubling dialect marker usage and reducing structural instability. The study introduces an interpretability framework that quantifies decision flexibility through dialect marker frequency, lexical proximity, and structural variance.

By Hasan Alkhder, Mohammad Abboush, Igor Tchappi, Ahmet Zengin, Amro Najjar
arXiv Computation and Language
Aug 31

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...

By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury