arXiv Machine Learning

MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games

arXiv:2602. 24188v2 Announce Type: replace-cross Abstract: We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games that require effective communication about private information.

arXiv Computation and Language
Sep 14

MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

arXiv:2609.13076v1 Announce Type: cross Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-en...

By Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath
arXiv Computation and Language
Sep 18

Evaluating Communicative Success in Machine-Translated Conversation

The paper introduces a three‑layer checklist-and-judge framework to evaluate interpreter agents that mediate live conversation across languages. It assesses semantic, pragmatic, and cultural‑social dimensions—naturalness, intent, and social appropriateness—rather than just fidelity, in both single‑turn and multi‑turn settings. Extensive validation shows that conventional MT metrics miss failures in stronger interpreters, and that context, structured instructions, and cultural cues influence communicative success.

By Faiz Ghifari Haznitrama, Alice Oh
Hugging Face Trending Papers
Jul 4

ProACT: Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration

Conversational agents are increasingly embedded in human collaborative work, yet they remain fundamentally passive and reactive: they respond to explicit user requests rather than proactively recognizing moments when a team would benefit from timely intervention as human collaborators often do. This reactive design substantially limits the use of agents as active participants in multi-user collaboration, where disagreements, ambiguous goals, forgotten constraints, underspecified plans, discussion loops, and imbalanced participation can gradually undermine group progress.

arXiv Computation and Language
Sep 1

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

arXiv:2607.26178v2 Announce Type: replace Abstract: Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current mo...

By Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-T\"ur
arXiv AI
Sep 25

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

PUBG Ally is an embodied, voice‑enabled AI teammate for PUBG: BATTLEGROUNDS that can perceive the game world, interpret player speech, and autonomously decide actions while keeping speech synchronized with gameplay. It combines a language‑model agent that uses a controlled interface to gather game information and a faster control layer for movement, combat, and recovery. The system was trained on nearly 39,000 real‑player sessions and evaluated through player feedback and preference comparisons, with live deployment requiring low‑latency on‑device execution and safety safeguards.

By Beomsoo Kim, Byeongju Kim, Dohyun Kim, Dongwon Kim, Eunchong Kim, Hongmin Kim, Hyeojung Im, Hyeonbin Hwang, Hyeonghwan Kim, Hyoseok Seol, Insub Im, Irene Chen, Jaeseung Jeon, Jimin Hong, Kiyoon Yoo, Minkyoung Park, Seohyeon Jung, Seungjun Chung, Sue Hyun Park, Sungwoo Kim, Youngin Cho, Yujeong Son, Kangwook Lee, Hyunseung Kim
arXiv AI
Sep 3

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

CivBench is an open‑source benchmark that evaluates language‑model agents in the long‑horizon, tool‑mediated game Civilization VI using the Model Context Protocol (MCP). Each episode lasts over 300 turns, generating thousands of tool calls across a 76‑tool action space, and includes a narration layer that translates visual game state into structured text. The study characterises agent behaviour across four model families, introducing Proactive Monitoring Rate (PMR) and RAG@10 as interface‑level metrics, and finds that agents often under‑monitor strategic state and fail to execute near‑term commitments despite tool access and explicit guidance.

By Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa
arXiv AI
Sep 1

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

arXiv:2605.27068v2 Announce Type: replace-cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Mo...

By Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu