arXiv AI By Xuzhao Geng, Haozhao Wang, Xuelian Li, Zhenyu Yang, Haonan Lu, Rui Zhang, Ruixuan Li

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

Read the original on arXiv AI →

arXiv:2607. 22635v1 Announce Type: new Abstract: Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

Hy‑MultiTurn is a Chinese benchmark designed to evaluate deep multi‑turn dialogue understanding over long interactions. It introduces six controlled evaluation modes—constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution—across 209 tasks ranging from 12 to 76 turns, incorporating dialogue length, irrelevant distractions, and colloquial phrasing. Testing 22 state‑of‑the‑art models shows the benchmark is highly challenging, with even the best model meeting all criteria only 41.1% of the time and no model excelling in every mode.

By Eileen Ye, Jiawen Tao, Yaoming Li, Chenxu Liu, Wenhan Yu, Yaxin Fan, Xiaokun Yuan, Mengzhou Wu, Yanbing Jiang, Maxm Pan
arXiv Computation and Language
Sep 25

Same Words, Different Actions: Paired Turn-Taking Evaluation under Rewritten Dialogue Contexts

The paper introduces ECHO, a paired diagnostic benchmark for evaluating Chinese real‑time spoken dialogue systems on turn‑taking. ECHO pairs examples that share the same overlap transcript but differ in preceding multi‑turn context, requiring either Yield or Keep actions, and also includes off‑talk cases to test unnecessary yielding. Experiments on four speech systems reveal that three systems over‑yield, correctly keeping the floor on fewer than 13% of backchannels, while the fourth system shows a more balanced performance, illustrating that interruption‑only evaluation can overestimate turn‑taking reliability.

By Shuofeng Zhao, Hongwei Cai, Wenke Fan, Qingxiang Guo, Zhou Wang, Dawei Yang, Zhiyang Zhou, Yingxin Shang, Weixu Wang, Lin Yang, Shuran Zhou, Yang Song
arXiv AI
Aug 25

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

CallScreenBench is a benchmark for evaluating small, on-device language models that act as phone secretaries, focusing on their ability to handle unknown inbound calls without a cooperative task. The benchmark measures owner endorsement through five call-and-note metrics, each paired with counter-metrics and uncertainty estimates, and includes guardedness diagnostics to identify safe, tool‑free proxies. Results across 4‑bit checkpoints of 0.6‑4 B parameter models show varying performance on service, recall, plausibility, and triage discrimination, highlighting trade‑offs between quality and guardedness.

By Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren
arXiv Computation and Language
Sep 1

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

arXiv:2607.26178v2 Announce Type: replace Abstract: Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current mo...

By Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-T\"ur