arXiv AI
Aug 25

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

CallScreenBench is a benchmark for evaluating small, on-device language models that act as phone secretaries, focusing on their ability to handle unknown inbound calls without a cooperative task. The benchmark measures owner endorsement through five call-and-note metrics, each paired with counter-metrics and uncertainty estimates, and includes guardedness diagnostics to identify safe, tool‑free proxies. Results across 4‑bit checkpoints of 0.6‑4 B parameter models show varying performance on service, recall, plausibility, and triage discrimination, highlighting trade‑offs between quality and guardedness.

By Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren
arXiv AI
22h ago

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

MTVA-Bench is a new benchmark designed to evaluate the language model component of cascaded voice agents under realistic conditions. It simulates callers with an LLM, mocks backend tool responses, and scores both tool‑call correctness and conversational quality using two LLM judges, covering 49 agents, 490 scenarios, and 7 languages. The benchmark reveals that while models perform similarly on tool selection, they differ widely in argument handling, rule compliance, and dialogue quality, highlighting the nuanced challenges of real‑world voice interactions.

By Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath