arXiv AI By Kanpat Vesessook, Saksorn Ruangtanusak

SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models

Read the original on arXiv AI →

SCB: SpeechConversationBench (SCB) evaluates multi-turn reasoning in speech-to-speech models by using 103 sharded GSM8K mathematical problems. The benchmark compares three delivery modes: a full single-turn problem, a concatenated multi-shard version, and an incremental spoken disclosure across turns. Results show that sharded accuracy drops by 5.0–25.3 percentage points for four commercial systems, while the internal LEGO pipeline maintains 77.5% accuracy across all conditions, outperforming GPT‑4o Realtime’s 76.6% sharded accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 14

MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

arXiv:2609.13076v1 Announce Type: cross Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-en...

By Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath
Hugging Face Trending Papers
Sep 3

Scalable Context Orchestration for Serving LLMs Over Voice

Scalable Context Orchestration for Serving LLMs Over Voice presents llmovoice, a middleware that explicitly models voice context—including speaking rate, background noise, and packet loss—to guide large language model responses. By constructing a bounded voice context at each turn, llmovoice improves alignment with user preferences and reduces errors, achieving a 52.4% drop in speaking‑rate alignment error and a 0.9% false‑interruption rate under packet loss. In addition, it cuts model usage costs dramatically, lowering per‑turn cost by up to 24.9× while maintaining 98.7% of baseline answer quality in long sessions.

arXiv Computation and Language
Sep 22

The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding

arXiv:2609.22214v1 Announce Type: new Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...

By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
arXiv AI
Sep 7

Scalable Context Orchestration for Serving LLMs Over Voice

The paper introduces llmovoice, a middleware that explicitly models voice context for large language model (LLM) serving in voice AI applications. By incorporating speaking rate, background noise, packet loss, and other paralinguistic factors into a bounded context, llmovoice guides the LLM to generate more aligned responses. Experiments show significant reductions in speaking‑rate errors, false interruptions, and model usage costs, especially in long voice sessions.

By Linyi Jiang, Silvery D. Fu, Yifei Zhu