arXiv AI By Yuejie Li, Ke Yang, Yueying Hua, Berlin Chen, Jianhao Nie, Yueping He, Caixin Kang

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

Read the original on arXiv AI →

arXiv:2602. 12783v3 Announce Type: replace-cross Abstract: Spoken query retrieval is an important interaction mode in modern information retrieval.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval

VoiceTrace introduces a new benchmark, VoiceTrace-Bench, for hybrid speech retrieval that combines a textual query specifying "what" to retrieve with a reference speech specifying "who" to retrieve. The authors propose a two‑stage framework: VoiceTrace‑Emb, which learns unified audio‑text embeddings for efficient large‑scale retrieval, and VoiceTrace‑Reranker, which fine‑grains relevance by jointly examining query‑candidate pairs. Experiments show VoiceTrace outperforms existing methods on both traditional semantic speech retrieval benchmarks and the new hybrid setting.

By Aaron Yee, Fengjie Lu, Jiarui Hai, Chenang Jiang, Helin Wang, Siwei Tu, Weitao You, Lingyun Sun
arXiv AI
Jul 17

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.

By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.