arXiv:2607. 23813v1 Announce Type: cross Abstract: We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions.
By Denglin Jiang, Haoran Zhou, Anshul Wadhawan, Brendan Fahy, Vinay Ramesh, David Weisberg, Dmitriy Derkachevskiy, Helen Sheehan, Srivas Prasad, Michele Franceschini
The paper introduces DualEvasion, a benchmark that evaluates evasion detection in earnings call Q&A using both textual transcripts and vocal cues. It contains 505 annotated question‑answer pairs from 60 calls, each labeled for textual evasion (direct vs. evasive) and speaker confidence (confident vs. unconfident). Experiments show that current multimodal models struggle to detect vocal confidence, especially in unconfident responses, and that providing speaker‑level references only modestly improves performance, leaving a significant gap compared to humans.
By Mirae Kim, Seonghun Jeong, Youngjun Kwak
arXiv:2609.38523v1 Announce Type: cross
Abstract: Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle co...
By Dong Shu, Yanguang Liu, Huopu Zhang, Saisai Hu, Haiyan Zhao, Hekun Huang, Mengnan Du
arXiv:2609.13893v1 Announce Type: new
Abstract: Earnings conference calls are a primary channel through which managers disclose information under analyst scrutiny. Prior work has linked vocal and lex...
By Huizhong Chen, Huan Zhang
arXiv:2610.00969v1 Announce Type: cross
Abstract: Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generati...
By Yingzhu Zhao, Vlad Pandelea, Han Yuan, Bo Hu, Wuqiong Luo, Li Zhang, Zheng Ma
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.
By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.
arXiv:2606. 28002v1 Announce Type: cross Abstract: Insurance fraud imposes substantial financial losses and operational inefficiencies, raising premiums and impacting trust among legitimate policyholders.
By Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly
The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.
By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
TRILOGUE is a new trilingual benchmark for spoken dialogue fact‑checking, covering English, Russian, and Kazakh. It includes almost 12,000 dialogues, 187,000 turns, and 390 hours of paired audio with ASR transcripts and word‑level timestamps, as well as nearly 5,000 human‑recorded Russian and Kazakh files. The dataset supports tasks such as claim check‑worthiness detection, evidence retrieval, and claim verification under various input conditions, and baseline experiments reveal challenges with ASR errors and cross‑lingual transfer, especially for Kazakh.
By Chaewan Chun, Meruyert Aristombayeva, Jiyoung Choi, Mahjabin Nahar, Delvin Ce Zhang, Dongwon Lee
arXiv:2608.22196v1 Announce Type: cross
Abstract: While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during sepa...
By Hermann Yepdjio Nkouanga, Minwei Luo, Maggie Wigness, Suresh Singh