The paper introduces DualEvasion, a benchmark that evaluates evasion detection in earnings call Q&A using both textual transcripts and vocal cues. It contains 505 annotated question‑answer pairs from 60 calls, each labeled for textual evasion (direct vs. evasive) and speaker confidence (confident vs. unconfident). Experiments show that current multimodal models struggle to detect vocal confidence, especially in unconfident responses, and that providing speaker‑level references only modestly improves performance, leaving a significant gap compared to humans.
By Mirae Kim, Seonghun Jeong, Youngjun Kwak
arXiv:2606. 24523v1 Announce Type: cross Abstract: Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages.
By Arda Eren, Micheal Cheung, Youqian Zhang, Grace Ngai, Eugene Yujun Fu
arXiv:2609.13893v1 Announce Type: new
Abstract: Earnings conference calls are a primary channel through which managers disclose information under analyst scrutiny. Prior work has linked vocal and lex...
By Huizhong Chen, Huan Zhang
The paper introduces MAD2, a synthetic benchmark of 1,000 two‑speaker dialogues with about 10 hours of audio and 1,230 check‑worthy sentence annotations for spoken claim verification. It proposes a calibrated multimodal fusion approach that combines a context‑aware audio encoder with a dialogue‑aware text model. Experiments show that adding dialogue context improves verification performance, though the gains differ across scenarios, and that fusion offers the largest advantage when full‑dialogue context is available, though it does not consistently outperform text alone.
By Chaewan Chun, Delvin Ce Zhang, Dongwon Lee
The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.
By Vignesh Ethiraj, Ashwath David
TRILOGUE is a new trilingual benchmark for spoken dialogue fact‑checking, covering English, Russian, and Kazakh. It includes almost 12,000 dialogues, 187,000 turns, and 390 hours of paired audio with ASR transcripts and word‑level timestamps, as well as nearly 5,000 human‑recorded Russian and Kazakh files. The dataset supports tasks such as claim check‑worthiness detection, evidence retrieval, and claim verification under various input conditions, and baseline experiments reveal challenges with ASR errors and cross‑lingual transfer, especially for Kazakh.
By Chaewan Chun, Meruyert Aristombayeva, Jiyoung Choi, Mahjabin Nahar, Delvin Ce Zhang, Dongwon Lee
The paper introduces a benchmark for topic matching in real-world ASR transcripts from contact centers, where noisy, punctuation‑free speech data must be classified into predefined topics. It presents a human‑annotated dataset of topic‑utterance judgments and evaluates three matcher types—regex, zero‑shot sentence embeddings, and Gemini‑based LLMs—using two topic representations: keyphrases and natural language descriptions. Experiments show that lightweight LLM matchers outperform the other methods, especially when natural language descriptions are used.
By Saman Rahbar, Xiliang Zhu, Irvin Cardoza, David Rossouw
arXiv:2606. 28048v1 Announce Type: cross Abstract: Insurance fraud remains costly and operationally difficult, particularly in call-centre workflows where many customer interactions begin at FNOL.
By Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly
arXiv:2609.01188v1 Announce Type: new
Abstract: Large Language Models (LLMs) are revolutionizing digital communication by powering conversational agents deployed across domains such as customer servi...
By Rohan Kirti, Akash Ghosh, Aryan Vats, Niladri Ghosh, Shipra Shriparn, Roshni Ramnani, Anutosh Maitra, Sriparna Saha
arXiv:2607. 23813v1 Announce Type: cross Abstract: We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions.
By Denglin Jiang, Haoran Zhou, Anshul Wadhawan, Brendan Fahy, Vinay Ramesh, David Weisberg, Dmitriy Derkachevskiy, Helen Sheehan, Srivas Prasad, Michele Franceschini
TeleAntiFraud 2.0 is a monthly‑frozen, audio‑based benchmark for telecom fraud detection that incorporates newly observed scam patterns while preserving earlier test sets. It uses a Mixed‑Tree Anti‑Fraud Generation Pipeline to create profile‑grounded scenarios, expands them into mixed‑tree dialogues, and renders validated speech for 900 Chinese calls (600 fraud, 300 near‑domain non‑fraud) each month. Experiments show that classifiers perform well against unrelated negatives but drop significantly against near‑domain negatives, highlighting the need for near‑domain construction and collapse‑aware reporting in realistic evaluation settings.
By Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen
The paper introduces VeriSpeak, a benchmark of 3,879 spoken claims for evaluating fact verification in Large Audio Language Models (LALMs). It shows a clear modality gap: models that verify written claims well often fail on spoken versions, and retrieval alone offers limited improvement. Combining retrieval with explicit reasoning yields the best performance, reaching 86.1% accuracy and demonstrating the need for grounded reasoning over retrieved evidence in speech misinformation detection.
By Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri