While text-based hallucination detection has been extensively studied, spoken hallucination detection remains largely unexplored, particularly for low-resource languages. We present the first multilin...
TRILOGUE is a new trilingual benchmark for spoken dialogue fact‑checking, covering English, Russian, and Kazakh. It includes almost 12,000 dialogues, 187,000 turns, and 390 hours of paired audio with ASR transcripts and word‑level timestamps, as well as nearly 5,000 human‑recorded Russian and Kazakh files. The dataset supports tasks such as claim check‑worthiness detection, evidence retrieval, and claim verification under various input conditions, and baseline experiments reveal challenges with ASR errors and cross‑lingual transfer, especially for Kazakh.
By Chaewan Chun, Meruyert Aristombayeva, Jiyoung Choi, Mahjabin Nahar, Delvin Ce Zhang, Dongwon Lee
The paper investigates why automatic speech recognition (ASR) systems sometimes generate fluent text that does not correspond to the input audio, a phenomenon termed hallucination. By examining two independently trained Conformer‑Large models—one using CTC and the other RNN‑T—under conditions of environmental noise and speaker‑background shift, the authors identify the final encoder stage as a critical boundary. Bypassing this final block leads to divergence on almost all utterances, while bypassing earlier blocks has minimal effect; at this stage, representations become compact, the decoder can read the text, and grapheme information becomes explicit, yet the output is garbled or repetitive rather than fluent fabrication. The study thus pinpoints a mechanistic precondition for hallucination—failure to produce adequately grounded output—though it does not fully explain naturally occurring hallucinations, and it highlights a consistent terminal‑stage dependency across decoder families and distribution shifts.
By Hamees Sayed, Apoorv Singh, Kumar Aman, Akshat Mandloi
VietPrism is a newly released, large‑scale Vietnamese speech corpus that combines 993.4 hours of real utterances from 1,262 verified speakers with 3.1 k hours of synthetic spoof speech. It uniquely offers transcripts, consistent speaker identities, five dialect groups, and extensive Vietnamese‑English code‑switching—nearly half of the corpus—while pairing each spoof with a matched bona fide utterance. The dataset enables controlled evaluation of deep‑fake detection models, revealing significant variability in detector performance across dialects and speaker similarity.
By Minh Hoang, Thai Le
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.
arXiv:2607. 21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks.
By Daniyal Kabir Dar, Arun Ross
The paper introduces a training‑free, inference‑time technique to curb hallucinated transcripts in Whisper by projecting decoder activations away from a low‑rank hallucination subspace derived from non‑speech data. Two variants are tested: an always‑on projection that dramatically lowers hallucination rates on non‑speech benchmarks, and a gated version that applies the projection only when non‑speech is predicted, achieving a smaller but still significant reduction. On LibriSpeech, the gated method slightly increases word error rate but keeps false‑rejection rates low, demonstrating a controllable trade‑off between hallucination suppression and recognition accuracy.
By Maryam Abbasihafshejani, Murtuza Jadliwala
The paper details a multi‑month engineering effort to build Sophea, a production Greek‑English automatic speech recognition system. It evaluates the system against nine production gates—including word error rate, language identification, and hallucinations—across twenty‑three training iterations and two model architectures, finding that no single training‑data composition satisfies all gates. The authors describe a six‑stage data pipeline that dramatically reduces discarded Greek audio, isolate a hallucination defect to a specific training package, and demonstrate that a three‑model ROVER ensemble achieves full gate coverage and a 29 % relative WER improvement. A separate learned arbiter model is released on the Open ASR Leaderboard with 4.26 % average WER on English tests and 25.88 % on live Greek noisy traffic.
By Christos Petrocheilos, Cleopatra Papadopoulou, Chris Porikis, Ioakeim Perros, Ayoub Kirouane, Themistoklis Nikolis
The paper investigates how encoder-decoder models in ASR and NMT can generate fluent text even when the input contains no recoverable message, a phenomenon known as message-free hallucination. By auditing the models’ reserved null tokens and manipulating their scores, the authors show that a higher null-token score can suppress fabrication but may also delete valid content or shorten translations. The study highlights that the null token can serve as a diagnostic tool for hallucination and suggests evaluating abstention methods by considering both suppression and deletion costs.
By Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev
arXiv:2609.18533v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representation...
By Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis
The paper introduces CAFNet, a lightweight cross‑attentive neural network that fuses MFCC, LFCC, and Chroma‑STFT features to detect and localise partially manipulated (half‑truth) speech. CAFNet achieves high ternary accuracy (97.55%) and low boundary mean absolute error (0.037 s) on the MLADDC benchmark, while demonstrating that cross‑corpus transfer depends on both capability and corpus characteristics. Ablation studies show that cross‑attention fusion is the most critical component, and removing a deeply supervised auxiliary head improves in‑domain performance and reduces variance.
By S. Sutharya, Remya K. Sasi
arXiv:2608. 00722v1 Announce Type: cross Abstract: Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text.
By Chenlin Liu, Minghui Fang, Zhonghao Bi, Zekai Su, Rong Wang, Jiqing Han