arXiv AI

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

arXiv:2606. 26901v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown.

arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
arXiv Computation and Language
Sep 22

Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations

arXiv:2609.24199v1 Announce Type: new Abstract: Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlle...

By Kaushal Santosh Bhogale, Srija Anand, Sadakopa Ramakrishnan Thothathiri, Tahir Javed, Sshubam Verma, Mitesh M. Khapra
Hugging Face Trending Papers
Sep 17

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

VākQA is a newly introduced benchmark for Telugu spoken factoid question answering, comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini‑as‑a‑judge best approximates human ratings but is inconsistently strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity loss in translation, phonetic confusions from speech input, and compounded errors from cascaded ASR‑MT pipelines.

arXiv AI
Aug 20

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

The paper titled "Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs" highlights that current safety alignment training for large language models is predominantly English-centric, leading to failures in non‑English languages. It introduces INCLUDE, a multilingual benchmark with 2,604 prompts in six languages (English, Hindi, Bengali, Marathi, Tamil, and Hinglish) to measure Indian‑centric socio‑cultural biases. Evaluation of ten open‑ and closed‑source LLMs shows that Bengali models exhibit the highest bias scores among open‑source models, while English shows the lowest bias in open‑source but the highest in closed‑source models.

By Namya Bhatnagar
arXiv Computation and Language
Sep 18

V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

VákQA is a new Telugu spoken factoid question answering benchmark comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini-as-a-judge best approximates human ratings but is unevenly strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity lost in translation, phonetic confusions from speech input, and cascading ASR‑MT errors.

By Bhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D, Srihari Bandarupalli, Santosh Kesiraju, Anil Vuppala
arXiv AI
Sep 25

Benchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages

This study evaluates automatic speech recognition (ASR) for adolescent health communication in Twi, Dagbani, and Ewe by benchmarking five ASR systems on a Bible corpus and a domain-specific ASRH dataset, then performing supervised domain adaptation with a fine‑tuned Qwen3-ASR-0.6B model. Fine‑tuning significantly lowered word and character error rates, especially for Ewe, and the adapted model was deployed in the KasaHealth voice‑first application, which received high user approval and highlighted remaining domain gaps. The work demonstrates that in‑domain data, rather than model size or computational resources, is the primary limitation for effective ASR in these languages.

By Stephen E. Moore, Akwasi Asare, Mich-Seth Owusu, Paul Azunre, Joel Budu, Lawrence A. Adu-Gyamfi
arXiv Computation and Language
Sep 10

False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UK

arXiv:2602.13047v2 Announce Type: replace Abstract: Conversational speech reveals early signs of cognitive decline, including dementia and mild cognitive impairment (MCI). AI models show promise for...

By Madhurananda Pahar, Caitlin Illingworth, Dorota Braun, Bahman Mirheidari, Lise Sproson, Daniel Blackburn, Heidi Christensen