VákQA is a new Telugu spoken factoid question answering benchmark comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini-as-a-judge best approximates human ratings but is unevenly strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity lost in translation, phonetic confusions from speech input, and cascading ASR‑MT errors.
By Bhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D, Srihari Bandarupalli, Santosh Kesiraju, Anil Vuppala
arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).
By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv:2606. 26901v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown.
By Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan, Astut Kurariya, Diptadhi Mukherjee, Prabhat Chand, Pratima Murthy, Koustav Rudra, Lekhansh Shukla, Animesh Mukherjee
arXiv:2305.12474v4 Announce Type: replace
Abstract: Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensiv...
By Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, Xipeng Qiu, Tianxiang Sun, Peng Li, Shiqiao Meng, Yanjun Zheng, Jun Zhan, Zhangyue Yin, Xiannian Hu, Guofeng Quan, Qixiang Wang
arXiv:2504.11582v3 Announce Type: replace
Abstract: How can a monolingual English speaker determine whether an automatic translation in French is good enough to be shared? Existing MT error detection...
By Dayeon Ki, Kevin Duh, Marine Carpuat
arXiv:2609.21663v1 Announce Type: new
Abstract: Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equal...
By Hritika Sharma, Thibault Ba\~neras-Roux, Alessandra Pinto, Petr Motlicek, Hyunggu Jung, Esa\'u Villatoro-Tello, Somang Nam