arXiv Computation and Language By Bhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D, Srihari Bandarupalli, Santosh Kesiraju, Anil Vuppala

V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

Read the original on arXiv Computation and Language →

VákQA is a new Telugu spoken factoid question answering benchmark comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini-as-a-judge best approximates human ratings but is unevenly strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity lost in translation, phonetic confusions from speech input, and cascading ASR‑MT errors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Sep 17

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

VākQA is a newly introduced benchmark for Telugu spoken factoid question answering, comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini‑as‑a‑judge best approximates human ratings but is inconsistently strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity loss in translation, phonetic confusions from speech input, and compounded errors from cascaded ASR‑MT pipelines.

arXiv Machine Learning
Aug 18

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).

By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv AI
Jun 26

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

arXiv:2606. 26901v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown.

By Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan, Astut Kurariya, Diptadhi Mukherjee, Prabhat Chand, Pratima Murthy, Koustav Rudra, Lekhansh Shukla, Animesh Mukherjee
arXiv AI
Sep 3

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

VakyArth is the first pragmatic benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam. It tests models on five pragmatic phenomena—deixis, speech acts, implicature, social pragmatics, and coherence—using multiple-choice questions, natural language inference, and translation tasks authored by native speakers. Evaluation of multilingual LLMs shows consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions, with systematic differences across languages and tasks.

By Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy L
arXiv Computation and Language
Sep 10

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

SWORD is a new benchmark that tests large language models’ ability to reject factually incorrect statements across eight major languages by distorting Wikidata triples. The benchmark reveals that models often perform better on semantically plausible distortions than on random ones, indicating a reliance on distributional familiarity rather than true factual verification. It also shows significant performance drops for East Asian languages, with gaps up to 28 percentage points, highlighting asymmetric multilingual factual reasoning capabilities.

By Sanghyeok Park, Minji Kang, Hosung Kwak, Jinhyuk Yun