arXiv AI By Aman Dalmia, Sanskriti Midha, Jigar Doshi

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

Read the original on arXiv AI →

arXiv:2608. 06027v1 Announce Type: cross Abstract: In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

MTVA-Bench is a new benchmark designed to evaluate the language model component of cascaded voice agents under realistic conditions. It simulates callers with an LLM, mocks backend tool responses, and scores both tool‑call correctness and conversational quality using two LLM judges, covering 49 agents, 490 scenarios, and 7 languages. The benchmark reveals that while models perform similarly on tool selection, they differ widely in argument handling, rule compliance, and dialogue quality, highlighting the nuanced challenges of real‑world voice interactions.

By Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath
arXiv Computation and Language
Sep 3

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

SpeakPay is a voice‑first digital wallet designed to make mobile payment apps in Nepal accessible to visually impaired users. The paper introduces NepFinSpeech‑403, a 403‑utterance Nepali financial voice command dataset, and demonstrates that fine‑tuning Whisper large‑v2 with LoRA reduces the Word Error Rate from 129.95% to 42.58% and improves Devanagari numeral recognition from 0.0% to 73.9%. Domain adaptation also boosts the Transaction Success Rate from 1.67% to 33.33%, with as few as 100 domain‑specific utterances halving the zero‑shot WER.

By Biraj Subedi
arXiv AI
6d ago

Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

The paper reviews the fragmented literature on real‑time voice agents, noting that architecture, turn‑taking, and agentic evaluation communities rarely cite each other. It presents three evidence‑based claims: (1) architecture choice is a deployment constraint rather than a definitive solution, (2) evaluation has shifted from component quality to grounded outcomes, and (3) the dyadic assumption is breaking down as multiparty turn‑taking and reasoning become essential. The authors propose the TRG reporting standard to characterize agents by timing, recovery, and state‑verified outcomes, with an optional fourth axis for multiparty contexts.

By Shivam Negi, Arpit Rawat, Rashi Jain
arXiv Computation and Language
Aug 28

Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit

Vagdhenu is a Sanskrit shloka‑to‑chant text‑to‑speech system that preserves meter (vrutta) and phonological nuances. It builds on an off‑the‑shelf flow‑matching backbone and a large‑scale neural vocoder, adding a Kannada‑based frontend to avoid schwa deletion, a phonology‑aware frontend handling visarga sandhi and sibilant distinctions, and a vrutta‑aware reference selection mechanism. The authors report that a text‑side prosody conditioner is ineffective in their architecture, while reference clips and voice‑steering retraining provide the necessary prosody control, and they demonstrate the system’s performance on a 32‑chapter video corpus and an audio app covering 18,000 verses. whyItMatters":"The system delivers high‑fidelity, meter‑aware Sanskrit chanting, enabling large‑scale deployment of authentic recitations for educational and cultural preservation purposes."

By Prathosh A P
arXiv AI
Sep 7

Scalable Context Orchestration for Serving LLMs Over Voice

The paper introduces llmovoice, a middleware that explicitly models voice context for large language model (LLM) serving in voice AI applications. By incorporating speaking rate, background noise, packet loss, and other paralinguistic factors into a bounded context, llmovoice guides the LLM to generate more aligned responses. Experiments show significant reductions in speaking‑rate errors, false interruptions, and model usage costs, especially in long voice sessions.

By Linyi Jiang, Silvery D. Fu, Yifei Zhu
arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji