arXiv AI

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

arXiv:2608. 06027v1 Announce Type: cross Abstract: In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write.

arXiv AI
Sep 18

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

MTVA-Bench is a new benchmark designed to evaluate the language model component of cascaded voice agents under realistic conditions. It simulates callers with an LLM, mocks backend tool responses, and scores both tool‑call correctness and conversational quality using two LLM judges, covering 49 agents, 490 scenarios, and 7 languages. The benchmark reveals that while models perform similarly on tool selection, they differ widely in argument handling, rule compliance, and dialogue quality, highlighting the nuanced challenges of real‑world voice interactions.

By Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath
arXiv Computation and Language
Sep 3

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

SpeakPay is a voice‑first digital wallet designed to make mobile payment apps in Nepal accessible to visually impaired users. The paper introduces NepFinSpeech‑403, a 403‑utterance Nepali financial voice command dataset, and demonstrates that fine‑tuning Whisper large‑v2 with LoRA reduces the Word Error Rate from 129.95% to 42.58% and improves Devanagari numeral recognition from 0.0% to 73.9%. Domain adaptation also boosts the Transaction Success Rate from 1.67% to 33.33%, with as few as 100 domain‑specific utterances halving the zero‑shot WER.

By Biraj Subedi
arXiv AI
6d ago

Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

The paper reviews the fragmented literature on real‑time voice agents, noting that architecture, turn‑taking, and agentic evaluation communities rarely cite each other. It presents three evidence‑based claims: (1) architecture choice is a deployment constraint rather than a definitive solution, (2) evaluation has shifted from component quality to grounded outcomes, and (3) the dyadic assumption is breaking down as multiparty turn‑taking and reasoning become essential. The authors propose the TRG reporting standard to characterize agents by timing, recovery, and state‑verified outcomes, with an optional fourth axis for multiparty contexts.

By Shivam Negi, Arpit Rawat, Rashi Jain
arXiv Computation and Language
Aug 28

Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit

Vagdhenu is a Sanskrit shloka‑to‑chant text‑to‑speech system that preserves meter (vrutta) and phonological nuances. It builds on an off‑the‑shelf flow‑matching backbone and a large‑scale neural vocoder, adding a Kannada‑based frontend to avoid schwa deletion, a phonology‑aware frontend handling visarga sandhi and sibilant distinctions, and a vrutta‑aware reference selection mechanism. The authors report that a text‑side prosody conditioner is ineffective in their architecture, while reference clips and voice‑steering retraining provide the necessary prosody control, and they demonstrate the system’s performance on a 32‑chapter video corpus and an audio app covering 18,000 verses. whyItMatters":"The system delivers high‑fidelity, meter‑aware Sanskrit chanting, enabling large‑scale deployment of authentic recitations for educational and cultural preservation purposes."

By Prathosh A P
arXiv AI
Sep 7

Scalable Context Orchestration for Serving LLMs Over Voice

The paper introduces llmovoice, a middleware that explicitly models voice context for large language model (LLM) serving in voice AI applications. By incorporating speaking rate, background noise, packet loss, and other paralinguistic factors into a bounded context, llmovoice guides the LLM to generate more aligned responses. Experiments show significant reductions in speaking‑rate errors, false interruptions, and model usage costs, especially in long voice sessions.

By Linyi Jiang, Silvery D. Fu, Yifei Zhu
arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
Hugging Face Trending Papers
Sep 3

Scalable Context Orchestration for Serving LLMs Over Voice

Scalable Context Orchestration for Serving LLMs Over Voice presents llmovoice, a middleware that explicitly models voice context—including speaking rate, background noise, and packet loss—to guide large language model responses. By constructing a bounded voice context at each turn, llmovoice improves alignment with user preferences and reduces errors, achieving a 52.4% drop in speaking‑rate alignment error and a 0.9% false‑interruption rate under packet loss. In addition, it cuts model usage costs dramatically, lowering per‑turn cost by up to 24.9× while maintaining 98.7% of baseline answer quality in long sessions.

arXiv AI
Aug 28

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

SpeechGym is an audio‑native environment that lets two omni‑modal models converse entirely in native audio, eliminating external ASR/TTS and API boundaries while preserving the tasks, tools, and success checks of a standard text‑based agent benchmark. By training end‑to‑end, the framework addresses perceptual failures—such as misheard arguments that cascade into failed calls—and behavioural failures, both of which are automatically labeled for free. Using per‑turn process rewards to overcome reward sparsity, agents trained in SpeechGym transfer to an independent voice benchmark, doubling task success and improving efficiency in turns and tokens.

By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren
arXiv AI
Sep 4

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

The paper presents a method for creating a compact fixed‑voice Thai text‑to‑speech system by training a student model on synthetic speech generated from a large voice‑cloning teacher. By using only a short 15‑second voice reference and carefully filtering synthetic data, the authors build an 82‑million‑parameter model, Wayu‑Paxa‑TTS‑Edge, that runs on device without reference audio. The system achieves strong performance—68.2 % challenge‑set keyword accuracy, 91.4 % pause precision, and low character error rates—while outperforming its teacher and approaching the quality of a larger Gemini 3.1 model.

By Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut