We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory. md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best.
SpeechGym is an audio‑native environment that lets two omni‑modal models converse entirely in native audio, eliminating external ASR/TTS and API boundaries while preserving the tasks, tools, and success checks of a standard text‑based agent benchmark. By training end‑to‑end, the framework addresses perceptual failures—such as misheard arguments that cascade into failed calls—and behavioural failures, both of which are automatically labeled for free. Using per‑turn process rewards to overcome reward sparsity, agents trained in SpeechGym transfer to an independent voice benchmark, doubling task success and improving efficiency in turns and tokens.
By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren
The paper presents an interpretable, fair, and accurately benchmarked automated system for assessing second‑language English speaking. Using a hybrid of feature‑based speech‑timing metrics and a large language model (LLM) fluency judgment, the system achieves a Spearman correlation of 0.818 with the ICNALE Global Rating Archive, outperforming 81 % of trained human raters. A controlled study shows that encoding pauses into the LLM prompt does not meaningfully affect fluency scores, indicating that the system’s fluency signal derives from measurable speech‑timing features.
By Eichi Uehara
The study investigates how two computational dimensions—model depth and refinement steps—affect intelligibility and speaker identity in masked-diffusion text‑to‑speech systems. Experiments with 15 models (19–133 M parameters) and up to 16 refinement steps show that refinement improves intelligibility more than identity, with a 1.86× asymmetry that persists even after retraining. Best‑of‑K search can recover identity when refinement fails, and analysis indicates that depth and steps target distinct bottlenecks, requiring separate optimization.
By Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh
arXiv:2609.38867v1 Announce Type: new
Abstract: Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular in...
By Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang
arXiv:2609.18533v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representation...
By Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis
The paper introduces ASCIL, a post‑ASR correction framework that re‑evaluates wake‑up intent by combining acoustic embeddings, linguistic cues, device context, and past misclassifications. ASCIL interprets both implicit (hesitation, disengagement, silence) and explicit (cancellation, repetition) signals as noisy indicators of misclassification, enabling online pattern updates without manual annotation. On a proprietary dataset of 3,667 interactions, ASCIL reduces errors by up to 54.27% relative on a session‑disjoint subset and 24.39% at a 0.90 threshold, while adding less than 60 ms of latency and improving intentional acceptance rates.
By Preeti Saraswat, Divya Neelagiri, Anil Yadav
arXiv:2608. 03970v1 Announce Type: new Abstract: Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools.
By Zizhao Hu, Nathan Elijah Segura, Mohammad Rostami, Jesse Thomason
arXiv:2606. 27472v1 Announce Type: cross Abstract: Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a plan is revised.
By Vedant Patel
GEPARD is a streaming text‑to‑speech model that uses a standard large language model backbone to generate speech autoregressively, decoding audio with an FSQ‑based neural codec. It streams audio chunk‑by‑chunk as text arrives, achieving a real‑time factor of about 0.067 and an aggregate speedup of roughly 204× on a single GPU with 256 concurrent streams. The design keeps all complex auxiliary mechanisms outside the decode loop, enabling deployment with a standard LLM engine (vLLM) without kernel modifications.
By Denis Pavlov, Ulanbek Abdurazakov, Nursultan Bakashov
The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.
By Nusrat Jahan Lia, Aritra Mazumder
arXiv:2608.28609v1 Announce Type: cross
Abstract: A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved...
By Bojie Li, Noah Shi