arXiv AI By Zizhao Hu, Nathan Elijah Segura, Mohammad Rostami, Jesse Thomason

Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations

Read the original on arXiv AI →

arXiv:2608. 03970v1 Announce Type: new Abstract: Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 31

Voice Memory for Agentic Speech Recognition

arXiv:2607. 26410v1 Announce Type: cross Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.

By Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg
arXiv AI
Aug 28

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

SpeechGym is an audio‑native environment that lets two omni‑modal models converse entirely in native audio, eliminating external ASR/TTS and API boundaries while preserving the tasks, tools, and success checks of a standard text‑based agent benchmark. By training end‑to‑end, the framework addresses perceptual failures—such as misheard arguments that cascade into failed calls—and behavioural failures, both of which are automatically labeled for free. Using per‑turn process rewards to overcome reward sparsity, agents trained in SpeechGym transfer to an independent voice benchmark, doubling task success and improving efficiency in turns and tokens.

By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren
arXiv AI
2d ago

What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

The paper introduces a Mixture-of-Agents (MoA) approach to quantify the per-token compute required by large language models. By having a panel of fifteen models of varying sizes attempt to reproduce each token, the authors define the smallest successful agent’s inference cost as the token’s sufficient compute, providing an upper bound on necessary computation. Experiments on benchmarks show that a 0.5B model can reproduce most tokens, and that the MoA-derived compute map can reduce latency in model routing and drafting tasks while improving or maintaining accuracy.

By Zhixu Du, Weijia Han, Hai Helen Li, Yiran Chen