Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.
arXiv:2609.23416v1 Announce Type: cross Abstract: Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape di...
arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...
The Eloquence team presents three methods for the Interspeech 2026 MLC‑SLM Task 2, a multilingual MCQA challenge covering 21 languages. They fine‑tune Voxtral‑Mini‑3B with LoRA and data augmentation, achieving 0.72 macro‑accuracy; they use multimodal in‑context learning on Voxtral‑24B to correct label bias, reaching 0.81; and they deploy a training‑free retrieval system with a voice‑anchored memory, scoring 0.68. All approaches surpass the official baseline.
arXiv:2609.22214v1 Announce Type: new Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...