VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
VoxReason introduces a listener‑free evaluation framework that measures whether a speech planning system’s delivery choices—such as pitch, energy, rate, pause, emphasis, and stance—are grounded in cited source records before any waveform is generated. The system outputs a source‑cited speaking plan and uses a deterministic verifier to check citation legality, slot agreement, unsupported states, schema validity, and counterfactual locality. Experiments on 1,440 source‑label cases show that simple slot accuracy can be misleading, while a 7B locality‑based repair model significantly improves plan‑slot accuracy and locality, and removing source records sharply reduces the grounded score. whyItMatters":"The framework provides a concrete, measurable way to ensure that expressive speech systems make source‑licensed planning decisions, addressing a source‑use failure that occurs before audio synthesis."
arXiv:2608.27783v3 Announce Type: replace-cross Abstract: Speech language models (speech LLMs) can generate plausible outputs from audio that contains no usable speech evidence. We study this failure...
The paper investigates how streaming emotion recognition models can be misled by their own prior predictions, a problem termed previous-belief contamination (PBC). Using a counterfactual diagnostic on CREMA-D-Stream, the authors show that feeding a model’s previous emotion label into its current prediction can drastically lower accuracy and flip many predictions, with the effect varying by label. To mitigate PBC, they propose EmoUpdate, a training‑free framework that isolates current audio perception from historical context through a prior‑blind firewall, a causal belief filter, and a decontamination operator, achieving significant gains across multiple SpeechLMs and benchmarks.
The paper investigates how to evaluate generative audio large language models (Audio‑LLMs) on known closed‑set tasks by separating the decision to call a generative model from the use of acoustic evidence. It introduces a controlled call‑decision framework where a policy can choose between a transcript label, encoder evidence from CLAP, AST, or WavLM, or a generative call to Qwen2‑Audio, Qwen2.5‑Omni, or MOSS‑Audio, and measures the impact of generative calls on accuracy. Results on the VocalSound dataset show that while transcript‑only accuracy is low (0.296), encoder‑based controls achieve high accuracy (≈0.85) without any generative calls, and adding generative calls yields only a marginal improvement (0.925 vs. 0.921).
The paper investigates how to evaluate audio‑language models by separating the use of acoustic evidence from the need to invoke a generative audio model. Using a controlled call‑decision framework, the authors compare policies that rely on transcript labels, encoder outputs from CLAP, AST, or WavLM, and optional calls to generative models such as Qwen2‑Audio, Qwen2.5‑Omni, or MOSS‑Audio. Results on the VocalSound dataset show that while transcript‑only accuracy is low (0.296), encoder‑only controls achieve high accuracy (≈0.85) without any generative calls, and adding generative calls yields only a marginal improvement (0.925 vs. 0.921).
The paper introduces the Speech-Unsupported Rejection Evaluation Challenge (SURE‑Challenge), a benchmark designed to test whether speech‑LLMs should accept or reject audio inputs before generating answers. Using LibriSpeech‑derived transcriptions paired with first‑word question answering, the authors evaluate various noise and silence conditions, and compare a simple energy‑plus‑Whisper‑score rule against a Qwen2‑Audio front‑end. On a 474‑row test set, the rule rejects 196 of 204 unsupported inputs while preserving accuracy on supported data, revealing a pre‑generation error mode that answer‑only scoring misses.