Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces a benchmark for topic matching in real-world ASR transcripts from contact centers, where noisy, punctuation‑free speech data must be classified into predefined topics. It presents a human‑annotated dataset of topic‑utterance judgments and evaluates three matcher types—regex, zero‑shot sentence embeddings, and Gemini‑based LLMs—using two topic representations: keyphrases and natural language descriptions. Experiments show that lightweight LLM matchers outperform the other methods, especially when natural language descriptions are used.
arXiv:2606.22473v2 Announce Type: replace-cross Abstract: Speech language models (SLMs) increasingly combine speech and text, often by interleaving their tokens within a single sequence. Yet how thes...
The paper introduces Hybrid Search, a method that refines warm-initialized large language model (LLM) based automatic speech recognition (ASR) systems by exploiting interactions between ASR hidden states and the base LLM’s hidden states. By identifying tokens with high semantic dependence and selectively correcting them, the approach surpasses traditional global LLM‑correction techniques such as rescoring and late fusion. The study demonstrates that even after warm initialization, LLM‑based ASR models can further benefit from their base LLM during inference.
arXiv:2609.22214v1 Announce Type: new Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
arXiv:2606. 10029v1 Announce Type: cross Abstract: Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream.