arXiv AI

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

arXiv:2606. 10029v1 Announce Type: cross Abstract: Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream.

arXiv AI
Jun 9

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

arXiv:2512. 20978v2 Announce Type: replace-cross Abstract: Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech.

By Haoyang Li, Xuyi Zhuang, Azmat Adnan, Ye Ni, Wei Rao, Shreyas Gopal, Eng Siong Chng, Boon Siew Han, Yuanjin Zheng
arXiv AI
Sep 2

VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.

By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
arXiv AI
Sep 18

Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

The paper introduces a discriminative adaptation for SpeechLLMs that reads the hidden state of the final prompt token via a simple classification head, enabling emotion recognition in a single forward pass without altering the backbone. This approach replaces the generative decoder, which can produce out‑of‑set labels and favor frequent classes, with a controlled comparison between generative and discriminative inference. Experiments on IEMOCAP show improved Macro F1 scores, elimination of hallucinations, and larger gains on realistic ASR transcripts, while revealing that emotion directions encode indirect associations reflecting web‑scale text biases.

By Hasindri Watawana, Sergio Burdisso, Esa\'u Villatoro-Tello, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
Hugging Face Trending Papers
Aug 13

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale.