DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.
arXiv:2602. 22431v2 Announce Type: replace-cross Abstract: Millimeter-wave (mmWave) radar captures are band-limited and noisy, making for difficult reconstruction of intelligible full-bandwidth speech.
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.
arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.
arXiv:2512. 15313v2 Announce Type: replace-cross Abstract: Deep learning has become a standard approach for the modeling of audio effects, yet strictly black-box modeling remains problematic for time-varying systems.
arXiv:2603. 09234v2 Announce Type: cross Abstract: Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE).
arXiv:2606. 05678v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription.
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
GrainSpeech is a compact speech synthesis model that uses a fixed‑receptive‑field convolutional encoder to reduce pitch, energy, and duration prediction errors by 36.0%, 17.3%, and 3.4% respectively. It introduces a Mel‑specific gradient‑variance supervision that improves fine‑scale variation while avoiding quality degradation. With only 264.8K parameters, GrainSpeech achieves 17.9× real‑time Mel generation on a microcontroller and attains UTMOS scores comparable to much larger models, using less than 1.5% of their parameters.
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM.
Generative models have recently demonstrated considerable promise in speech super-resolution (SSR). Nevertheless, the majority of existing work has concentrated on standard or versatile SSR configurat...
MeanVoiceFlow2 is a new voice conversion framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. It is trained via conversion distillation from MeanVoiceFlow and real data reconstruction, and further enhanced with diffusion-GAN training, sample mixing, and teacher-guided conditioning augmentation. Experiments on zero-shot voice conversion show that MeanVoiceFlow2 delivers higher perceptual quality and about nine times faster inference than its predecessor while preserving speaker similarity.
arXiv:2609.24138v1 Announce Type: cross Abstract: Generative models have recently demonstrated considerable promise in speech super-resolution (SSR). Nevertheless, the majority of existing work has c...