DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.
arXiv:2602. 22431v2 Announce Type: replace-cross Abstract: Millimeter-wave (mmWave) radar captures are band-limited and noisy, making for difficult reconstruction of intelligible full-bandwidth speech.
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.
arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.
arXiv:2512. 15313v2 Announce Type: replace-cross Abstract: Deep learning has become a standard approach for the modeling of audio effects, yet strictly black-box modeling remains problematic for time-varying systems.
arXiv:2603. 09234v2 Announce Type: cross Abstract: Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE).
arXiv:2606. 05678v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription.
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM.
arXiv:2511. 11686v4 Announce Type: replace Abstract: Speech enhancement (SE) requires high-fidelity reconstruction of clean speech that preserves linguistic and paralinguistic cues while maintaining high perceptual quality.
arXiv:2512. 20978v2 Announce Type: replace-cross Abstract: Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech.
arXiv:2607. 05196v1 Announce Type: cross Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech.
arXiv:2607. 09134v1 Announce Type: cross Abstract: Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity.