Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
arXiv:2506. 01503v2 Announce Type: replace Abstract: With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear.
arXiv:2602. 19778v4 Announce Type: replace-cross Abstract: Automatic Chord Recognition (ACR) is constrained by the scarcity of aligned chord labels, as well-aligned annotations are costly to acquire.
arXiv:2506. 01503v2 Announce Type: replace Abstract: With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear.
arXiv:2606. 07387v1 Announce Type: new Abstract: State-of-the-art text-to-music generation systems rely on massive proprietary datasets and industrial-scale compute, making it impossible to disentangle architectural contributions from resource advantages.
The paper introduces Reward‑Tilted On‑Policy Distillation (RT‑OPD), a method that enhances acoustic grounding in audio‑language models by using a frozen teacher to generate a reward based on the contrast between token predictions with and without audio. This reward reshapes the teacher distribution for reverse‑KL distillation, encouraging students to rely more on acoustic evidence. Experiments on two compact students across three benchmarks show that RT‑OPD consistently outperforms vanilla OPD, and a 3B model trained with RT‑OPD achieves 72.72% accuracy on the MMAU benchmark, surpassing other 3B models and rivaling larger 7B and 8B models.
arXiv:2601. 19919v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is one of the most effective paradigms for compressing large-scale foundation models into deployable architectures.
arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.
arXiv:2609.27389v1 Announce Type: cross Abstract: Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is c...
arXiv:2607. 06179v1 Announce Type: cross Abstract: There are some datasets of varying scales for audio classification (AC) applied to different tasks.
arXiv:2608.30974v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in la...
arXiv:2607. 05196v1 Announce Type: cross Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech.
arXiv:2508. 10409v3 Announce Type: replace-cross Abstract: This paper constructs a textual dataset for training large language models (LLMs) to learn analog circuit knowledge and customizes LLM training techniques.
arXiv:2512.07571v3 Announce Type: replace Abstract: This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tune...
The paper introduces a domain‑specific parameter‑isolation architecture for domain‑incremental learning (DIL) in audio classification, aiming to preserve knowledge from earlier domains without accessing their data. By employing data‑free generative replay and cross‑domain feature generation, the method constructs new experts conditioned on all previously frozen models, thereby mitigating catastrophic forgetting. Applied to the DCASE 2026 Challenge Task 7, the approach achieves micro and macro accuracies of 78.4 % and 78.9 %, outperforming the baseline by 33 and 25 percentage points, respectively, with ablation studies confirming the contribution of each component.