arXiv Machine Learning By Lee Seung-woo, Bowen Qi, Kim Min-jun, Jang Won-young

SkillFormer: Skill-Decomposed Adaptation for Audio Language Models

Read the original on arXiv Machine Learning →

SkillFormer is a method for audio language models that decomposes audio understanding into skill‑specific low‑rank adapters and uses a learned router to activate the appropriate adapters at inference time. The router selects which adapters to engage based on the question, allowing different parameters to be used for tasks such as pitch comparison versus genre classification. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, reducing gradient conflicts and adding fewer than 4% of the base model’s parameters. The approach improves average accuracy by 2.5 to 4.1 points across three distinct models on MMSU, MMAU‑Pro, and MMAR, achieving balanced gains across perception, reasoning, and semantic subcategories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Oct 1

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.

By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu
arXiv Machine Learning
1d ago

AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models

AdaLoop is a lightweight recurrent module that adaptively determines how many latent refinement steps are needed for audio–question pairs, allowing deeper reasoning only when necessary. It shares a transformer block that iterates over the audio representation guided by the question, with a learned halting mechanism that exits the loop once the representation is ready. Adding fewer than 3 % of the base model’s parameters, AdaLoop improves average accuracy by 2.9 to 3.8 points across three distinct models, especially on perception-heavy subtasks.

By Lee Seung-woo, Bowen Qi
arXiv Computation and Language
Sep 28

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

The paper introduces a symbiotic architecture that equips large language models with audio‑understanding abilities without fine‑tuning their weights. It uses an injector module to write audio‑conditioned vectors into the LLM’s key‑value cache, allowing the model to act as an audio language model while keeping the backbone unchanged. The approach improves scalability—since injection cost depends on the injector width—and preserves the LLM’s original text performance, outperforming conventional frozen‑LLM methods and approaching fine‑tuned ALM results on audio tasks.

By Yotaro Kubo, Qi Sun, Yujin Tang