arXiv AI

Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models

arXiv Computation and Language
3d ago

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

The paper introduces a symbiotic architecture that equips large language models with audio‑understanding abilities without fine‑tuning their weights. It uses an injector module to write audio‑conditioned vectors into the LLM’s key‑value cache, allowing the model to act as an audio language model while keeping the backbone unchanged. The approach improves scalability—since injection cost depends on the injector width—and preserves the LLM’s original text performance, outperforming conventional frozen‑LLM methods and approaching fine‑tuned ALM results on audio tasks.

By Yotaro Kubo, Qi Sun, Yujin Tang
arXiv Computation and Language
Sep 24

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

Mizar is a 159.3‑million‑parameter audio‑language model designed for devices with limited memory and computation. It couples a compact CED‑Small audio encoder with SmolLM2‑135M via a frequency‑merging mapper and is trained in three stages—audio‑language alignment, audio‑dependent fine‑tuning, and post‑training—to improve performance on audio‑question tasks. Across five random seeds, Mizar outperforms all other sub‑200M‑parameter ALMs on MMAU, MMAR, and ADQA‑clean, achieving mean accuracies of 52.92%, 42.42%, and 36.02% respectively, while enabling local inference on a single CPU with an average latency of 1.09 seconds for MMAU questions.

By Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji
arXiv Computation and Language
Sep 21

I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance

arXiv:2609.21183v1 Announce Type: cross Abstract: Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM...

By Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang, Seungwhan Moon, Shashank Jain, Pinar Donmez, Babak Damavandi
arXiv AI
Aug 10

MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning

arXiv:2601. 18904v3 Announce Type: replace-cross Abstract: Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data.

By Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson