Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces a symbiotic architecture that equips large language models with audio‑understanding abilities without fine‑tuning their weights. It uses an injector module to write audio‑conditioned vectors into the LLM’s key‑value cache, allowing the model to act as an audio language model while keeping the backbone unchanged. The approach improves scalability—since injection cost depends on the injector width—and preserves the LLM’s original text performance, outperforming conventional frozen‑LLM methods and approaching fine‑tuned ALM results on audio tasks.
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
arXiv:2510.12851v2 Announce Type: replace-cross Abstract: Large Audio-Language Models (LALMs) excel in Audio QA but often suffer from hallucinations ungrounded in the audio. To our knowledge, we are...
arXiv:2510. 20441v2 Announce Type: replace-cross Abstract: Neural audio codecs have largely promoted the application of language models (LMs) for speech applications.
Mizar is a 159.3‑million‑parameter audio‑language model designed for devices with limited memory and computation. It couples a compact CED‑Small audio encoder with SmolLM2‑135M via a frequency‑merging mapper and is trained in three stages—audio‑language alignment, audio‑dependent fine‑tuning, and post‑training—to improve performance on audio‑question tasks. Across five random seeds, Mizar outperforms all other sub‑200M‑parameter ALMs on MMAU, MMAR, and ADQA‑clean, achieving mean accuracies of 52.92%, 42.42%, and 36.02% respectively, while enabling local inference on a single CPU with an average latency of 1.09 seconds for MMAU questions.
arXiv:2609.21183v1 Announce Type: cross Abstract: Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM...