arXiv Machine Learning By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan

AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

Read the original on arXiv Machine Learning →

AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

Tracing Audio Grounding and Answer Selection in Audio LLMs

The paper investigates how Audio Large Language Models (Audio LLMs) actually use audio input to determine answers, rather than relying on textual cues. It finds that replacing audio with silence or unrelated audio degrades performance more after training than before, that acoustic information shapes representations in early-to-middle layers and influences final predictions in middle-to-late layers, and that training impacts specific layer bands most strongly. These observations offer a mechanistic view of how training enhances the use of acoustic evidence in Audio LLMs.

By Hyebin Cho, Suho Yoo, Jihoo Jung, Joon Son Chung
arXiv AI
Aug 18

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.

By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov