arXiv AI By Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji, Dinesh Manocha

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

Read the original on arXiv AI →

arXiv:2607. 28896v1 Announce Type: cross Abstract: Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 28

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

FireRedAudio is a 9‑billion‑parameter audio language model that separates continuous input representations for audio understanding and speech generation, enabling a single autoregressive LLM to perform tasks such as ASR, zero‑shot TTS, Instruct TTS, and semantic/acoustic speech editing. The model uses a dedicated Audio Encoder for recognition and a RedAE‑based pathway for generation, with the LLM directly generating text or conditioning a flow‑matching DiT to produce acoustic latents. Evaluations show competitive or leading performance in multilingual ASR, content‑accurate zero‑shot TTS, strong instruction following, and significant improvements in speech editing over prior work.

By Feiyu Shen, Fenglong Xie, Junjie Li, Kun Xie, Lei Xie, Xu Tang, Xuelong Geng, Yan Jia, Yao Hu, Yichen Han, Yichen Wu, Ziqi Dai, Junjie Chen, Kai Huang, Manzhen Wei, Yixuan Li
arXiv AI
Sep 10

Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation

arXiv:2609.08275v1 Announce Type: new Abstract: Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute e...

By Tianyi Zeng, Junchao Liao, Yujie Wei, Ziying Zhang, Litao Li, Tianyi Wang, Zhichao Wei, Shuyao Xu, Wenwen Qiang, Siyu Zhu, Zhenghao Zhang, Long Qin
arXiv AI
Sep 7

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

PRISM‑Bench is an audio‑centric diagnostic benchmark for text‑to‑audio‑video generation, built from 900 human‑verified samples. It evaluates audio along two axes—audio type (speech, music, sound) and sound‑source visibility (on‑screen vs. off‑screen)—across four perceptual dimensions (audio‑visual coherence, audio quality, audio expressiveness, and prompt following) using 35 fine‑grained criteria. The benchmark employs an enhanced MLLM‑as‑a‑Judge protocol that aligns strongly with human raters, revealing a performance gap between frontier and open‑source T2AV models and highlighting overfitting to perceptual fidelity while struggling with complex grounding and control tasks, especially for music and synchronized on‑screen audio.

By Yuchen Sun, Qian Yang, Jun Wang, Detai Xin, Guoqiao Yu, Guanglu Wan, Qi Jia
arXiv Machine Learning
1d ago

AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.

By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan