arXiv:2603. 15685v2 Announce Type: replace-cross Abstract: Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive.
By Bingzhou Li, Tao Huang
Cephalonauts One is a 3 Tesla fMRI dataset featuring 30 hours of brain activity recorded from three healthy subjects while they listened to native-language audio podcasts. The release includes the raw fMRI data, corresponding podcast audio, transcript annotations, and derived stimulus embeddings, making it the deepest naturalistic speech fMRI dataset available. A brain‑decoding benchmark is introduced, framing audio segment retrieval as a task where decoders must match fMRI activity to the correct time‑aligned podcast segment, with standardized splits, metrics, and baseline models provided.
By Antoine Collas, Louis Jalouzot, G\'eraud Ilinca, Corentin Caris, Romain Valabr\`egue, Ahmed Hassayoune, David Goncalves, Madeleine Hueber, Thadd\'ee Delebarre, Julien Savatovsky, Clara Fonteneau, Charles Maussion, Bertrand Thirion, Alexis Thual
arXiv:2512. 10120v2 Announce Type: replace-cross Abstract: General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting.
By Maris Basha, Anja Zai, Sabine Stoll, Richard Hahnloser
arXiv:2606. 05173v1 Announce Type: cross Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are strongly anchored to surface-form token identity rather than deeper semantic structure.
By Aimen Boukhari
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
By Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin
The paper introduces Alignment-Free Text‑Audiobox (Text‑AB), a unified diffusion‑based framework that performs high‑quality voice dubbing and full‑duplex dialogue synthesis without requiring forced alignment. Text‑AB uses a latent diffusion model with DAC‑VAE features, achieving over 10× compression compared to prior EnCodec representations, and learns text‑speech alignment via cross‑attention. The authors pretrain a 3B‑parameter model on 480k hours of monolingual speech and fine‑tune it for cross‑lingual dubbing, full‑duplex dialogue, and emotional dialogue, reporting significant improvements in prosody, voice similarity, naturalness, and emotional expressivity over existing internal systems.
By Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu
arXiv:2610.03125v1 Announce Type: cross
Abstract: Speech delivery varies with both the requested paralinguistic attribute and the linguistic content. We introduce ParaGeo, a matched-content decomposi...
By Yuhan Liu, Yuxuan Ou, Ruoxi Su, Mohamed Ahmed Zaki, Yunbo Long
The paper introduces Align Then Reason (ATR), a multilingual lip‑sync judge that aligns frame‑level lip representations with phonetic units of a candidate text line and then uses a language model to evaluate both content and timing. ATR achieves significant improvements over existing baselines on a seven‑language benchmark, with mean AUC gains of up to 59.4% for 2B reasoners and similar gains across other LLM families. The method also transfers well to unseen languages and outperforms lip‑reading baselines on real dubbing tasks such as dub‑line reranking and script‑to‑clip assignment.
By Rui Liu, Bhavin Jawade, Haoqi Li, Shivam Mehta, Karan Saxena, Yinghong Lan, Cameron R. Wolfe
arXiv:2605.06582v4 Announce Type: replace
Abstract: Modern learning systems represent perceptual signals with continuous vectors, but comparison, retrieval, memory, alignment, and reasoning are often...
By Adhiraj Banerjee, Vipul Arora
arXiv:2605. 06582v2 Announce Type: replace Abstract: Many operations on sensory data -- comparison, memory, retrieval, and reasoning -- are naturally expressed over discrete symbolic structures.
By Adhiraj Banerjee, Vipul Arora
arXiv:2603. 01006v3 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth.
By Pengfei Zhang, Tianxin Xie, Minghao Yang, Li Liu
arXiv:2604. 19635v2 Announce Type: replace-cross Abstract: While generative models have set new benchmarks for Target Speaker Extraction (TSE), their inherent reliance on global context precludes deployment in real-time applications.
By Shuhai Peng, Hui Lu, Jinjiang Liu, Liyang Chen, Guiping Zhong, Jiakui Li, Huimeng Wang, Haiyun Li, Liang Cao, Shiyin Kang, Zhiyong Wu