arXiv AI

BAT-CLIP: Trimodal Alignment of Brain, Audio and Text

BAT-CLIP is a trimodal alignment framework that jointly aligns intracranial EEG (iEEG) neural embeddings to both pretrained audio and text anchors within a shared, frozen audio‑text manifold. Unlike existing CLIP‑style brain‑speech models that anchor neural activity to a single modality, BAT‑CLIP leverages both audio and text to preserve temporal structure and linguistic separability. On the naturalistic Podcast benchmark, BAT‑CLIP produces more robust representations than bimodal CLIP baselines and demonstrates the value of self‑supervised foundation models for CLIP training.

arXiv AI
5d ago

Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

Cephalonauts One is a 3 Tesla fMRI dataset featuring 30 hours of brain activity recorded from three healthy subjects while they listened to native-language audio podcasts. The release includes the raw fMRI data, corresponding podcast audio, transcript annotations, and derived stimulus embeddings, making it the deepest naturalistic speech fMRI dataset available. A brain‑decoding benchmark is introduced, framing audio segment retrieval as a task where decoders must match fMRI activity to the correct time‑aligned podcast segment, with standardized splits, metrics, and baseline models provided.

By Antoine Collas, Louis Jalouzot, G\'eraud Ilinca, Corentin Caris, Romain Valabr\`egue, Ahmed Hassayoune, David Goncalves, Madeleine Hueber, Thadd\'ee Delebarre, Julien Savatovsky, Clara Fonteneau, Charles Maussion, Bertrand Thirion, Alexis Thual
arXiv Computation and Language
Sep 4

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

The paper introduces Alignment-Free Text‑Audiobox (Text‑AB), a unified diffusion‑based framework that performs high‑quality voice dubbing and full‑duplex dialogue synthesis without requiring forced alignment. Text‑AB uses a latent diffusion model with DAC‑VAE features, achieving over 10× compression compared to prior EnCodec representations, and learns text‑speech alignment via cross‑attention. The authors pretrain a 3B‑parameter model on 480k hours of monolingual speech and fine‑tune it for cross‑lingual dubbing, full‑duplex dialogue, and emotional dialogue, reporting significant improvements in prosody, voice similarity, naturalness, and emotional expressivity over existing internal systems.

By Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu
arXiv Computer Vision
Oct 2

Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

The paper introduces Align Then Reason (ATR), a multilingual lip‑sync judge that aligns frame‑level lip representations with phonetic units of a candidate text line and then uses a language model to evaluate both content and timing. ATR achieves significant improvements over existing baselines on a seven‑language benchmark, with mean AUC gains of up to 59.4% for 2B reasoners and similar gains across other LLM families. The method also transfers well to unseen languages and outperforms lip‑reading baselines on real dubbing tasks such as dub‑line reranking and script‑to‑clip assignment.

By Rui Liu, Bhavin Jawade, Haoqi Li, Shivam Mehta, Karan Saxena, Yinghong Lan, Cameron R. Wolfe
arXiv AI
Jul 24

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

arXiv:2603. 01006v3 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth.

By Pengfei Zhang, Tianxin Xie, Minghao Yang, Li Liu
arXiv AI
Jul 7

StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model

arXiv:2604. 19635v2 Announce Type: replace-cross Abstract: While generative models have set new benchmarks for Target Speaker Extraction (TSE), their inherent reliance on global context precludes deployment in real-time applications.

By Shuhai Peng, Hui Lu, Jinjiang Liu, Liyang Chen, Guiping Zhong, Jiakui Li, Huimeng Wang, Haiyun Li, Liang Cao, Shiyin Kang, Zhiyong Wu