arXiv AI By Suhyun Kim, Jinmo Han, Danny Dongyeop Han, Ahhyun Lucy Lee, Jewoon Lee, Yonghyeon Gwon, Zach Paris, Chun Kee Chung, Saewoong Bahk, Nam Soo Kim, Seong Jae Hwang, Jiook Cha

BAT-CLIP: Trimodal Alignment of Brain, Audio and Text

Read the original on arXiv AI →

BAT-CLIP is a trimodal alignment framework that jointly aligns intracranial EEG (iEEG) neural embeddings to both pretrained audio and text anchors within a shared, frozen audio‑text manifold. Unlike existing CLIP‑style brain‑speech models that anchor neural activity to a single modality, BAT‑CLIP leverages both audio and text to preserve temporal structure and linguistic separability. On the naturalistic Podcast benchmark, BAT‑CLIP produces more robust representations than bimodal CLIP baselines and demonstrates the value of self‑supervised foundation models for CLIP training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
5d ago

Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

Cephalonauts One is a 3 Tesla fMRI dataset featuring 30 hours of brain activity recorded from three healthy subjects while they listened to native-language audio podcasts. The release includes the raw fMRI data, corresponding podcast audio, transcript annotations, and derived stimulus embeddings, making it the deepest naturalistic speech fMRI dataset available. A brain‑decoding benchmark is introduced, framing audio segment retrieval as a task where decoders must match fMRI activity to the correct time‑aligned podcast segment, with standardized splits, metrics, and baseline models provided.

By Antoine Collas, Louis Jalouzot, G\'eraud Ilinca, Corentin Caris, Romain Valabr\`egue, Ahmed Hassayoune, David Goncalves, Madeleine Hueber, Thadd\'ee Delebarre, Julien Savatovsky, Clara Fonteneau, Charles Maussion, Bertrand Thirion, Alexis Thual
arXiv Computation and Language
Sep 4

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

The paper introduces Alignment-Free Text‑Audiobox (Text‑AB), a unified diffusion‑based framework that performs high‑quality voice dubbing and full‑duplex dialogue synthesis without requiring forced alignment. Text‑AB uses a latent diffusion model with DAC‑VAE features, achieving over 10× compression compared to prior EnCodec representations, and learns text‑speech alignment via cross‑attention. The authors pretrain a 3B‑parameter model on 480k hours of monolingual speech and fine‑tune it for cross‑lingual dubbing, full‑duplex dialogue, and emotional dialogue, reporting significant improvements in prosody, voice similarity, naturalness, and emotional expressivity over existing internal systems.

By Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu