arXiv Computation and Language

Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond

The article surveys brain‑to‑language decoding, covering tasks, neural signals, methods, and evaluation across invasive and non‑invasive modalities. It traces the field’s evolution from constrained recognition to text generation, streaming speech, and facial animation, linking tasks to neural populations and decoder representations. The review highlights complementary decoding targets, shared representations, and the growing importance of calibration, feedback, and user control for online communication, while proposing a five‑level future trajectory toward bidirectional cognitive exchange.

arXiv Computation and Language
1d ago

From Neurons to Conversation: Speech Brain-Computer Interfaces

arXiv:2609.36736v1 Announce Type: cross Abstract: Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative in...

By Moein Khajehnejad, Forough Habibollahi, Tommaso Boccato, Margarida Sousa, Michal Olak, Francesco Jamal Sheiban, Matteo Ferrante
arXiv AI
Jun 9

Liberating LLM Capabilities in Full-Duplex Speech Models

arXiv:2606. 07547v1 Announce Type: cross Abstract: Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs.

By Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu, Hanyu Liu, Yuan Yao
arXiv Computation and Language
Sep 4

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

The paper introduces Alignment-Free Text‑Audiobox (Text‑AB), a unified diffusion‑based framework that performs high‑quality voice dubbing and full‑duplex dialogue synthesis without requiring forced alignment. Text‑AB uses a latent diffusion model with DAC‑VAE features, achieving over 10× compression compared to prior EnCodec representations, and learns text‑speech alignment via cross‑attention. The authors pretrain a 3B‑parameter model on 480k hours of monolingual speech and fine‑tune it for cross‑lingual dubbing, full‑duplex dialogue, and emotional dialogue, reporting significant improvements in prosody, voice similarity, naturalness, and emotional expressivity over existing internal systems.

By Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu
arXiv Machine Learning
Aug 21

Decoding silent reading from non-invasive EEG

arXiv:2608. 20186v1 Announce Type: new Abstract: Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable.

By Ingo Marquardt, Anthilia Alchanat, Priyanka Jain
arXiv Computation and Language
Sep 22

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...

By Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue
arXiv Computer Vision
Aug 28

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.

By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
arXiv Computation and Language
Aug 28

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

FireRedAudio is a 9‑billion‑parameter audio language model that separates continuous input representations for audio understanding and speech generation, enabling a single autoregressive LLM to perform tasks such as ASR, zero‑shot TTS, Instruct TTS, and semantic/acoustic speech editing. The model uses a dedicated Audio Encoder for recognition and a RedAE‑based pathway for generation, with the LLM directly generating text or conditioning a flow‑matching DiT to produce acoustic latents. Evaluations show competitive or leading performance in multilingual ASR, content‑accurate zero‑shot TTS, strong instruction following, and significant improvements in speech editing over prior work.

By Feiyu Shen, Fenglong Xie, Junjie Li, Kun Xie, Lei Xie, Xu Tang, Xuelong Geng, Yan Jia, Yao Hu, Yichen Han, Yichen Wu, Ziqi Dai, Junjie Chen, Kai Huang, Manzhen Wei, Yixuan Li
arXiv Computer Vision
Sep 23

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.

By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
arXiv Machine Learning
Aug 27

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

LibriBrain100 is a new large‑scale MEG dataset for speech decoding that contains over 100 hours of high‑quality recordings while subjects listened to naturalistic continuous speech. The dataset more than doubles the size of the original LibriBrain release, with a record 80 hours from a single subject and additional 40‑minute recordings from 32 subjects. The authors demonstrate the value of deep within‑subject data and broad multi‑subject data by achieving state‑of‑the‑art word‑classification performance and showing that supervised fine‑tuning can compensate for limited per‑subject data, all supported by open‑source tools and a public competition leaderboard.

By Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran \"Ozdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones