arXiv Computation and Language

Beyond Short Segments : Expanding Speaker Embeddings with Vector Archives

arXiv AI
Jun 18

Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation

arXiv:2603. 10827v2 Announce Type: replace-cross Abstract: Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity.

By Thomas Thebaud, Yuzhe Wang, Laureano Moro-Velazquez, Jesus Villalba-Lopez, Najim Dehak
arXiv AI
Jul 15

An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge

arXiv:2607. 12468v1 Announce Type: cross Abstract: We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time.

By Shuming Fang, Shuifei Zeng
arXiv Machine Learning
Aug 28

Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units

The paper introduces ABX-Accent, a benchmark built on the AESRC dataset that evaluates how well representation learning models adapt to 10 different English accents with less than 10 hours of unlabeled data per accent. It adapts the Zero Resources Challenge ABX metrics for each accent and demonstrates a baseline using adaptive domain normalization to fine‑tune a Contrastive Predictive Coding model, achieving a 23.6% relative improvement on across‑speaker ABX scores compared to non‑adapted models. The dataset and evaluation metrics will be released publicly after the paper is accepted.

By Robin San Roman, Manel Khentout, Tu Anh Nguyen, Paul Michel, Yossi Adi, Emmanuel Dupoux
arXiv AI
Jul 14

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

arXiv:2607. 10191v1 Announce Type: cross Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely.

By Shuhai Peng, Jinjiang Liu, Hui Lu, Liyang Chen, Guiping Zhong, Jiakui Li, Shiyin Kang, Zhiyong Wu
arXiv Computation and Language
Sep 21

Transsion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge

The paper describes Transsion Speech Team’s submission to Task 1 of the MLC‑SLM 2026 Challenge, aiming at speaker‑attributed transcription for multilingual conversational speech. Their cascaded framework includes a DiariZen‑based speaker diarization module, a Qwen3‑Omni‑based long‑form multilingual ASR module with CTC alignment for precise timestamps, and a fusion module that merges diarization and transcription outputs into speaker‑attributed STM results. On the official evaluation set, the system achieved a tcpMER of 15.41% and secured second place among all participants.

By Zhecheng Ren, Xuanji He, Xiaoxiao Li, Zhichen Han, Gaoyang Dong, Gaosheng Zhang, Minchuan Chen, Fengjie Zhu