arXiv Computer Vision
Sep 17

Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge Evaluation Plan

arXiv:2609.17913v1 Announce Type: new Abstract: Face--voice association models may rely on language or gender cues in the voice rather than on speaker-specific voice characteristics, which can lead t...

By Marta Moscati, Swapnil Khandoker, Muhammad Saad Saeed, Shah Nawaz, Fatima Noor, Rohan Kumar Das, Mubashir Noman, Junaid Mir, Muhammad Haroon Yousaf, Khalid Malik, Markus Schedl
arXiv AI
Aug 6

MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

arXiv:2608. 04433v1 Announce Type: cross Abstract: We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages.

By Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen, Ying Lay Chiu, Yang Ding, Yingxu He, Ridong Jiang, Zhuohan Liu, Yanfeng Lu, Yi Ma, Muhammad Huzaifah, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pham Minh Duc, Sailor Hardik Bhupendra, Siti Umairah Binte Mohammad Salleh, Shuo Sun, Tarun Kumar Vangani, Jeremy H. M. Wong, Jinyang Wu, Longyin Zhang
arXiv AI
Sep 11

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

The study investigates how speech‑to‑speech (S2S) models handle gender, distinguishing between the acoustic voice and the content’s gender cues. Experiments across five models in English, Spanish, and Mandarin show that while the rendered voice remains unbiased, the models consistently attribute speaker gender based on textual content rather than voice. When content and voice disagree, misgendering rates soar to 90%, whereas agreement yields only 2% misgendering.

By Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji