The study investigates how speech‑to‑speech (S2S) models handle gender, distinguishing between the acoustic voice and the content’s gender cues. Experiments across five models in English, Spanish, and Mandarin show that while the rendered voice remains unbiased, the models consistently attribute speaker gender based on textual content rather than voice. When content and voice disagree, misgendering rates soar to 90%, whereas agreement yields only 2% misgendering.
By Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji
arXiv:2510. 25577v2 Announce Type: replace-cross Abstract: Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation.
By Harm Lameris, Shree Harsha Bokkahalli Satish, Joakim Gustafson, \'Eva Sz\'ekely
arXiv:2609.16366v1 Announce Type: cross
Abstract: When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., "sh...
By Yingjia Wan, Lin Lin, Elisa Kreiss
The study introduces GAPA, a dataset of 316 physical attributes with 14,706 gender-association ratings from 304 US annotators, showing that such descriptions carry structured gender associations. It evaluates 16 LLMs, finding they partially mirror human ratings but exhibit biases such as compressed distributions, weaker alignment for men, and asymmetric abstention toward non‑binary identities. A proxy model trained on these data is released and applied to analyze character descriptions in LitBank, illustrating the persistence of gendered interpretations in ostensibly neutral language.
By Yingjia Wan, Lin Lin, Elisa Kreiss
arXiv:2609.17913v1 Announce Type: new
Abstract: Face--voice association models may rely on language or gender cues in the voice rather than on speaker-specific voice characteristics, which can lead t...
By Marta Moscati, Swapnil Khandoker, Muhammad Saad Saeed, Shah Nawaz, Fatima Noor, Rohan Kumar Das, Mubashir Noman, Junaid Mir, Muhammad Haroon Yousaf, Khalid Malik, Markus Schedl
Face--voice association models may rely on language or gender cues in the voice rather than on speaker-specific voice characteristics, which can lead to a performance deterioration when the model has...
arXiv:2606. 09589v1 Announce Type: cross Abstract: AI minidramas (also known as fruit dramas) are short, algorithmically distributed generative AI video series featuring anthropomorphized characters that have recently emerged as a widespread phenomenon on social media platforms.
By Piera Riccio
arXiv:2605. 08093v2 Announce Type: replace-cross Abstract: The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emotional dependence, and psychological harm.
By Maribeth Rauh, Dick A. H. Blankvoort, Matias Duran, Caoilfhionn N\'i Dheor\'ain, Harshvardhan J. Pandit, Syrine Enneifer, Siddharth D. Jaiswal, Anthony Ventresque, Abeba Birhane
arXiv:2609.13168v1 Announce Type: cross
Abstract: Speech AI, any AI system that recognizes, transforms, or generates speech, is built and evaluated across two communities with only a small overlap: t...
By Maria Teleki, Kimi Wenzel, Anna Seo Gyeong Choi, Tobias Weinberg, Shree Harsha Bokkahalli Satish, Stephanny Sanchez, Belu Ticona, Ariadna Sanchez, Yash Sonkar, Aarti Mathur, Christoph Minixhofer, Abraham Glasser, Raja Kushalnagar, James Caverlee, Minha Lee, Shaomei Wu, Alyssa Hillary Zisk, \'Eva Sz\'ekely, Dylan Gaines, Angelika Seeschaaf Veres, Seray Ibrahim, Nicholas Cummins, Allison Koenecke
The paper introduces a German-English benchmark dataset to evaluate anti‑LGBTQ biases in language models, combining community‑sourced stereotypes from German‑speaking queer individuals with a German translation of WinoQueer. Eight language models of varying sizes and architectures were assessed, revealing that they reproduce anti‑queer stereotypes with differences across identities and models. Fine‑tuning on community and progressive media content reduced bias on average, though the effect was not consistent across all models and identities.
By Melina Morch, Daniel Braun
arXiv:2502.10577v2 Announce Type: replace-cross
Abstract: Instruct-based large language models (LLMs) have been shown to propagate and even amplify gender bias when prompted with contextually constra...
By Enzo Doyen, Amalia Todirascu
The paper introduces WinoQueer‑NL, a Dutch adaptation of the English WinoQueer benchmark, designed to assess anti‑queer bias in Dutch language models. After validating the dataset with 43 queer Dutch participants, the authors released 42,906 sentences and evaluated several Dutch‑specific and multilingual models, finding that while overall bias scores were neutral, certain models disproportionately favored stereotypical statements for transgender and non‑binary identities. The study underscores the need for culturally grounded datasets to identify and mitigate biases that affect marginalized groups in Dutch NLP systems.