The paper investigates gender bias in machine translation evaluation metrics using an occupation-balanced subset of GAMBIT+ across seven English‑source language pairs, including a new German extension. It finds that masculine translations tend to receive higher scores and that biases align with stereotypical gender representations, though the strength varies by evaluator and language. The study highlights that assessing bias requires multiple dimensions beyond a single aggregate measure.
By Orfeas Menis Mastromichalakis, Giorgos Filandrianos, Wafaa Mohammed, Giuseppe Attanasio, Chrysoula Zerva
arXiv:2608. 04433v1 Announce Type: cross Abstract: We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages.
By Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen, Ying Lay Chiu, Yang Ding, Yingxu He, Ridong Jiang, Zhuohan Liu, Yanfeng Lu, Yi Ma, Muhammad Huzaifah, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pham Minh Duc, Sailor Hardik Bhupendra, Siti Umairah Binte Mohammad Salleh, Shuo Sun, Tarun Kumar Vangani, Jeremy H. M. Wong, Jinyang Wu, Longyin Zhang
arXiv:2606. 30152v1 Announce Type: cross Abstract: Contextual language models conflate grammatical gender and social semantic bias in gendered languages such as Spanish.
By Huanping Xiao, Yingji Li
The paper introduces a unified framework that simultaneously measures intrinsic (encoded) and extrinsic (expressed) gender bias in large language models using identical neutral prompts. It finds a consistent link between latent gender information and output bias, but shows that alignment via supervised fine‑tuning reduces expressed bias while leaving internal gender associations largely intact and reactivatable by adversarial prompts. The study also demonstrates that debiasing gains on structured benchmarks may not transfer to realistic tasks such as story generation.
By Nour Bouchouchi, Thibault Laugel, Xavier Renard, Christophe Marsala, Marie-Jeanne Lesot, Marcin Detyniecki
arXiv:2606. 31718v1 Announce Type: cross Abstract: Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotated corpora.
By Dragos-Mitrut Vasile, Elena-Simona Apostol, Stefan-Adrian Toma, Adrian Paschke, Ciprian-Octavian Truica
PolERo presents a new dataset of 3,574 Romanian question‑answer pairs from presidential transcripts, annotated for political evasion using a two‑level taxonomy of response clarity and fine‑grained evasion strategies. The study evaluates various classification methods—including TF‑IDF baselines, fine‑tuned encoders, a sliding‑window encoder, and zero/few‑shot LLM prompting—under matched conditions. Cross‑lingual transfer experiments via joint bilingual training and machine‑translation augmentation reveal that fine‑tuned encoders perform competitively, transfer is asymmetric, and ambivalent evasion categories with pragmatic cues remain the most challenging across all models.
By Gabriel Stefan, Sergiu Nisioi
arXiv:2502. 11603v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit strong natural language understanding capabilities but also inherit and amplify societal biases, particularly gender bias, raising fairness concerns.
By Hongye Qiu, Yue Xu, Yi Wang, Meikang Qiu, Wenjie Wang
arXiv:2603. 23485v2 Announce Type: replace-cross Abstract: Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalent discourses.
By Sagar Kumar, Ariel Flint, Luca Maria Aiello, Andrea Baronchelli
Fine‑tuning large language models on parallel data can improve translation quality but also causes catastrophic forgetting of general capabilities. The study evaluates several forgetting‑mitigation methods—anchored to auxiliary data, model outputs, and base model parameters—using Llama 3.2 1B Instruct and Llama 3.1 8B Instruct on Arabic‑English and Spanish‑English translation tasks. Elastic Weight Consolidation best preserves general benchmark performance, yet only data mixing with control‑task examples maintains instruction‑following abilities such as formality and grammatical gender control, though these gains do not generalize to unseen prompts.
By Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney
arXiv:2509. 07829v4 Announce Type: replace-cross Abstract: Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small open models remains an open problem, particularly for low-resource languages such as Romanian.
By Mihai Nadas, Laura Diosan, Andreea Tomescu, Andrei Piscoran
arXiv:2607. 20073v1 Announce Type: new Abstract: AI-based recruitment systems that rely on machine learning models trained on historical CV data, risk perpetuating and amplifying social biases.
By Farnaz Faramarzi Lighvan, Lynn Houthuys
arXiv:2510. 07074v2 Announce Type: replace-cross Abstract: Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts.
By Fred Philippy, Laura Bernardy, Siwen Guo, Jacques Klein, Tegawend\'e F. Bissyand\'e